SAVRN
Search Contact SAVRN

Open-weight model · Image classification

CLIP-ViT-large-patch14-FT-SUN397

by Yanggangu yanggangu/CLIP-ViT-large-patch14-FT-SUN397

CLIP-ViT-large-patch14-FT-SUN397 is an open-weight model for image classification from Yanggangu, released under MIT License. Its published files total 1.2 GB.

A FT expert from Table 2 of SMAT: Simple and Efficient Merge-Aware Training (seed 42). SMAT trains experts with model merging in mind.

Parameters—
Context—
Weights1.2 GB
Licensemit
AccessOpen weights
Monthly Downloads—

Model Card

By Yanggangu, published under mit, revision 1d9994200b29.

A FT expert from Table 2 of SMAT: Simple and Efficient Merge-Aware Training (seed 42). SMAT trains experts with model merging in mind. encoder.pt is the original FP32 vision-encoder state dictionary; load it with the SMAT code and the matching OpenAI CLIP base model. Dataset terms apply separately.

Read Yanggangu's full model card

A FT expert from Table 2 of SMAT: Simple and Efficient Merge-Aware Training (seed 42). SMAT trains experts with model merging in mind.

Paper · GitHub & usage

encoder.pt is the original FP32 vision-encoder state dictionary; load it with the SMAT code and the matching OpenAI CLIP base model. Dataset terms apply separately.

Identity and Version

Repository
yanggangu/CLIP-ViT-large-patch14-FT-SUN397
Publisher
Yanggangu
Task
Image classification
Modality
Image
Library
Not stated by the source
Parameters
Not stated by the source
Languages
Not stated by the source
Revision
1d9994200b29f5fffe9662babd98c4984d41e5dd
First published
2026-09-29
Last updated
2026-09-30

Files and Weights

5 files, 1.2 GB in total. The weights are 1 file totalling 1.2 GB in pt.

Weights1 file · 1.2 GB
Configuration1 file · 412 B
Documentation2 files · 1.7 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
encoder.ptWeights1.2 GB 71469c0a9b54
metadata.jsonConfiguration412 B —
LICENSEDocumentation1.1 KB —
README.mdDocumentation610 B —
.gitattributesRepository1.5 KB —

License and Download

License
mit
Access
Open weights, no gate
Download size
1.2 GB
Download from Yanggangu

Released by Yanggangu through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published1.2 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About CLIP-ViT-large-patch14-FT-SUN397

Can I use CLIP-ViT-large-patch14-FT-SUN397 commercially?

Yes. CLIP-ViT-large-patch14-FT-SUN397 is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

Similar Models

Model · Image classification

swinv2-tiny-patch4-window16-256

Microsoft

Swin Transformer v2 model pre-trained on ImageNet-1k at resolution 256x256. It was introduced in the paper Swin Transformer V2: Scaling Up Capacity and Resolution by Liu et al. and first released in this repository. Disclaimer: The team releasing Swin Transformer v2 did not write a model card for this model so this model card has been written by the Hugging Face team. The Swin Transformer is a type of Vision Transformer. It builds hierarchical feature maps by merging image patches (shown in gray) in deeper layers and has linear computation complexity to input image size due to computation of self-attention only within each local window (shown in red). It can thus serve as a general-purpose…

Open weights apache-2.0 transformers

Model · Image classification

AI-image-detector

Matthew Maybe

NOTE: Unless you are trying to detect imagery generated using older models such as VQGAN+CLIP, please use the updated version of this detector instead. This model is a proof-of-concept demonstration of using a ViT model to predict whether an artistic image was generated using AI. It was created in October 2022, and as such, the training data did not include any samples generated by Midjourney 5, SDXL, or DALLE-3. It still may be able to correctly identify samples from these more recent models due to being trained on outputs of their predecessors. Furthermore the intended scope of this tool is artistic images; that is to say, it is not a deepfake photo detector, and general computer imagery…

Open weights cc-by-4.0 transformers

Measured on device (edge-compat): Raspberry Pi 5 · LiteRT 2.2.0.dev20260804 · CPU/XNNPACK, 4 threads · 34.1 ms p50 (2026-08-31); browser · Chromium 151 (2026-08-11). Record: https://github.com/john-rocky/edge-compat/blob/main/cards/plantnet-300k-resnet18/CARD.md On-device fine-grained plant species identification — 1081 species — running fully on the LiteRT CompiledModel GPU delegate (no CPU fallback). A PlantNet-300K (NeurIPS 2021) ResNet18. ~16 ms/frame on a Pixel 8a. (mean [0.485,0.456,0.406], std [0.229,0.224,0.225]; center-crop then resize 224). Labels: class index i maps to the i-th species when the PlantNet-300K species-id strings are sorted (torchvision ImageFolder order); names…

Open weights apache-2.0 litert

Model · Image classification

traffic-sign-adverse-weather

Yy

Official model checkpoints for the solution in the Traffic Sign Recognition under Adverse Weather Competition. See classes.txt for the 25 traffic sign classes. For inference scripts, training code, and in-depth engineering retrospective, visit the GitHub Repository.

Open weights mit timm

Model · Image classification

tinyvit-5m-int8-imagenet

Core Epoch

TinyViT-5M (timm/tinyvit5m224.distin22kftin1k, Apache-2.0) quantized to INT8 with Kenosis — 128-image calibration, no retraining. 80.53% top-1 from a 9.2 MB single file, on ONNX Runtime or OpenVINO, CPU or GPU, no accelerator required. ImageNet-1K validation, 49,872 images (disjoint from the 128 calibration images). Measured on a CPU with AVX-VNNI; on CPUs without VNNI this model's INT8 top-1 sits ~0.9 below FP32 rather than 0.34. Input 1x3x224x224, RGB, /255, ImageNet mean/std. Output logits [1,1000], sorted-synset order. runclassify.py / evalimagenet.py reproduce the demo and table. tinyvit5m224int8kenosis.onnx (9,228,567 B) — SHA-256…

Open weights apache-2.0 onnx