SAVRN
Search Contact SAVRN

Open-weight model · Zero-shot classification

aplomb-1

by EmpirioLabs AI empiriolabsai/aplomb-1

aplomb-1 is a model for zero-shot classification from EmpirioLabs AI, released under other (access requested at publisher). It has 5.3B parameters. At 16-bit it needs about 12.7 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 30 downloads a month.

Aplomb 1 is our first model, a decision model. It reads text, JSON, images, video and audio, and answers typed questions about them with calibrated probabilities instead of generated text.

Parameters5.3B
Context—
Weights12.0 GB
Licenseother
AccessAccess requested at publisher
Monthly Downloads30

Runs On

What it takes to serve aplomb-1 (5.3B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 10.6 GB 12.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 5.3 GB 6.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 2.6 GB 3.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

aplomb-1 on every accelerator the SAVRN Index prices, at every precision

Model Card

Aplomb 1 is our first model, a decision model. It reads text, JSON, images, video and audio, and answers typed questions about them with calibrated probabilities instead of generated text. Send a state of up to 1M tokens with up to 128 questions, and every answer comes back as a probability distribution with a confidence value, in one request. Aplomb 1 has 5.3 billion parameters. - Every input type in one model: text, JSON objects and arrays, images, video and audio. selection (which function to call, with distributions over its enum and boolean arguments). contain what the question needs. - Zero data retention by default on the hosted API: EmpirioLabs does not retain the content of…

Excerpt from the card by EmpirioLabs AI, licensed other.

Identity and Version

Repository
empiriolabsai/aplomb-1
Publisher
EmpirioLabs AI
Task
Zero-shot classification
Modality
Text
Library
transformers
Parameters
5.3B parameters
Languages
af, am, ar, az, bn, cy, da, de
Revision
6b0d39e0b53f9192cb04313e473ee40a466f8d17
First published
2026-09-29
Last updated
2026-10-07

Files and Weights

32 files, 12.0 GB in total. The weights are 4 files totalling 12.0 GB in safetensors.

Weights4 files · 12.0 GB
Configuration18 files · 172.3 KB
Tokenizer4 files · 30.1 MB
Documentation4 files · 35.0 KB
Other1 file · 7.8 KB
Repository1 file · 62 B
Every file
FileTypeSizeSHA-256
audio_attach.safetensorsWeights94.4 MB —
audio_encoder/model.safetensorsWeights1.3 GB —
model.safetensors-00001-of-00002.safetensorsWeights6.6 GB —
model.safetensors-00002-of-00002.safetensorsWeights4.0 GB —
aplomb/__init__.pyConfiguration62 B —
aplomb/adapter.pyConfiguration10.4 KB —
aplomb/audio_attach.pyConfiguration21.4 KB —
aplomb/compile.pyConfiguration2.7 KB —
aplomb/decision_head.pyConfiguration7.6 KB —
aplomb/readout.pyConfiguration3.0 KB —
aplomb/render.pyConfiguration6.2 KB —
aplomb/schema.pyConfiguration5.9 KB —
aplomb/targets.pyConfiguration2.1 KB —
aplomb/tools.pyConfiguration6.2 KB —
audio_config.jsonConfiguration1.8 KB —
audio_encoder/config.jsonConfiguration13.7 KB —
config.jsonConfiguration2.8 KB —
edm_config.jsonConfiguration3.3 KB —
model.safetensors.index.jsonConfiguration76.3 KB —
preprocessor_config.jsonConfiguration390 B —
run_aplomb.pyConfiguration8.4 KB —
video_preprocessor_config.jsonConfiguration385 B —
LICENSEDocumentation6.2 KB —
LICENSE-APACHE-2.0Documentation11.5 KB —
NOTICEDocumentation438 B —
README.mdDocumentation16.8 KB —
chat_template.jinjaOther7.8 KB —
.gitattributesRepository62 B —
merges.txtTokenizer3.4 MB —
tokenizer.jsonTokenizer20.0 MB —
tokenizer_config.jsonTokenizer1.1 KB —
vocab.jsonTokenizer6.7 MB —

License and Download

License
other
Access
Access requested at publisher
Download size
12.0 GB
Request access from EmpirioLabs AI

EmpirioLabs AI grants access through its official repository on Hugging Face.

Built From

Memory Requirements

PrecisionWeights in memory
As published12.0 GB
16-bit10.6 GB
8-bit5.3 GB
4-bit2.6 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About aplomb-1

How much GPU memory does aplomb-1 need?

About 12.7 GB at 16-bit and 3.2 GB at 4-bit: the weights (5.3B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run aplomb-1 on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

What license is aplomb-1 released under?

other, as its publisher declares it. Read the license text before commercial use.

Similar Models

Model · Zero-shot classification

Vela-2.0-4B

vLLM Semantic Router

Open Foundation Routing Models Routing decisions. Safety checks. Precise text spans. The balanced hybrid member of Vela 2.0: strong multilingual safety and general decisions, with router and open-label spans through one interface. Define options, labels and rubrics at request time. Ask multiple named questions about a request, context and answer, and receive structured decisions with the text spans that support your workflow. 1. Balanced routing capability. Safety macro AUC of 0.921 across 14 public sets, alongside open-label extraction and general decisions. 2. Routing and safety together. Use one request for routing, prompt-attack checks, PII and unsupported-claim detection. 3. Decisions…

Open weights apache-2.0 4.2B parameters

Model · Zero-shot classification

LLM2CLIP-Llama-3-8B-Instruct-CC-Finetuned

Microsoft

Weiquan Huang 1, Aoqi Wu 1, Yifan Yang 2†, Xufang Luo 2, Yuqing Yang 2, Liang Hu 1, Qi Dai 2, Xiyang Dai 2, Dongdong Chen 2, Chong Luo 2, Lili Qiu 2 In this paper, we propose LLM2CLIP, a novel approach that embraces the power of LLMs to unlock CLIP’s potential. By fine-tuning the LLM in the caption space with contrastive learning, we extract its textual capabilities into the output embeddings, significantly improving the output layer’s textual discriminability. We then design an efficient training process where the fine-tuned LLM acts as a powerful teacher for CLIP’s visual encoder. Thanks to the LLM’s presence, we can now incorporate longer and more complex captions without being…

Open weights apache-2.0 7.5B parameters 8,192 tokens

Model · Zero-shot classification

Vela-2.0-9B

vLLM Semantic Router

Open Foundation Routing Models Routing decisions. Safety checks. Precise text spans. The largest hybrid member of Vela 2.0: high-accuracy multilingual routing, long-document PII and hallucination detection, with open-label spans through one interface. Define options, labels and rubrics at request time. Ask multiple named questions about a request, context and answer, and receive structured decisions with the text spans that support your workflow. 1. High-accuracy decisions. A 41.63 Jev Decision Index with optional Noul calibration, plus 0.989 AUC on unseen prompt-attack families. 2. Routing and safety together. Use one request for routing, prompt-attack checks, PII and unsupported-claim…

Open weights apache-2.0 7.9B parameters

Model · Zero-shot classification

Vela-2.0-0.8B

vLLM Semantic Router

Open Foundation Routing Models Routing decisions. Safety checks. Precise text spans. The compact hybrid member of Vela 2.0: a 756M-parameter model for multilingual routing, safety checks and span-level decisions through one interface. Define options, labels and rubrics at request time. Ask multiple named questions about a request, context and answer, and receive structured decisions with the text spans that support your workflow. 1. Compact deployment. About 3 GB of GPU memory for FP32 parameters, with a 16,384-token input limit. 2. Routing and safety together. Use one request for routing, prompt-attack checks, PII and unsupported-claim detection. 3. Decisions at span resolution. Return…

Open weights apache-2.0 755M parameters

Model · Zero-shot classification

OpenThai-SystemOne

iApp Technology

An open Thai + English "System One" decision model. It does not generate text. Given a state (any text or JSON) and typed questions, it returns calibrated probabilities over the options in one forward pass: The request/response contract mirrors TypeSafe's POST /v1/systemone so code written for the TypeSafe SDK can be pointed at this model unchanged. Typical uses: ticket routing, moderation, intent detection, RAG relevance judging, LLM-output verification, and computer-use / browser-agent action selection (which element to click, which tool to call). Gated-DeltaNet / attention, 262k context), vision encoder removed, then continued-pretrained on ~5B tokens of Thai (web, Wikipedia, parallel…

Open weights apache-2.0 753M parameters 262,144 tokens

Model · Zero-shot classification

LLM2CLIP-Openai-L-14-336

Microsoft

Weiquan Huang 1, Aoqi Wu 1, Yifan Yang 2†, Xufang Luo 2, Yuqing Yang 2, Liang Hu 1, Qi Dai 2, Xiyang Dai 2, Dongdong Chen 2, Chong Luo 2, Lili Qiu 2 In this paper, we propose LLM2CLIP, a novel approach that embraces the power of LLMs to unlock CLIP’s potential. By fine-tuning the LLM in the caption space with contrastive learning, we extract its textual capabilities into the output embeddings, significantly improving the output layer’s textual discriminability. We then design an efficient training process where the fine-tuned LLM acts as a powerful teacher for CLIP’s visual encoder. Thanks to the LLM’s presence, we can now incorporate longer and more complex captions without being…

Open weights apache-2.0 579M parameters