SAVRN
Search Contact SAVRN

Open-weight model · Text classification

Jev-LCT-Qwen2.5-1.5B

by CaoHaoWei CaoHaoWei/Jev-LCT-Qwen2.5-1.5B

Jev-LCT-Qwen2.5-1.5B is an open-weight model for text classification from CaoHaoWei, released under Apache License 2.0. It has 1.5B parameters and a 131,072-token context. At 16-bit it needs about 3.7 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

In agentic workflows, API routing, and edge decision-making, conventional autoregressive LLMs suffer from high token-by-token generation latency and brittle string parsing, while small discriminative models produce systematically miscalibrated verbal…

Parameters1.5B
Context131,072
Weights3.1 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve Jev-LCT-Qwen2.5-1.5B (1.5B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 3.1 GB 3.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 1.5 GB 1.9 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.8 GB 0.9 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.

Jev-LCT-Qwen2.5-1.5B on every accelerator the SAVRN Index prices, at every precision

Model Card

By CaoHaoWei, published under apache-2.0, revision 5e86eb6acab2.

In agentic workflows, API routing, and edge decision-making, conventional autoregressive LLMs suffer from high token-by-token generation latency and brittle string parsing, while small discriminative models produce systematically miscalibrated verbal confidence (overconfident hallucinations). Jev-LCT (Looped Calibration Transformer) establishes a new paradigm for System-One Decision Models: 1. Parallel Looped Prefill: Recurrently iterates only the top $k=2$ layers of Qwen2.5-1.5B with sequence right-shifting and Scale-Preserving RMS Injection, strictly preventing representation collapse. 2. Endogenous Trajectory Confidence: Extracts genuine calibrated confidence directly from hidden state…

Read CaoHaoWei's full model card

Jev-LCT-Qwen2.5-1.5B: Open System-One Decision Engine

"Decisions, Not Strings" meets "Free Calibrated Confidence"
Free Calibrated Confidence from Recurrent Computation Trajectories for Small Decision Models
《Jev-LCT:下一代开源系统一决策引擎——从小模型循环计算轨迹中免费提取校准置信度》


English Description

1. Key Innovations

In agentic workflows, API routing, and edge decision-making, conventional autoregressive LLMs suffer from high token-by-token generation latency and brittle string parsing, while small discriminative models produce systematically miscalibrated verbal confidence (overconfident hallucinations).

Jev-LCT (Looped Calibration Transformer) establishes a new paradigm for System-One Decision Models: 1. Parallel Looped Prefill: Recurrently iterates only the top $k=2$ layers of Qwen2.5-1.5B with sequence right-shifting and Scale-Preserving RMS Injection, strictly preventing representation collapse. 2. Endogenous Trajectory Confidence: Extracts genuine calibrated confidence directly from hidden state convergence rates ($\Delta\cos$), decision entropy reduction ($\Delta H$), and softmax margins—without requiring Reinforcement Learning (RLCD) or verbal introspection tokens. - Error Detection AUROC reaches 0.9043 (In-Distribution) and 0.9524 (Out-of-Distribution). - Expected Calibration Error (ECE-15) drops from 0.1646 to 0.1259 (a 23.5% relative reduction). 3. Dual-Channel Adaptive Early Exit: - Confidence Margin Gate ($m_1 \ge 0.90$): 85%+ of simple queries exit on loop 1 at single-pass latency (~45ms), completely eliminating overthinking. - Cosine Attractor Stability ($\Delta_t < 0.30$): Ambiguous problems dynamically iterate up to 4 loops before terminating cleanly. - Deterministic SLA Guarantee: Keeps average latency at 61.9 ms on an RTX 3090 Ti.


2. Comprehensive 300-Item Benchmark (NVIDIA RTX 3090 Ti Verified)

Cross-domain benchmark comparison across financial intent (Banking77), scientific reasoning (AI2 ARC), factual verification (TruthfulQA), reading comprehension (BoolQ), and multi-task academic reasoning (MMLU):

Model Architecture Parameters Intent (Banking77) Science (ARC) Factuality (TQA) Reading (BoolQ) Academic (MMLU) Overall Accuracy Avg Latency Avg Loops
Convai Laya Base (ModernBERT) 421M 95.0% 28.3% 23.3% 76.7% 33.3% 51.3% 49.3 ms 1.00 (Single-Pass)
Convai Laya Typed (ModernBERT) 421M 95.0% 30.0% 18.3% 81.7% 31.7% 51.3% 52.9 ms 1.00 (Single-Pass)
Open-Jev (DeBERTa-v3) 435M 100.0% 45.0% 25.0% 91.7% 35.0% 59.3% 59.7 ms 1.00 (Single-Pass)
Qwen-0.5B Adaptive LCT 0.49B 91.7% 40.0% 13.3% 53.3% 50.0% 49.7% 50.8 ms 1.36 loops
Jev-LCT-Qwen2.5-1.5B (This Model) 1.54B 93.3% 71.7% 55.0% 73.3% 58.3% 70.3% 61.9 ms 1.23 loops
Jev-LCT-Qwen3-8B 7.61B 95.0% 85.0% 75.0% 83.3% 70.0% 81.7% 89.2 ms 1.09 loops

Takeaway: At an ultra-fast 61.9 ms latency, Jev-LCT-1.5B surpasses Convai Laya by +41.7% and Open-Jev by +26.7% on scientific reasoning (ARC), delivering generative-grade reasoning at discriminative-grade speed.


3. Quickstart (3 Lines of Code)

This repository contains fully merged standalone weights (no separate adapter loading needed):

from lct_qwen_standalone import LCTQwen

# 1. Load standalone model (bfloat16 on CUDA)
model = LCTQwen.from_pretrained("CaoHaoWei/Jev-LCT-Qwen2.5-1.5B")

# 2. Predict multiple-choice decision
result = model.predict_choice(
    prompt="Patient reports severe chest pain radiating to left shoulder. Determine triage urgency level:",
    choices=["emergency", "urgent", "routine", "elective"]
)

print(f"Top Decision: {result['choice']}")
print(f"Calibrated Confidence: {result['confidence']:.2%}")
print(f"Loops Executed: {result['loops']}")

中文详细说明

核心亮点

  1. 系统一决策新范式:结合 TypeSafe AI 的 Jev 理念,摆脱生成文本的高延迟,以毫秒级直出结构化决策与校准置信度。
  2. 通识与科学推理大幅领先:在 ARC 科学推理达到 71.7%(远超 Laya 30.0% 与 Open-Jev 45.0%),MMLU 达 58.3%。
  3. 免 RL 零额外开销校准:隐状态动力学余弦收敛与熵变轨迹特征,检错 AUROC 达 0.904 ~ 0.952,ECE 校准误差降低 23.5%。
  4. 自适应早停秒级直出:简单题单遍直出(~45ms),复杂题深思 2~4 轮,平均延迟仅 61.9 ms。

Citation & License

Apache License 2.0. Please cite as:

@article{cao2026lct,
  title={Looped Calibration Transformer: Free Calibrated Confidence from Recurrent Computation Trajectories for Small Decision Models},
  author={Cao, Haowei},
  year={2026},
  publisher={GitHub},
  journal={GitHub repository},
  howpublished={\url{https://github.com/gitchw/LCT}}
}

Configuration

Architecture
Qwen2ForCausalLM
Context length (tokens)
131,072
Layers
28
Hidden size
1,536
Feed-forward size
8,960
Attention heads
12
Key/value heads
2
Vocabulary size
151,936
RoPE base
1e+06
Model type
qwen2

Identity and Version

Repository
CaoHaoWei/Jev-LCT-Qwen2.5-1.5B
Publisher
CaoHaoWei
Task
Text classification
Modality
Text
Library
transformers
Parameters
1.5B parameters
Languages
en, zh
Revision
5e86eb6acab29859b196597d98cde3f6151fce30
First published
2026-09-25
Last updated
2026-09-26

Files and Weights

14 files, 3.1 GB in total. The weights are 2 files totalling 3.1 GB in pt, safetensors.

Weights2 files · 3.1 GB
Configuration5 files · 3.2 KB
Tokenizer4 files · 15.9 MB
Documentation1 file · 6.5 KB
Other1 file · 2.5 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
lct_auxiliary.ptWeights20.9 KB 7277d1c11de0
model.safetensorsWeights3.1 GB 087d2804476a
added_tokens.jsonConfiguration629 B —
config.jsonConfiguration1.4 KB —
generation_config.jsonConfiguration123 B —
lct_config.jsonConfiguration399 B —
special_tokens_map.jsonConfiguration647 B —
README.mdDocumentation6.5 KB —
chat_template.jinjaOther2.5 KB —
.gitattributesRepository1.6 KB —
merges.txtTokenizer1.7 MB —
tokenizer.jsonTokenizer11.4 MB 9c5ae00e602b
tokenizer_config.jsonTokenizer4.9 KB —
vocab.jsonTokenizer2.8 MB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
3.1 GB
Download from CaoHaoWei

Released by CaoHaoWei through its official repository on Hugging Face. Read the license.

Built From

  • Trained on (disclosed) ai2_arc
  • Trained on (disclosed) banking77
  • Trained on (disclosed) boolq
  • Trained on (disclosed) mmlu
  • Trained on (disclosed) truthful_qa

Memory Requirements

PrecisionWeights in memory
As published3.1 GB
16-bit3.1 GB
8-bit1.5 GB
4-bit0.8 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Jev-LCT-Qwen2.5-1.5B

How much GPU memory does Jev-LCT-Qwen2.5-1.5B need?

About 3.7 GB at 16-bit and 0.9 GB at 4-bit: the weights (1.5B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Jev-LCT-Qwen2.5-1.5B on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Jev-LCT-Qwen2.5-1.5B commercially?

Yes. Jev-LCT-Qwen2.5-1.5B is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is Jev-LCT-Qwen2.5-1.5B's context length?

131,072 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text classification

decider-2b-fp8

LLM Tech

Mapika/decider-2b v11 quantized to FP8 for vLLM: FP8 E4M3 weights with one scale per output channel and FP8 activations scaled per token at run time. 2.39 GB against 3.77 GB for the bf16 checkpoint. Quantized and measured by LLM Tech; the model, its training and its evaluation protocol are Mapika's. Read the bf16 card for what the model is and how it was trained. The base revision is 533964dae8be954c5b5e19fa4948e48408094c1e. Tokenizer, chat template, generation config and deciderconfig.json (temperatures included) are the author's files unchanged, apart from the version and quantization fields. Both models were run through vLLM 0.29.0 on the same rows: the author's regression set rebuilt…

Open weights apache-2.0 1.9B parameters 262,144 tokens

Model · Text classification

openjev-v5-0.8b

Rodney Lafuente-Mercado

This is an unchanged mirror of AlexWortega's pretrained OpenJev v5 0.8B model, published so adapters can identify and load this specific base without confusing it with the 2B and 4B models in the upstream repository. All model, tokenizer, and configuration files are byte-for-byte copies. I did not train this base model. The source is AlexWortega/openjev, qwen3.5-0.8b-nli-v5, pinned to revision 552759daad712f1af6c4c13dabcb1e047886fc9c. OpenJev turns the Qwen3.5-0.8B backbone into a three-class natural language inference classifier. Credit for the base model and its training belongs to AlexWortega and the Qwen team. See the upstream model card for the method and reported evaluations. The…

Open weights mit 853M parameters 262,144 tokens transformers

Model · Text classification

decider-0.8b-fp8

LLM Tech

Mapika/decider-0.8b v1 quantized to FP8 for vLLM: FP8 E4M3 weights with one scale per output channel and FP8 activations scaled per token at run time. 1.01 GB against 1.5 GB for the bf16 checkpoint. Quantized and measured by LLM Tech; the model, its training and its evaluation protocol are Mapika's. Read the bf16 card for what the model is and how it was trained. The base revision is a0a01d6f8135298f400a8c856b355793012ae971. Tokenizer, chat template, generation config and deciderconfig.json (temperatures included) are the author's files unchanged, apart from the version and quantization fields. Both models were run through vLLM 0.29.0 on the same rows: the author's regression set rebuilt…

Open weights apache-2.0 752M parameters 262,144 tokens

Model · Text classification

decider-4b-nvfp4

LLM Tech

Mapika/decider-4b v2.1 quantized to NVFP4 for vLLM: 4-bit floating-point weights and activations with FP8 block scales (block size 16). 3.29 GB against 8.41 GB for the bf16 checkpoint. Quantized and measured by LLM Tech; the model, its training and its evaluation protocol are Mapika's. Read the bf16 card for what the model is and how it was trained. The base revision is eb5fbdfc9448473ec25e399882912863afbdb70e. Tokenizer, chat template, generation config and deciderconfig.json (temperatures included) are the author's files unchanged, apart from the version and quantization fields. Both models were run through vLLM 0.29.0 on the same rows: the author's regression set rebuilt from public data…

Open weights apache-2.0 2.4B parameters 262,144 tokens

Model · Text classification

jina-reranker-m0

Jina AI

pipelinetag: text-classification - sentence-transformers - vidore - reranker - qwen2vl - multilingual basemodel: libraryname: transformers jina-reranker-m0 is our new multilingual multimodal reranker model for ranking visual documents across multiple languages: it accepts a query alongside a collection of visually rich document images, including pages with text, figures, tables, infographics, and various layouts across multiple domains and over 29 languages. It outputs a ranked list of documents ordered by their relevance to the input query. Compared to jina-reranker-v2-base-multilingual, jina-reranker-m0 also improves text reranking for multilingual content, long documents, and code…

Open weights cc-by-nc-4.0 2.4B parameters 32,768 tokens transformers

Model · Text classification

Zircon-0.6B-v2-mlx

Fahrenheit Research

Zircon v2 is a 0.6B-parameter decision model from Fahrenheit Research. It runs fully on-device on Apple silicon. You give it an email, a message or a pending task plus a set of options, and it returns a calibrated probability for each option in under 50 ms per decision. (1) MacBook Pro (Apple M5), 8-bit weights, median. Speed varies by hardware. (2) Fahrenheit Research internal testing, September 2026, on held-out emails not seen in training. (3) Fahrenheit Research internal testing, September 2026. 400 cases (2,000 decisions) from the public LocalLLaMA typed-decisions test set. (4) Fahrenheit Research internal testing, September 2026, on game seeds not seen in training. Built on Qwen3-0.6B…

Open weights apache-2.0 596M parameters 40,960 tokens mlx