Jev-LCT-Qwen2.5-1.5B: Open System-One Decision Engine
"Decisions, Not Strings" meets "Free Calibrated Confidence"
Free Calibrated Confidence from Recurrent Computation Trajectories for Small Decision Models
《Jev-LCT:下一代开源系统一决策引擎——从小模型循环计算轨迹中免费提取校准置信度》
English Description
1. Key Innovations
In agentic workflows, API routing, and edge decision-making, conventional autoregressive LLMs suffer from high token-by-token generation latency and brittle string parsing, while small discriminative models produce systematically miscalibrated verbal confidence (overconfident hallucinations).
Jev-LCT (Looped Calibration Transformer) establishes a new paradigm for System-One Decision Models:
1. Parallel Looped Prefill: Recurrently iterates only the top $k=2$ layers of Qwen2.5-1.5B with sequence right-shifting and Scale-Preserving RMS Injection, strictly preventing representation collapse.
2. Endogenous Trajectory Confidence: Extracts genuine calibrated confidence directly from hidden state convergence rates ($\Delta\cos$), decision entropy reduction ($\Delta H$), and softmax margins—without requiring Reinforcement Learning (RLCD) or verbal introspection tokens.
- Error Detection AUROC reaches 0.9043 (In-Distribution) and 0.9524 (Out-of-Distribution).
- Expected Calibration Error (ECE-15) drops from 0.1646 to 0.1259 (a 23.5% relative reduction).
3. Dual-Channel Adaptive Early Exit:
- Confidence Margin Gate ($m_1 \ge 0.90$): 85%+ of simple queries exit on loop 1 at single-pass latency (~45ms), completely eliminating overthinking.
- Cosine Attractor Stability ($\Delta_t < 0.30$): Ambiguous problems dynamically iterate up to 4 loops before terminating cleanly.
- Deterministic SLA Guarantee: Keeps average latency at 61.9 ms on an RTX 3090 Ti.
2. Comprehensive 300-Item Benchmark (NVIDIA RTX 3090 Ti Verified)
Cross-domain benchmark comparison across financial intent (Banking77), scientific reasoning (AI2 ARC), factual verification (TruthfulQA), reading comprehension (BoolQ), and multi-task academic reasoning (MMLU):
| Model Architecture |
Parameters |
Intent (Banking77) |
Science (ARC) |
Factuality (TQA) |
Reading (BoolQ) |
Academic (MMLU) |
Overall Accuracy |
Avg Latency |
Avg Loops |
| Convai Laya Base (ModernBERT) |
421M |
95.0% |
28.3% |
23.3% |
76.7% |
33.3% |
51.3% |
49.3 ms |
1.00 (Single-Pass) |
| Convai Laya Typed (ModernBERT) |
421M |
95.0% |
30.0% |
18.3% |
81.7% |
31.7% |
51.3% |
52.9 ms |
1.00 (Single-Pass) |
| Open-Jev (DeBERTa-v3) |
435M |
100.0% |
45.0% |
25.0% |
91.7% |
35.0% |
59.3% |
59.7 ms |
1.00 (Single-Pass) |
| Qwen-0.5B Adaptive LCT |
0.49B |
91.7% |
40.0% |
13.3% |
53.3% |
50.0% |
49.7% |
50.8 ms |
1.36 loops |
| Jev-LCT-Qwen2.5-1.5B (This Model) |
1.54B |
93.3% |
71.7% |
55.0% |
73.3% |
58.3% |
70.3% |
61.9 ms |
1.23 loops |
| Jev-LCT-Qwen3-8B |
7.61B |
95.0% |
85.0% |
75.0% |
83.3% |
70.0% |
81.7% |
89.2 ms |
1.09 loops |
Takeaway: At an ultra-fast 61.9 ms latency, Jev-LCT-1.5B surpasses Convai Laya by +41.7% and Open-Jev by +26.7% on scientific reasoning (ARC), delivering generative-grade reasoning at discriminative-grade speed.
3. Quickstart (3 Lines of Code)
This repository contains fully merged standalone weights (no separate adapter loading needed):
from lct_qwen_standalone import LCTQwen
# 1. Load standalone model (bfloat16 on CUDA)
model = LCTQwen.from_pretrained("CaoHaoWei/Jev-LCT-Qwen2.5-1.5B")
# 2. Predict multiple-choice decision
result = model.predict_choice(
prompt="Patient reports severe chest pain radiating to left shoulder. Determine triage urgency level:",
choices=["emergency", "urgent", "routine", "elective"]
)
print(f"Top Decision: {result['choice']}")
print(f"Calibrated Confidence: {result['confidence']:.2%}")
print(f"Loops Executed: {result['loops']}")
中文详细说明
核心亮点
- 系统一决策新范式:结合 TypeSafe AI 的 Jev 理念,摆脱生成文本的高延迟,以毫秒级直出结构化决策与校准置信度。
- 通识与科学推理大幅领先:在 ARC 科学推理达到 71.7%(远超 Laya 30.0% 与 Open-Jev 45.0%),MMLU 达 58.3%。
- 免 RL 零额外开销校准:隐状态动力学余弦收敛与熵变轨迹特征,检错 AUROC 达 0.904 ~ 0.952,ECE 校准误差降低 23.5%。
- 自适应早停秒级直出:简单题单遍直出(~45ms),复杂题深思 2~4 轮,平均延迟仅 61.9 ms。
Citation & License
Apache License 2.0. Please cite as:
@article{cao2026lct,
title={Looped Calibration Transformer: Free Calibrated Confidence from Recurrent Computation Trajectories for Small Decision Models},
author={Cao, Haowei},
year={2026},
publisher={GitHub},
journal={GitHub repository},
howpublished={\url{https://github.com/gitchw/LCT}}
}