100x faster than generative LLMs • Runs on laptops & cloud CPUs • Global #1 on JevBench When you ask ChatGPT or Claude a question, it generates words one token at a time, like a person typing out an essay. That takes 2 to 5 seconds and burns expensive GPU compute. That is great for writing a story, but it is painfully slow and expensive for simple decisions: - "Did the AI make up this answer, or is it actually in the PDF?" - "Should this customer's message go to billing, shipping, or technical support?" - "Does the revenue bar chart support this financial claim?" - "Did the student get the math problem right according to the answer key?" Psychologist Daniel Kahneman described human thinking…
Jev-LCT-Qwen3-8B is an open-weight model for text classification from CaoHaoWei, released under Apache License 2.0. It has 8.2B parameters and a 40,960-token context. At 16-bit it needs about 19.7 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.
Jev-LCT-Qwen3-8B is the flagship enterprise-grade decision engine of the Jev-LCT family.
Runs On
What it takes to serve Jev-LCT-Qwen3-8B (8.2B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 16.4 GB | 19.7 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 8.2 GB | 9.8 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 4.1 GB | 4.9 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.
Jev-LCT-Qwen3-8B on every accelerator the SAVRN Index prices, at every precision
Model Card
By CaoHaoWei, published under apache-2.0, revision de343787fa15.
Jev-LCT-Qwen3-8B is the flagship enterprise-grade decision engine of the Jev-LCT family. Combining Qwen3-8B's extensive foundational capabilities with Looped Calibration and adaptive early exit, it provides frontier generative reasoning capabilities with deterministic sub-100ms decision latency. - 85.0% 科学推理 + 70.0% MMLU:媲美中大型生成模型的复杂逻辑推理能力,但单次推断控制在 89.2 ms 内。 - 企业级智能体中枢:支持高风险场景的“选择性预测(Selective Prediction)”,在 80% 覆盖率下实现近乎零差错审核。 - 全量独立权重:开箱即用,支持多 GPU 分片或单张 24GB 显卡(RTX 3090 / 4090)bfloat16 全速推断。 Apache License 2.0. Full repository at GitHub.
Read CaoHaoWei's full model card
Jev-LCT-Qwen3-8B: Flagship Open System-One Decision Engine
"Decisions, Not Strings" meets "Free Calibrated Confidence"
《Jev-LCT-8B:旗舰级系统一决策引擎,81.7% 全量高准确率,85% 科学推理,89ms 极低延迟》
English Overview
Jev-LCT-Qwen3-8B is the flagship enterprise-grade decision engine of the Jev-LCT family. Combining Qwen3-8B's extensive foundational capabilities with Looped Calibration and adaptive early exit, it provides frontier generative reasoning capabilities with deterministic sub-100ms decision latency.
Benchmark Highlights (300 Items on RTX 3090 Ti)
- Overall Accuracy: 81.7% (vs 51.3% for Laya and 59.3% for Open-Jev)
- Scientific Reasoning (ARC): 85.0% (vs 28.3% for Laya and 45.0% for Open-Jev)
- Factuality Verification (TruthfulQA): 75.0%
- Academic Multi-Task (MMLU): 70.0%
- Average Decision Latency: 89.2 ms
Quickstart
from lct_qwen_standalone import LCTQwen
# Load 8B flagship model (requires ~16GB VRAM in bfloat16)
model = LCTQwen.from_pretrained("CaoHaoWei/Jev-LCT-Qwen3-8B")
result = model.predict_choice(
prompt="A scientist observes that an unknown mineral scratches glass but cannot scratch quartz. Which mineral is it most likely to be?",
choices=["Gypsum (Hardness 2)", "Calcite (Hardness 3)", "Feldspar (Hardness 6)", "Topaz (Hardness 8)"]
)
print(f"Decision: {result['choice']} (Confidence: {result['confidence']:.2%})")
中文简介
- 81.7% 全量 300 题超高准确率:在意图识别、科学常识、通识、阅读理解与事实性全领域表现卓越。
- 85.0% 科学推理 + 70.0% MMLU:媲美中大型生成模型的复杂逻辑推理能力,但单次推断控制在 89.2 ms 内。
- 企业级智能体中枢:支持高风险场景的“选择性预测(Selective Prediction)”,在 80% 覆盖率下实现近乎零差错审核。
- 全量独立权重:开箱即用,支持多 GPU 分片或单张 24GB 显卡(RTX 3090 / 4090)bfloat16 全速推断。
Citation & License
Apache License 2.0. Full repository at GitHub.
Configuration
- Architecture
- Qwen3ForCausalLM
- Context length (tokens)
- 40,960
- Layers
- 36
- Hidden size
- 4,096
- Feed-forward size
- 12,288
- Attention heads
- 32
- Key/value heads
- 8
- Head dimension
- 128
- Vocabulary size
- 151,936
- RoPE base
- 1,000,000
- Model type
- qwen3
Identity and Version
- Repository
- CaoHaoWei/Jev-LCT-Qwen3-8B
- Publisher
- CaoHaoWei
- Task
- Text classification
- Modality
- Text
- Library
- transformers
- Parameters
- 8.2B parameters
- Languages
- en, zh
- Revision
- de343787fa15440e7f152980338d43fbe64a51ee
- First published
- 2026-09-26
- Last updated
- 2026-09-26
Files and Weights
18 files, 16.4 GB in total. The weights are 5 files totalling 16.4 GB in pt, safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| lct_auxiliary.pt | Weights | 18.3 KB | 08a70a463968 |
| model-00001-of-00004.safetensors | Weights | 4.9 GB | c0dc64934ae0 |
| model-00002-of-00004.safetensors | Weights | 4.9 GB | d58533b468c3 |
| model-00003-of-00004.safetensors | Weights | 5.0 GB | 11a3510beb93 |
| model-00004-of-00004.safetensors | Weights | 1.6 GB | f0e7d5c069f7 |
| added_tokens.json | Configuration | 735 B | — |
| config.json | Configuration | 1.6 KB | — |
| generation_config.json | Configuration | 227 B | — |
| lct_config.json | Configuration | 394 B | — |
| model.safetensors.index.json | Configuration | 33.3 KB | — |
| special_tokens_map.json | Configuration | 644 B | — |
| README.md | Documentation | 3.4 KB | — |
| chat_template.jinja | Other | 4.3 KB | — |
| .gitattributes | Repository | 1.6 KB | — |
| merges.txt | Tokenizer | 1.7 MB | — |
| tokenizer.json | Tokenizer | 11.4 MB | aeb13307a71a |
| tokenizer_config.json | Tokenizer | 5.6 KB | — |
| vocab.json | Tokenizer | 2.8 MB | — |
License and Download
- License
- apache-2.0
- Access
- Open weights, no gate
- Download size
- 16.4 GB
Released by CaoHaoWei through its official repository on Hugging Face. Read the license.
Built From
- Trained on (disclosed) ai2_arc
- Trained on (disclosed) banking77
- Trained on (disclosed) boolq
- Trained on (disclosed) mmlu
- Trained on (disclosed) truthful_qa
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 16.4 GB |
| 16-bit | 16.4 GB |
| 8-bit | 8.2 GB |
| 4-bit | 4.1 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About Jev-LCT-Qwen3-8B
How much GPU memory does Jev-LCT-Qwen3-8B need?
About 19.7 GB at 16-bit and 4.9 GB at 4-bit: the weights (8.2B parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run Jev-LCT-Qwen3-8B on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use Jev-LCT-Qwen3-8B commercially?
Yes. Jev-LCT-Qwen3-8B is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.
What is Jev-LCT-Qwen3-8B's context length?
40,960 tokens, from the maximum position embeddings in its published configuration.
Similar Models
100x faster than generative LLMs • Runs on laptops & cloud CPUs • Global #1 on JevBench When you ask ChatGPT or Claude a question, it generates words one token at a time, like a person typing out an essay. That takes 2 to 5 seconds and burns expensive GPU compute. That is great for writing a story, but it is painfully slow and expensive for simple decisions: - "Did the AI make up this answer, or is it actually in the PDF?" - "Should this customer's message go to billing, shipping, or technical support?" - "Does the revenue bar chart support this financial claim?" - "Did the student get the math problem right according to the answer key?" Psychologist Daniel Kahneman described human thinking…
Experimental open-weights judgment model by Kitani OpenJudgement is unfinished. We're releasing this checkpoint for people to experiment with, inspect, and build on. It still needs work on judgment quality, calibration, and inference efficiency. It is not as good as Jev overall in our internal task comparisons. It does show a meaningful improvement over untouched Qwen on our recorded validation comparison: 75.4% versus 64.2% annotation agreement. That is a result on a particular evaluation set, not a claim that we beat the base model on every task. There are questions it handles well and questions it confidently gets wrong. Please judge the preview by your own examples rather than assuming…
A calibrated typed-decision model: give it a state (text, ticket, policy, JSON) and a typed question — choice, boolean, or rubric score — and it returns a probability for every option in a single forward pass (~24 ms). No generation, no parsing, nothing to hallucinate. 12 of 13 subsets exceed Bespoke Nimble-9B — a model 2.2× its size — same prompt format, same scoring protocol. The primary suite: BoolQ, MultiNLI, PAWS, PubMedQA, SQuAD-2, VitaminC, Civil Comments, Aegis 2.0, MASSIVE (en/de), HelpSteer-2, SummEval (consistency / relevance). Every item human-labeled; byte-reproducible (manifest-locked ids + sha256); same protocol as the Bespoke Nimble evaluation. Wins: verification-style noul…
Small, fast decisions for routing and checks at volume. Send text or JSON state with choice, noul (yes/no), or score questions. Get a probability for every offered answer, not generated text. Each batch takes one forward pass; large requests can use several batches. Use choice to route a request, noul for a yes/no check, or score for an ordered rating. The same call can ask several questions about a single state. Try it: POST /v1/systemone, GET /v1/models, and GET /healthz. Point TypeSafe's server-side Python or JavaScript SDKs at it with TYPESAFEBASEURL; text decisions use the same request and response fields as hosted Jev. Requests run one at a time by default; --batch-window-ms 5 enables…
Mapika/decider-4b v2.1 quantized to FP8 for vLLM: FP8 E4M3 weights with one scale per output channel and FP8 activations scaled per token at run time. 4.85 GB against 8.41 GB for the bf16 checkpoint. Quantized and measured by LLM Tech; the model, its training and its evaluation protocol are Mapika's. Read the bf16 card for what the model is and how it was trained. The base revision is eb5fbdfc9448473ec25e399882912863afbdb70e. Tokenizer, chat template, generation config and deciderconfig.json (temperatures included) are the author's files unchanged, apart from the version and quantization fields. Both models were run through vLLM 0.29.0 on the same rows: the author's regression set rebuilt…