SAVRN
Search Contact SAVRN

Open-weight model · Text classification

Jev-LCT-Qwen3-8B

by CaoHaoWei CaoHaoWei/Jev-LCT-Qwen3-8B

Jev-LCT-Qwen3-8B is an open-weight model for text classification from CaoHaoWei, released under Apache License 2.0. It has 8.2B parameters and a 40,960-token context. At 16-bit it needs about 19.7 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

Jev-LCT-Qwen3-8B is the flagship enterprise-grade decision engine of the Jev-LCT family.

Parameters8.2B
Context40,960
Weights16.4 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve Jev-LCT-Qwen3-8B (8.2B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 16.4 GB 19.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 8.2 GB 9.8 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 4.1 GB 4.9 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.

Jev-LCT-Qwen3-8B on every accelerator the SAVRN Index prices, at every precision

Model Card

By CaoHaoWei, published under apache-2.0, revision de343787fa15.

Jev-LCT-Qwen3-8B is the flagship enterprise-grade decision engine of the Jev-LCT family. Combining Qwen3-8B's extensive foundational capabilities with Looped Calibration and adaptive early exit, it provides frontier generative reasoning capabilities with deterministic sub-100ms decision latency. - 85.0% 科学推理 + 70.0% MMLU:媲美中大型生成模型的复杂逻辑推理能力,但单次推断控制在 89.2 ms 内。 - 企业级智能体中枢:支持高风险场景的“选择性预测(Selective Prediction)”,在 80% 覆盖率下实现近乎零差错审核。 - 全量独立权重:开箱即用,支持多 GPU 分片或单张 24GB 显卡(RTX 3090 / 4090)bfloat16 全速推断。 Apache License 2.0. Full repository at GitHub.

Read CaoHaoWei's full model card

Jev-LCT-Qwen3-8B: Flagship Open System-One Decision Engine

"Decisions, Not Strings" meets "Free Calibrated Confidence"
《Jev-LCT-8B:旗舰级系统一决策引擎,81.7% 全量高准确率,85% 科学推理,89ms 极低延迟》


English Overview

Jev-LCT-Qwen3-8B is the flagship enterprise-grade decision engine of the Jev-LCT family. Combining Qwen3-8B's extensive foundational capabilities with Looped Calibration and adaptive early exit, it provides frontier generative reasoning capabilities with deterministic sub-100ms decision latency.

Benchmark Highlights (300 Items on RTX 3090 Ti)

  • Overall Accuracy: 81.7% (vs 51.3% for Laya and 59.3% for Open-Jev)
  • Scientific Reasoning (ARC): 85.0% (vs 28.3% for Laya and 45.0% for Open-Jev)
  • Factuality Verification (TruthfulQA): 75.0%
  • Academic Multi-Task (MMLU): 70.0%
  • Average Decision Latency: 89.2 ms

Quickstart

from lct_qwen_standalone import LCTQwen

# Load 8B flagship model (requires ~16GB VRAM in bfloat16)
model = LCTQwen.from_pretrained("CaoHaoWei/Jev-LCT-Qwen3-8B")

result = model.predict_choice(
    prompt="A scientist observes that an unknown mineral scratches glass but cannot scratch quartz. Which mineral is it most likely to be?",
    choices=["Gypsum (Hardness 2)", "Calcite (Hardness 3)", "Feldspar (Hardness 6)", "Topaz (Hardness 8)"]
)
print(f"Decision: {result['choice']} (Confidence: {result['confidence']:.2%})")

中文简介

  • 81.7% 全量 300 题超高准确率:在意图识别、科学常识、通识、阅读理解与事实性全领域表现卓越。
  • 85.0% 科学推理 + 70.0% MMLU:媲美中大型生成模型的复杂逻辑推理能力,但单次推断控制在 89.2 ms 内。
  • 企业级智能体中枢:支持高风险场景的“选择性预测(Selective Prediction)”,在 80% 覆盖率下实现近乎零差错审核。
  • 全量独立权重:开箱即用,支持多 GPU 分片或单张 24GB 显卡(RTX 3090 / 4090)bfloat16 全速推断。

Citation & License

Apache License 2.0. Full repository at GitHub.

Configuration

Architecture
Qwen3ForCausalLM
Context length (tokens)
40,960
Layers
36
Hidden size
4,096
Feed-forward size
12,288
Attention heads
32
Key/value heads
8
Head dimension
128
Vocabulary size
151,936
RoPE base
1,000,000
Model type
qwen3

Identity and Version

Repository
CaoHaoWei/Jev-LCT-Qwen3-8B
Publisher
CaoHaoWei
Task
Text classification
Modality
Text
Library
transformers
Parameters
8.2B parameters
Languages
en, zh
Revision
de343787fa15440e7f152980338d43fbe64a51ee
First published
2026-09-26
Last updated
2026-09-26

Files and Weights

18 files, 16.4 GB in total. The weights are 5 files totalling 16.4 GB in pt, safetensors.

Weights5 files · 16.4 GB
Configuration6 files · 36.9 KB
Tokenizer4 files · 15.9 MB
Documentation1 file · 3.4 KB
Other1 file · 4.3 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
lct_auxiliary.ptWeights18.3 KB 08a70a463968
model-00001-of-00004.safetensorsWeights4.9 GB c0dc64934ae0
model-00002-of-00004.safetensorsWeights4.9 GB d58533b468c3
model-00003-of-00004.safetensorsWeights5.0 GB 11a3510beb93
model-00004-of-00004.safetensorsWeights1.6 GB f0e7d5c069f7
added_tokens.jsonConfiguration735 B —
config.jsonConfiguration1.6 KB —
generation_config.jsonConfiguration227 B —
lct_config.jsonConfiguration394 B —
model.safetensors.index.jsonConfiguration33.3 KB —
special_tokens_map.jsonConfiguration644 B —
README.mdDocumentation3.4 KB —
chat_template.jinjaOther4.3 KB —
.gitattributesRepository1.6 KB —
merges.txtTokenizer1.7 MB —
tokenizer.jsonTokenizer11.4 MB aeb13307a71a
tokenizer_config.jsonTokenizer5.6 KB —
vocab.jsonTokenizer2.8 MB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
16.4 GB
Download from CaoHaoWei

Released by CaoHaoWei through its official repository on Hugging Face. Read the license.

Built From

  • Trained on (disclosed) ai2_arc
  • Trained on (disclosed) banking77
  • Trained on (disclosed) boolq
  • Trained on (disclosed) mmlu
  • Trained on (disclosed) truthful_qa

Memory Requirements

PrecisionWeights in memory
As published16.4 GB
16-bit16.4 GB
8-bit8.2 GB
4-bit4.1 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Jev-LCT-Qwen3-8B

How much GPU memory does Jev-LCT-Qwen3-8B need?

About 19.7 GB at 16-bit and 4.9 GB at 4-bit: the weights (8.2B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Jev-LCT-Qwen3-8B on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Jev-LCT-Qwen3-8B commercially?

Yes. Jev-LCT-Qwen3-8B is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is Jev-LCT-Qwen3-8B's context length?

40,960 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text classification

gevva-e2b-multimodal

David Burhans

100x faster than generative LLMs • Runs on laptops & cloud CPUs • Global #1 on JevBench When you ask ChatGPT or Claude a question, it generates words one token at a time, like a person typing out an essay. That takes 2 to 5 seconds and burns expensive GPU compute. That is great for writing a story, but it is painfully slow and expensive for simple decisions: - "Did the AI make up this answer, or is it actually in the PDF?" - "Should this customer's message go to billing, shipping, or technical support?" - "Does the revenue bar chart support this financial claim?" - "Did the student get the math problem right according to the answer key?" Psychologist Daniel Kahneman described human thinking…

Open weights apache-2.0 5.1B parameters 131,072 tokens transformers

Model · Text classification

gevva-e2b

David Burhans

100x faster than generative LLMs • Runs on laptops & cloud CPUs • Global #1 on JevBench When you ask ChatGPT or Claude a question, it generates words one token at a time, like a person typing out an essay. That takes 2 to 5 seconds and burns expensive GPU compute. That is great for writing a story, but it is painfully slow and expensive for simple decisions: - "Did the AI make up this answer, or is it actually in the PDF?" - "Should this customer's message go to billing, shipping, or technical support?" - "Does the revenue bar chart support this financial claim?" - "Did the student get the math problem right according to the answer key?" Psychologist Daniel Kahneman described human thinking…

Open weights apache-2.0 5.1B parameters 131,072 tokens transformers

Model · Text classification

OpenJudgement-4B-Preview

Kitani

Experimental open-weights judgment model by Kitani OpenJudgement is unfinished. We're releasing this checkpoint for people to experiment with, inspect, and build on. It still needs work on judgment quality, calibration, and inference efficiency. It is not as good as Jev overall in our internal task comparisons. It does show a meaningful improvement over untouched Qwen on our recorded validation comparison: 75.4% versus 64.2% annotation agreement. That is a result on a particular evaluation set, not a claim that we beat the base model on every task. There are questions it handles well and questions it confidently gets wrong. Please judge the preview by your own examples rather than assuming…

Open weights apache-2.0 4.5B parameters 262,144 tokens transformers

Model · Text classification

metask-jev-4b-policy-mix

Raymond wei

A calibrated typed-decision model: give it a state (text, ticket, policy, JSON) and a typed question — choice, boolean, or rubric score — and it returns a probability for every option in a single forward pass (~24 ms). No generation, no parsing, nothing to hallucinate. 12 of 13 subsets exceed Bespoke Nimble-9B — a model 2.2× its size — same prompt format, same scoring protocol. The primary suite: BoolQ, MultiNLI, PAWS, PubMedQA, SQuAD-2, VitaminC, Civil Comments, Aegis 2.0, MASSIVE (en/de), HelpSteer-2, SummEval (consistency / relevance). Every item human-labeled; byte-reproducible (manifest-locked ids + sha256); same protocol as the Bespoke Nimble evaluation. Wins: verification-style noul…

Open weights apache-2.0 4.5B parameters 262,144 tokens transformers

Model · Text classification

blink-4b

Govind Kamtamneni

Small, fast decisions for routing and checks at volume. Send text or JSON state with choice, noul (yes/no), or score questions. Get a probability for every offered answer, not generated text. Each batch takes one forward pass; large requests can use several batches. Use choice to route a request, noul for a yes/no check, or score for an ordered rating. The same call can ask several questions about a single state. Try it: POST /v1/systemone, GET /v1/models, and GET /healthz. Point TypeSafe's server-side Python or JavaScript SDKs at it with TYPESAFEBASEURL; text decisions use the same request and response fields as hosted Jev. Requests run one at a time by default; --batch-window-ms 5 enables…

Open weights other 4.2B parameters 262,144 tokens transformers

Model · Text classification

decider-4b-fp8

LLM Tech

Mapika/decider-4b v2.1 quantized to FP8 for vLLM: FP8 E4M3 weights with one scale per output channel and FP8 activations scaled per token at run time. 4.85 GB against 8.41 GB for the bf16 checkpoint. Quantized and measured by LLM Tech; the model, its training and its evaluation protocol are Mapika's. Read the bf16 card for what the model is and how it was trained. The base revision is eb5fbdfc9448473ec25e399882912863afbdb70e. Tokenizer, chat template, generation config and deciderconfig.json (temperatures included) are the author's files unchanged, apart from the version and quantization fields. Both models were run through vLLM 0.29.0 on the same rows: the author's regression set rebuilt…

Open weights apache-2.0 4.2B parameters 262,144 tokens