Lodestar-4B is a 4-billion-parameter decision model. You give it a state (plain text, JSON or a long document) and one
or more typed questions; it returns a probability for every listed option. Each answer is read from a single forward
pass. The model never generates free text, so there is nothing to parse and no output length to budget for.
|
|
| Base model |
Qwen3.5-4B-Base (32 layers: 24 Gated DeltaNet, 8 full attention) |
| Question types |
choice (one of up to 26 options per pass; longer lists are resolved in rounds), noul (yes/no), score (ordinal levels) |
| Output |
a probability for every option, calibrated per question type |
| Interface |
Python API and a TypeSafe-compatible POST /v1/systemone server (included) |
| Latency |
11 ms per short question, 21 ms median on the multi-thousand-token hard items (one H200, bf16, CUDA graphs) |
| Context |
tested up to 64k tokens |
| Precision |
bf16, about 9 GB of weights |
| License |
Apache-2.0 |
Quick start
hf download startlux-models/Lodestar-4B --local-dir Lodestar-4B
cd Lodestar-4B
pip install -r requirements.txt # flash-linear-attention is optional but much faster
python -m lodestar.server --model . --port 8090
curl -s localhost:8090/v1/systemone -H 'Content-Type: application/json' -d '{
"state": {"ticket": "I was charged twice for order #4411 and the app still shows it as unpaid."},
"questions": {
"team": {"type": "choice", "instructions": "Which team should handle this ticket?",
"criteria": {"billing": "Payments, refunds and invoices",
"shipping": "Delivery and tracking",
"technical": "App, login and account problems"}},
"urgent": {"type": "noul", "instructions": "Should this ticket be answered today?"}
}}'
{"answers": {"team": {"type": "choice", "choice": "billing", "probabilities": {"billing": 0.965, "shipping": 0.002, "technical": 0.034}},
"urgent": {"type": "noul", "noul": 0.623}},
"usage": {"input_tokens": 191, "output_tokens": 0}, "model": "Lodestar-4B"}
(Probabilities rounded to three decimals.) The same call from Python, run inside the downloaded folder:
from lodestar import Lodestar
model = Lodestar(".") # or "startlux-models/Lodestar-4B"; needs one CUDA GPU
answers, usage = model.decide(state, questions)
A score question takes its levels as a list, lowest first, and returns
{"type": "score", "score": <level index>, "probabilities": {"0": p0, "1": p1, ...}}.
How it answers
Every question is rendered into one chat prompt (thinking disabled):
system: Apply the criterion to the evidence. Choose exactly one listed option. Answer with its letter only.
user: Evidence:
<state>
Question: <instructions>
Options:
A) <option id>: <description>
B) ...
The probability of each option is the softmax of the next-token logits of its letter, taken at the last prompt
position and divided by the temperature of the question type (lodestar_config.json). Yes/no questions are shown as
the options yes and no. Choice lists longer than 26 options are split into groups of 25; the top three of each group
go to a final round, and the options that miss it share a small residual probability. The renderer is
lodestar/jevfmt.py; prompts built another way will not reproduce the published behaviour.
Training
- Supervised fine-tuning, all parameters. About 2.2 million typed decisions, converted into the prompt format
above from publicly available datasets: intent and topic classification, natural-language inference, reading
comprehension and multiple-choice knowledge questions, policy and rule application, tool and routing selection,
answer verification, preference and quality judgments, field extraction, and numeric and temporal reasoning. Where
a source provides a label distribution rather than a single label, the model is trained toward that distribution.
One epoch, learning rate 2e-5 with cosine decay.
- Targeted refinement. A rank-64 LoRA over the attention, DeltaNet and MLP projections, trained for two epochs on
61k rows: decisions the first-stage model got wrong or was unsure about, human-labelled sets, and replay rows whose
target is the first-stage model's own distribution, so that the refinement does not erode what the model already
does well. The adapter is merged into the released weights.
- Calibration. One temperature per question type, fitted by minimising NLL on the development splits of the three Kev
evaluation sets (devtools-v1, documents-v1, hard-v1; 3,077 rows; their test splits were used only to check). No
JevBench item was used: none of the fitting rows shares text with any JevBench public item.
noul 1.79, choice 1.47,
score 1.51. Temperatures change confidence, never the chosen option.
Training data was decontaminated against the benchmark items used for evaluation (13-gram overlap and exact match):
the Decision Index 0.2 test splits, the JevBench public set and Humanity's Last Exam. A separate exhaustive 13-gram
check of both training stages against the 231 JevBench public items finds no overlapping example.
Evaluation
JevBench public split (231 items), run through JevBench's own typesafe adapter against the bundled server:
| Tier |
Items |
Correct |
| Easy |
48 |
48 |
| Standard |
72 |
70 |
| Hard |
111 |
80 |
| All |
231 |
198 |
These are self-reported results on the public items. They are not an official leaderboard entry.
Limitations
- One pass, no chain of thought: multi-step arithmetic, date calculations and subtle answer checking are the weakest
areas, especially inside long documents.
- The temperatures were fitted on Kev development items, which are harder than typical inputs, so on easy inputs the
probabilities are on the cautious side.
- Training data is predominantly English.
- The model decides among the options it is given. It does not add options, and it answers
noul questions even when
the evidence is insufficient; give it an explicit "cannot tell" option when that matters.
- It is a decision component, not a chat assistant.
License
The weights are released under Apache-2.0, following the base model. The training data comes from public datasets,
each under its own license.