Open-weight model
jevify-qwen3.5-4b-readout
by Praveen Raj U S Praveenrajus/jevify-qwen3.5-4b-readout
jevify-qwen3.5-4b-readout is an open-weight model from Praveen Raj U S, released under Apache License 2.0. Its published files total 85.1 MB.
A System One decision model: it does not write text. It reads a state, answers typed questions (choice, score, noul), and returns calibrated probability distributions your code can branch on.
Model Card
By Praveen Raj U S, published under apache-2.0, revision 9036639a41da.
A System One decision model: it does not write text. It reads a state, answers typed questions (choice, score, noul), and returns calibrated probability distributions your code can branch on. This repo is a rank-16 LoRA (21,233,664 parameters) on Qwen/Qwen3.5-4B, merged into the weights at load. How it was trained (readout fine-tuning). The model is trained on its own decision readout — the distribution over the allowed answers read at the answer position, one forward pass, no decoding — with the primitive's proper scoring rule. Options are shuffled per family. Training data: the train splits of the 16 non-held-out jev-bench sources (5,885 families, at most 400 records per source); lr…
Read Praveen Raj U S's full model card
Qwen3.5-4B, readout fine-tuned (LoRA, supervised)
A System One decision model: it does not write text. It reads a state, answers typed questions (choice, score,
noul), and returns calibrated probability distributions your code can branch on. This repo is a rank-16 LoRA (21,233,664 parameters) on Qwen/Qwen3.5-4B, merged into the weights at load.
How it was trained (readout fine-tuning). The model is trained on its own decision readout — the distribution over
the allowed answers read at the answer position, one forward pass, no decoding — with the primitive's proper scoring
rule.
Options are shuffled per family. Training data: the train splits of the 16 non-held-out jev-bench sources
(5,885 families, at most 400 records per source); lr 3e-05, 2 epochs,
best epoch by validation loss (epoch 0), seed 0.
A Tier 0 recipe (temperature per primitive, Noul bias, option-order permutations) was then fitted on validation splits.
The six held-out sources (clinc150, arc_challenge, yelp5, measuring_hate_speech, fever_evidence, strategyqa_grounded) never appeared in training.
from jevify import load_jevified
model = load_jevified("Praveenrajus/jevify-qwen3.5-4b-readout")
model.ask({"text": "The battery lasted two days on a single charge."},
{"q": {"type": "noul", "instructions": "Is the review positive?"}})
jevify-serve --model Praveenrajus/jevify-qwen3.5-4b-readout serves it as a drop-in for the TypeSafe SDK (TYPESAFE_BASE_URL=http://localhost:8000).
The backbone is pulled from its own repo at load, pinned to commit 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a.
Results
Every number is on the jev-bench test splits (22,773 records) or the study's other test suites, scored the same way for every model; rows below this model are references from the same study.
Decisions and calibration
| model | acc | ECE | Brier | held-out acc | TVD to human labels |
|---|---|---|---|---|---|
| this model | 0.743 | 0.059 | 0.321 | 0.774 | 0.326 |
| Qwen3.5-4B, untuned (Tier 0) | 0.662 | 0.090 | 0.401 | 0.719 | 0.438 |
| same recipe + coherence | 0.751 | 0.058 | 0.316 | 0.792 | 0.303 |
| Jev 1.13.0 (TypeSafe API) | 0.733 | 0.113 | 0.349 | 0.835 | 0.432 |
Coherence and invariance — sure loss: mean d² over 4,749 question families (0 = perfectly coherent); order flip: how often the top answer changes when options are shuffled; tag TVD: how much the distribution moves when option tags change from A–J to other identifiers.
| model | sure loss | share incoherent | order flip | tag TVD | K=2→max acc drop |
|---|---|---|---|---|---|
| this model | 0.283 | 0.917 | 0.072 | 0.023 | 0.228 |
| Qwen3.5-4B, untuned (Tier 0) | 0.151 | 0.959 | 0.140 | 0.038 | 0.299 |
| same recipe + coherence | 0.029 | 0.347 | 0.057 | 0.023 | 0.230 |
| Jev 1.13.0 (TypeSafe API) | 0.081 | 0.725 | 0.046 | — | 0.246 |
Out of distribution — stated rules (LegalBench, rule given in the question), none-of-the-above when the gold option is removed, injected-instruction hijack rate, and three community Jev benchmarks.
| model | stated rule | 'none' when gone | hijack | phishing AUROC | tool risk |
|---|---|---|---|---|---|
| this model | 0.742 | 0.614 | 0.116 | 0.864 | 0.833 |
| Qwen3.5-4B, untuned (Tier 0) | 0.619 | 0.484 | 0.394 | 0.784 | 0.867 |
| same recipe + coherence | 0.723 | 0.682 | 0.089 | 0.929 | 0.900 |
| Jev 1.13.0 (TypeSafe API) | 0.924 | 0.744 | 0.205 | 0.688 | 0.933 |
Reproduction check. Loading this folder with load_jevified and re-scoring 72 jev-bench test records from six sources reproduced the training run's own test predictions: 0 changed choice answers, mean largest |Δp| 0.004, max 0.024 (the adapter is merged into bf16 weights at load).
results/ holds the raw files: test_metrics.json (per source), recipe.json, coherence.json, probes.json,
tags.json, train.json (the training log) and summary.json (this model's row of the study table).
Limitations
- One training seed per repo branch; out-of-distribution numbers in particular vary between identical runs, so compare arms across seeds before drawing conclusions.
- The phishing benchmark's decision threshold shifts after fine-tuning (ranking, AUROC, is preserved); a one-number log-odds shift fitted on a handful of labelled emails repairs it.
- English only; the recipe was fitted on jev-bench validation splits and may need refitting on a very different domain.
Method, benchmark and findings: github.com/uspraveen/Jevify (docs/FINDINGS.md).
Identity and Version
- Repository
- Praveenrajus/jevify-qwen3.5-4b-readout
- Publisher
- Praveen Raj U S
- Task
- Not stated by the source
- Modality
- Other
- Library
- jevify
- Parameters
- Not stated by the source
- Languages
- Not stated by the source
- Revision
- 9036639a41da1d7e0608d2ef4723cc0042863521
- First published
- 2026-09-25
- Last updated
- 2026-09-25
Files and Weights
13 files, 85.1 MB in total. The weights are 1 file totalling 85.0 MB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| lora/adapter_model.safetensors | Weights | 85.0 MB | f00d2f255c55 |
| jevify_config.json | Configuration | 1.4 KB | — |
| lora/adapter_config.json | Configuration | 1.2 KB | — |
| results/coherence.json | Configuration | 27.7 KB | — |
| results/probes.json | Configuration | 14.4 KB | — |
| results/recipe.json | Configuration | 2.4 KB | — |
| results/summary.json | Configuration | 1.9 KB | — |
| results/tags.json | Configuration | 2.5 KB | — |
| results/test_metrics.json | Configuration | 43.9 KB | — |
| results/train.json | Configuration | 848 B | — |
| results/verification.json | Configuration | 139 B | — |
| README.md | Documentation | 4.8 KB | — |
| .gitattributes | Repository | 1.5 KB | — |
License and Download
- License
- apache-2.0
- Access
- Open weights, no gate
- Download size
- 85.0 MB
Released by Praveen Raj U S through its official repository on Hugging Face. Read the license.
Built From
- Derived from Qwen/Qwen3.5-4B
- Trained on (disclosed) Praveenrajus/jev-bench
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 85.0 MB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About jevify-qwen3.5-4b-readout
Can I use jevify-qwen3.5-4b-readout commercially?
Yes. jevify-qwen3.5-4b-readout is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.