SAVRN
Search Contact SAVRN

Open-weight model

jevify-qwen3.5-4b-base-readout-coh

by Praveen Raj U S Praveenrajus/jevify-qwen3.5-4b-base-readout-coh

jevify-qwen3.5-4b-base-readout-coh is an open-weight model from Praveen Raj U S, released under Apache License 2.0. Its published files total 85.1 MB.

A System One decision model: it does not write text. It reads a state, answers typed questions (choice, score, noul), and returns calibrated probability distributions your code can branch on.

Parameters—
Context—
Weights85.0 MB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Model Card

By Praveen Raj U S, published under apache-2.0, revision 1a90e7ba296a.

A System One decision model: it does not write text. It reads a state, answers typed questions (choice, score, noul), and returns calibrated probability distributions your code can branch on. This repo is a rank-16 LoRA (21,233,664 parameters) on Qwen/Qwen3.5-4B-Base, merged into the weights at load. How it was trained (readout fine-tuning). The model is trained on its own decision readout — the distribution over the allowed answers read at the answer position, one forward pass, no decoding — with the primitive's proper scoring rule, plus a coherence penalty (weight 1.0): every training question comes with automatically derived siblings (the options as yes/no questions, the negation, the…

Read Praveen Raj U S's full model card

Qwen3.5-4B-Base, readout fine-tuned (LoRA, + coherence)

A System One decision model: it does not write text. It reads a state, answers typed questions (choice, score, noul), and returns calibrated probability distributions your code can branch on. This repo is a rank-16 LoRA (21,233,664 parameters) on Qwen/Qwen3.5-4B-Base, merged into the weights at load.

How it was trained (readout fine-tuning). The model is trained on its own decision readout — the distribution over the allowed answers read at the answer position, one forward pass, no decoding — with the primitive's proper scoring rule, plus a coherence penalty (weight 1.0): every training question comes with automatically derived siblings (the options as yes/no questions, the negation, the threshold questions of a scale), and the de Finetti sure loss of the family's answers is penalised, so the model's answers to related questions stay mutually consistent. Options are shuffled per family. Training data: the train splits of the 16 non-held-out jev-bench sources (5,885 families, at most 400 records per source); lr 3e-05, 2 epochs, best epoch by validation loss (epoch 0), seed 0. A Tier 0 recipe (temperature per primitive, Noul bias, option-order permutations) was then fitted on validation splits. The six held-out sources (clinc150, arc_challenge, yelp5, measuring_hate_speech, fever_evidence, strategyqa_grounded) never appeared in training.

from jevify import load_jevified

model = load_jevified("Praveenrajus/jevify-qwen3.5-4b-base-readout-coh")
model.ask({"text": "The battery lasted two days on a single charge."},
          {"q": {"type": "noul", "instructions": "Is the review positive?"}})

jevify-serve --model Praveenrajus/jevify-qwen3.5-4b-base-readout-coh serves it as a drop-in for the TypeSafe SDK (TYPESAFE_BASE_URL=http://localhost:8000). The backbone is pulled from its own repo at load, pinned to commit 1001bb4d826a52d1f399e183466143f4da7b741b.

Results

Every number is on the jev-bench test splits (22,773 records) or the study's other test suites, scored the same way for every model; rows below this model are references from the same study.

Decisions and calibration

model acc ECE Brier held-out acc TVD to human labels
this model 0.747 0.058 0.318 0.783 0.314
Qwen3.5-4B-Base, untuned (Tier 0) 0.647 0.095 0.424 0.710 0.456
same recipe, supervised only 0.741 0.064 0.324 0.769 0.324
same recipe from the instruct model 0.751 0.058 0.316 0.792 0.303
Jev 1.13.0 (TypeSafe API) 0.733 0.113 0.349 0.835 0.432

Coherence and invariance — sure loss: mean d² over 4,749 question families (0 = perfectly coherent); order flip: how often the top answer changes when options are shuffled; tag TVD: how much the distribution moves when option tags change from A–J to other identifiers.

model sure loss share incoherent order flip tag TVD K=2→max acc drop
this model 0.033 0.396 0.058 0.021 0.224
Qwen3.5-4B-Base, untuned (Tier 0) 0.188 0.982 0.191 0.030 0.371
same recipe, supervised only 0.271 0.915 0.064 0.020 0.237
same recipe from the instruct model 0.029 0.347 0.057 0.023 0.230
Jev 1.13.0 (TypeSafe API) 0.081 0.725 0.046 — 0.246

Out of distribution — stated rules (LegalBench, rule given in the question), none-of-the-above when the gold option is removed, injected-instruction hijack rate, and three community Jev benchmarks.

model stated rule 'none' when gone hijack phishing AUROC tool risk
this model 0.734 0.612 0.098 0.845 0.900
Qwen3.5-4B-Base, untuned (Tier 0) 0.627 0.356 0.301 0.945 0.867
same recipe, supervised only 0.730 0.656 0.133 0.886 0.817
same recipe from the instruct model 0.723 0.682 0.089 0.929 0.900
Jev 1.13.0 (TypeSafe API) 0.924 0.744 0.205 0.688 0.933

Reproduction check. Loading this folder with load_jevified and re-scoring 72 jev-bench test records from six sources reproduced the training run's own test predictions: 0 changed choice answers, mean largest |Δp| 0.005, max 0.022 (the adapter is merged into bf16 weights at load).

results/ holds the raw files: test_metrics.json (per source), recipe.json, coherence.json, probes.json, tags.json, train.json (the training log) and summary.json (this model's row of the study table).

Limitations

  • One training seed per repo branch; out-of-distribution numbers in particular vary between identical runs, so compare arms across seeds before drawing conclusions.
  • The phishing benchmark's decision threshold shifts after fine-tuning (ranking, AUROC, is preserved); a one-number log-odds shift fitted on a handful of labelled emails repairs it.
  • English only; the recipe was fitted on jev-bench validation splits and may need refitting on a very different domain.

Method, benchmark and findings: github.com/uspraveen/Jevify (docs/FINDINGS.md).

Identity and Version

Repository
Praveenrajus/jevify-qwen3.5-4b-base-readout-coh
Publisher
Praveen Raj U S
Task
Not stated by the source
Modality
Other
Library
jevify
Parameters
Not stated by the source
Languages
Not stated by the source
Revision
1a90e7ba296a1c8dca3955bfbec67b1ea0597ab8
First published
2026-09-25
Last updated
2026-09-25

Files and Weights

13 files, 85.1 MB in total. The weights are 1 file totalling 85.0 MB in safetensors.

Weights1 file · 85.0 MB
Configuration10 files · 95.5 KB
Documentation1 file · 5.5 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
lora/adapter_model.safetensorsWeights85.0 MB 12f3ad3424c0
jevify_config.jsonConfiguration1.4 KB —
lora/adapter_config.jsonConfiguration1.2 KB —
results/coherence.jsonConfiguration27.6 KB —
results/probes.jsonConfiguration14.3 KB —
results/recipe.jsonConfiguration2.4 KB —
results/summary.jsonConfiguration1.5 KB —
results/tags.jsonConfiguration2.6 KB —
results/test_metrics.jsonConfiguration43.4 KB —
results/train.jsonConfiguration896 B —
results/verification.jsonConfiguration148 B —
README.mdDocumentation5.5 KB —
.gitattributesRepository1.5 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
85.0 MB
Download from Praveen Raj U S

Released by Praveen Raj U S through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published85.0 MB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About jevify-qwen3.5-4b-base-readout-coh

Can I use jevify-qwen3.5-4b-base-readout-coh commercially?

Yes. jevify-qwen3.5-4b-base-readout-coh is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.