SAVRN
Search Contact SAVRN

Open-weight model · Text generation

Jev-Style-Qwen3.5-2B-Decision-GGUF

by Chaoliang Yan chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-GGUF

Jev-Style-Qwen3.5-2B-Decision-GGUF is an open-weight model for text generation from Chaoliang Yan, released under Apache License 2.0. Its published files total 7.3 GB. It draws 6k downloads a month.

Website: jevstyle.com — all JevStyle decision models, benchmarks and quickstart in one place. A Jev-style decision model: it does not write text.

Parameters—
Context—
Weights7.3 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads6k

Model Card

By Chaoliang Yan, published under apache-2.0, revision adc565674171.

Website: jevstyle.com — all JevStyle decision models, benchmarks and quickstart in one place. A Jev-style decision model: it does not write text. Give it a state, a question and a list of options, and it returns the decision with calibrated probabilities from a single token position. Runs in LM Studio and llama.cpp. MLX bf16 build for Apple Silicon: chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-MLX-bf16 Everything below is measured on data the model never trained on, with the probabilities exactly as the released weights produce them (no post-processing). - Calibrated out of the box. An ECE of 0.017 on 1,500 examples is statistically indistinguishable from a perfectly calibrated model…

Read Chaoliang Yan's full model card

Jev-Style v3 is available — smaller and stronger: Jev-Style-0.8B-Decision-v3-GGUF scores 79.2% on the 2,000 typed decisions (v1: 53.4%, v2: 73.5%, as reported on the v2 card), takes 25,600-token inputs, works across 51 languages and scores options without the 26-letter cap (tested with 77 options), all at 0.8B parameters. This repository preserves v1; v2 is here.

Jev-Style-Qwen3.5-2B-Decision (GGUF)

Website: jevstyle.com — all JevStyle decision models, benchmarks and quickstart in one place.

A Jev-style decision model: it does not write text. Give it a state, a question and a list of options, and it returns the decision with calibrated probabilities from a single token position. Runs in LM Studio and llama.cpp.

File Size Same decision as bf16 Accuracy (500 held-out)
Jev-Style-Qwen3.5-2B-Decision-Q4_K_M.gguf 1.3 GB 94.4% 82.4%
Jev-Style-Qwen3.5-2B-Decision-Q8_0.gguf 2.1 GB 99.4% 81.6%
Jev-Style-Qwen3.5-2B-Decision-BF16.gguf 3.9 GB 99.8% 81.6%

MLX bf16 build for Apple Silicon: chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-MLX-bf16

Results

Everything below is measured on data the model never trained on, with the probabilities exactly as the released weights produce them (no post-processing).

Qwen3.5-2B-Base, zero-shot This model
Accuracy, 5 decision tasks (1,500 held-out examples) 65.9% 82.3%
Calibration error (ECE) on those tasks 0.065 0.017
Negative log-likelihood / Brier score 0.786 / 0.446 0.418 / 0.242
Calibration error on task types never seen in training 0.155 0.075
Latency per decision (M1 Max, MLX bf16) 76 ms 77 ms (no added cost)
  • Calibrated out of the box. An ECE of 0.017 on 1,500 examples is statistically indistinguishable from a perfectly calibrated model: simulating labels from the model's own probabilities gives an expected ECE of 0.017 (95th percentile 0.025) from sampling noise alone. When this model says 80%, it is right about 80% of the time.
  • Large accuracy gains where the base model struggled: MNLI 52.3% -> 86.7%, SST-5 32.0% -> 61.7%, BoolQ 73.0% -> 82.7%, SST-2 87.3% -> 92.7%, AG News 84.7% -> 87.7%.
  • Calibration transfers to new task types: on emotion classification and RTE (never seen in training) the calibration error is halved (0.155 -> 0.075) at unchanged accuracy (64.5%).
  • Zero-cost calibration. A temperature fitted on 4,366 held-out examples is folded into the final RMSNorm weight, so every logit is already calibrated. Nothing to apply at inference time.
  • Quantisation-friendly. Q8_0 makes the same decision as bf16 on 99.4% of examples; Q4_K_M (1.3 GB) keeps 82.4% accuracy.
  • Efficient training recipe. LoRA rank 16 on all linear layers with a log-score loss, built on a custom chunk-parallel, differentiable Gated DeltaNet forward that matches the per-token training path to 1e-6 (outputs, state and all gradients) and is 6.5x faster per step (measured on the 0.8B sibling model).

What "Jev-style" means

Jev (TypeSafe AI, 2026) introduced System One models: instead of generating text, the model takes a state plus a typed question and returns a decision with calibrated probabilities in a single pass. This model follows that pattern on top of an open base model:

  • Choice - pick one of N declared options, with a probability for each
  • Bool - probability that a proposition is true
  • Score - a distribution over ordered levels, and its expectation as a continuous score

It cannot answer outside the declared options, it does not decode text, and one prefill pass gives the whole distribution.

This is an independent, from-scratch reproduction of the publicly described idea. It is not affiliated with TypeSafe AI and is not the Jev model.

Quick start

Important: you must use the prompt format below. Plain chat messages produce meaningless text continuation. This is a decision function, not a chat model. In a chat window, paste the full prompt (it must end with Answer:) and start a new chat for every decision.

LM Studio - download a file from this repo, load it, start the local server (Developer tab), then:

python jev_style_client.py --url http://localhost:1234 --model jev-style-qwen3.5-2b-decision

llama.cpp

llama-server -hf chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-GGUF:Q8_0 --port 8080
python jev_style_client.py --url http://localhost:8080

jev_style_client.py (standard library only) asks for one token with top_logprobs on /v1/chat/completions and renormalises the option letters:

from jev_style_client import decide, decide_bool, decide_score

decide("http://localhost:1234",
       "Shares of the chipmaker jumped 8% after it raised its revenue forecast.",
       "Which news section does this article belong to?",
       ["World", "Sports", "Business", "Science/Technology"],
       model="jev-style-qwen3.5-2b-decision")
# [('Business', 0.68), ('Science/Technology', 0.31), ('World', 0.005), ('Sports', 0.002)]

Verified end to end through LM Studio's server on 500 held-out examples: 81.6% accuracy, ECE 0.028, about 110 ms per decision over HTTP on an M1 Max. In the LM Studio chat window you can also paste the prompt below and the model replies with the option letter.

Prompt format

You are a decision function. Read the state, then answer the question by choosing exactly one option.

[State]
{state}

[Question]
{question}

[Options]
A. {option 1}
B. {option 2}

Answer:

The next token is the option letter (A, B, ...). Its probability, renormalised over the declared letters, is the decision distribution. For Score, list the levels in order; for Bool, use yes / no. The repository ships a pass-through chat template, so chat endpoints and the LM Studio chat window pass this text to the model verbatim.

Scope

  • A decision function, not a chat model: send the prompt format above.
  • Trained on five English task families (sentiment, natural-language inference, topic, yes/no question answering, 5-level rating). On unseen task types it keeps the base model's accuracy with better, though not perfect, calibration.
  • Up to 26 options (20 when probabilities are read through a server's top_logprobs).

Training data and licence

SST-2 and MNLI (GLUE), AG News, BoolQ and SST-5, 22k examples converted to typed decisions; 80% for LoRA training, 20% held out for the calibration temperature. AG News is distributed for research / non-commercial use. Weights: Apache-2.0, same as Qwen/Qwen3.5-2B-Base.

Contact

I welcome internship, employment, and research collaboration opportunities. Please contact me at [email protected].

欢迎提供实习、工作及科研合作机会,请邮件联系:[email protected]。

Identity and Version

Repository
chaoliangUNSW/Jev-Style-Qwen3.5-2B-Decision-GGUF
Publisher
Chaoliang Yan
Task
Text generation
Modality
Text
Library
Not stated by the source
Parameters
Not stated by the source
Languages
en
Revision
adc5656741715ddd3e40dac44c6294cb888bc655
First published
2026-09-21
Last updated
2026-09-27

Files and Weights

7 files, 7.3 GB in total. The weights are 3 files totalling 7.3 GB in gguf.

Weights3 files · 7.3 GB
Configuration1 file · 4.6 KB
Documentation1 file · 7.8 KB
Other1 file · 221.9 KB
Repository1 file · 1.8 KB
Every file
FileTypeSizeSHA-256
Jev-Style-Qwen3.5-2B-Decision-BF16.ggufWeights3.9 GB ccc6f44df660
Jev-Style-Qwen3.5-2B-Decision-Q4_K_M.ggufWeights1.3 GB f3c14cd9d6d3
Jev-Style-Qwen3.5-2B-Decision-Q8_0.ggufWeights2.1 GB 470aa63b87fe
jev_style_client.pyConfiguration4.6 KB —
README.mdDocumentation7.8 KB —
calibration.pngOther221.9 KB fedaafb359d8
.gitattributesRepository1.8 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
7.3 GB
Download from Chaoliang Yan

Released by Chaoliang Yan through its official repository on Hugging Face. Read the license.

Built From

  • Derived from Qwen/Qwen3.5-2B-Base
  • Quantized from Qwen/Qwen3.5-2B-Base
  • Trained on (disclosed) SetFit/sst5
  • Trained on (disclosed) fancyzhx/ag_news
  • Trained on (disclosed) google/boolq
  • Trained on (disclosed) nyu-mll/glue

Memory Requirements

PrecisionWeights in memory
As published7.3 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Jev-Style-Qwen3.5-2B-Decision-GGUF

Can I use Jev-Style-Qwen3.5-2B-Decision-GGUF commercially?

Yes. Jev-Style-Qwen3.5-2B-Decision-GGUF is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Fine-tune Qwen3 (14B) for free using our Google Colab notebook! - Read our Blog about Qwen3 support: unsloth.ai/blog/qwen3 - View the rest of our notebooks in our docs here. Qwen3-Coder is available in multiple sizes. Today, we're excited to introduce Qwen3-Coder-30B-A3B-Instruct. This streamlined model maintains impressive performance and efficiency, featuring the following key enhancements: - Significant Performance among open models on Agentic Coding, Agentic Browser-Use, and other foundational coding tasks. - Long-context Capabilities with native support for 256K tokens, extendable up to 1M tokens using Yarn, optimized for repository-scale understanding. - Agentic Coding supporting for…

Open weights apache-2.0 transformers

Model · Text generation

opt-125m

AI at Meta

OPT was first introduced in Open Pre-trained Transformer Language Models and first released in metaseq's repository on May 3rd 2022 by Meta AI. Disclaimer: The team releasing OPT wrote an official model card, which is available in Appendix D of the paper. Content from this model card has been written by the Hugging Face team. To quote the first two paragraphs of the official paper OPT was predominantly pretrained with English text, but a small amount of non-English data is still present within the training corpus via CommonCrawl. The model was pretrained using a causal language modeling (CLM) objective. OPT belongs to the same family of decoder-only models like GPT-3. As such, it was…

Open weights other 2,048 tokens transformers

Model · Text generation

Ornith-1.5-9B-GGUF

Ornith

Chirp Chirp! We are introducing Ornith-1.5, a major step toward building foundation models through end-to-end self-improvement. Ornith-1.5 extends Ornith-1.0 (which was developed on top of Qwen3.5 and Gemma4 with additional continued pretraining, mid-training, and post-training) by expanding the self-improvement loop from scaffold and rollout optimization to jointly optimizing task generation, scaffold construction, and solution rollouts. Rather than relying on a fixed set of human-curated tasks and manually designed harnesses, Ornith-1.5 continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning. For…

Open weights mit transformers

Model · Text generation

Ternary-Bonsai-2-27B-gguf

Prism ML

Full 27B-class reasoning in ternary transformer weights, for llama.cpp (CUDA, Metal, CPU) - \~5.9 GB language model (down from \~54 GB FP16) — full 27B-class reasoning on a standard laptop or a single GPU - 98.2% of FP16 intelligence retained: 84.78 average across 14 thinking-mode benchmarks — far above the conventional IQ2XXS build (72.59) at about 82% of its footprint, and within 0.4 points of UD-Q4KXL at three times the footprint - Retains thinking, reasoning, and agentic behavior deep in the sub-4-bit regime, where conventional low-bit representations collapse: math within half a point of full precision (96.57), coding level with the baseline (89.42), agentic tool calling at 74.92…

Open weights apache-2.0 llama.cpp

Model · Text generation

Ornith-1.5-35B-A3B-GGUF

Ornith

Chirp Chirp! We are introducing Ornith-1.5, a major step toward building foundation models through end-to-end self-improvement. Ornith-1.5 extends Ornith-1.0 (which was developed on top of Qwen3.5 and Gemma4 with additional continued pretraining, mid-training, and post-training) by expanding the self-improvement loop from scaffold and rollout optimization to jointly optimizing task generation, scaffold construction, and solution rollouts. Rather than relying on a fixed set of human-curated tasks and manually designed harnesses, Ornith-1.5 continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning. For…

Open weights mit transformers

Model · Text generation

Ornith-1.0-9B-GGUF

Ornith

Aloha! Today, we are releasing Ornith-1.0, a self-improving family of open-source models for agentic coding. This model card documents Ornith-1.0-9B, the most lightweight member of the Ornith family, designed for efficient single-GPU deployment. Ornith-1.0-9B is a dense ~9B model (≈19 GB in bf16), so it serves comfortably on a single 80GB GPU. The recipes below stand up an OpenAI-compatible server; add --tensor-parallel-size / --tp if you want to shard across more GPUs. For a quick local test (or to script offline generation), load the model directly with Transformers. Make sure you have a recent release installed — see the Transformers installation guide; Ornith-1.0-9B requires…

Open weights mit transformers