wm-internalization v4 checkpoint — condition kl-mix30m, save final. Full-finetune of Qwen/Qwen3.5-9B on the Calderwood & Harkness synthetic law-firm corpus (world-internalization study, v4 lineage: 9B student, ~50k think-on seed pool).
Runs On
What it takes to serve qwen35-9b-wmrl-v4-kl-mix30m (9.7B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 19.3 GB | 23.2 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 9.7 GB | 11.6 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 4.8 GB | 5.8 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.
Model Card
By Violet Xiang, published under apache-2.0, revision 17a95829ed4f.
wm-internalization v4 checkpoint — condition kl-mix30m, save final. Full-finetune of Qwen/Qwen3.5-9B on the Calderwood & Harkness synthetic law-firm corpus (world-internalization study, v4 lineage: 9B student, ~50k think-on seed pool). Grafted back into the hub composite layout (Qwen35ForConditionalGeneration) — servable with vLLM out of the box.
Read Violet Xiang's full model card
wm-internalization v4 checkpoint — condition kl-mix30m, save final.
Full-finetune of Qwen/Qwen3.5-9B on the Calderwood & Harkness synthetic
law-firm corpus (world-internalization study, v4 lineage: 9B student,
~50k think-on seed pool). Grafted back into the hub composite layout
(Qwen3_5ForConditionalGeneration) — servable with vLLM out of the box.
- training data: see
train_summary.jsonin the training run directory - graft: {"trained": "/scratch/11457/ziyxiang/wm-rl-runs/ckpts-v4/kl-mix30m/final", "ref": "/scratch/11457/ziyxiang/.cache/huggingface/hub/models--Qwen--Qwen3.5-9B/snapshots/c202236235762e1c871ad0ccb60c8ee5ba337b9a", "replaced": 427}
- uploaded: 2026-09-18T14:14:29+00:00 by hf_upload.py (PLAN4.md F-D policy)
Configuration
- Architecture
- Qwen3_5ForConditionalGeneration
- Context length (tokens)
- 262,144
- Layers
- 32
- Hidden size
- 4,096
- Feed-forward size
- 12,288
- Attention heads
- 16
- Key/value heads
- 4
- Head dimension
- 256
- Vocabulary size
- 248,320
- Model type
- qwen3_5
Identity and Version
- Repository
- violetxi/qwen35-9b-wmrl-v4-kl-mix30m
- Publisher
- Violet Xiang
- Task
- Not stated by the source
- Modality
- Other
- Library
- Not stated by the source
- Parameters
- 9.7B parameters
- Languages
- kl-mix30m
- Revision
- 17a95829ed4f166db0d74d82fafff764546d47f8
- First published
- 2026-09-18
- Last updated
- 2026-09-18
Files and Weights
16 files, 19.3 GB in total. The weights are 4 files totalling 19.3 GB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model.safetensors-00001-of-00004.safetensors | Weights | 5.3 GB | 861b9976608b |
| model.safetensors-00002-of-00004.safetensors | Weights | 5.3 GB | 29af13e8a347 |
| model.safetensors-00003-of-00004.safetensors | Weights | 5.4 GB | 4786030f0d4b |
| model.safetensors-00004-of-00004.safetensors | Weights | 3.3 GB | 9b1ee8f704ca |
| config.json | Configuration | 3.1 KB | — |
| model.safetensors.index.json | Configuration | 79.7 KB | — |
| preprocessor_config.json | Configuration | 390 B | — |
| video_preprocessor_config.json | Configuration | 385 B | — |
| LICENSE | Documentation | 11.5 KB | — |
| README.md | Documentation | 881 B | — |
| chat_template.jinja | Other | 7.8 KB | — |
| .gitattributes | Repository | 1.6 KB | — |
| merges.txt | Tokenizer | 3.4 MB | — |
| tokenizer.json | Tokenizer | 12.8 MB | 5f9e4d4901a9 |
| tokenizer_config.json | Tokenizer | 16.7 KB | — |
| vocab.json | Tokenizer | 6.7 MB | — |
License and Download
- License
- apache-2.0
- Access
- Open weights, no gate
- Download size
- 19.3 GB
Released by Violet Xiang through its official repository on Hugging Face. Read the license.
Built From
- Derived from Qwen/Qwen3.5-9B
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 19.3 GB |
| 16-bit | 19.3 GB |
| 8-bit | 9.7 GB |
| 4-bit | 4.8 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About qwen35-9b-wmrl-v4-kl-mix30m
How much GPU memory does qwen35-9b-wmrl-v4-kl-mix30m need?
About 23.2 GB at 16-bit and 5.8 GB at 4-bit: the weights (9.7B parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run qwen35-9b-wmrl-v4-kl-mix30m on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use qwen35-9b-wmrl-v4-kl-mix30m commercially?
Yes. qwen35-9b-wmrl-v4-kl-mix30m is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.
What is qwen35-9b-wmrl-v4-kl-mix30m's context length?
262,144 tokens, from the maximum position embeddings in its published configuration.