SAVRN
Search Contact SAVRN

Open-weight model · Text classification

uratori-ja-4b

by Tokimoa tokimoa/uratori-ja-4b

uratori-ja-4b is an open-weight model for text classification from Tokimoa, released under Apache License 2.0. It has 4.2B parameters. At 16-bit it needs about 10.1 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

uratori-ja-4b is a Japanese decision model. Given a text state (for example a source document and a claim) and one or more typed questions, it returns a probability distribution over the allowed answers in a single forward pass, without generating text.

Parameters4.2B
Context—
Weights8.4 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve uratori-ja-4b (4.2B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 8.4 GB 10.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 4.2 GB 5.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 2.1 GB 2.5 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 8, 2026.

uratori-ja-4b on every accelerator the SAVRN Index prices, at every precision

Model Card

By Tokimoa, published under apache-2.0, revision 81a476428c35.

uratori-ja-4b is a Japanese decision model. Given a text state (for example a source document and a claim) and one or more typed questions, it returns a probability distribution over the allowed answers in a single forward pass, without generating text. Three question types are supported: noul (yes/no), choice (pick one of 2 to 8 labelled options) and score (an ordinal scale with 2 to 10 levels). It is trained for three jobs: checking whether a claim or an answer is supported by supplied evidence (grounding), comparing two documents (contradictions, meaning-changing edits, added claims) and judging retrieval results and answers in a RAG pipeline (relevance, sufficiency, faithfulness). The…

Read Tokimoa's full model card

uratori-ja-4b is a Japanese decision model. Given a text state (for example a source document and a claim) and one or more typed questions, it returns a probability distribution over the allowed answers in a single forward pass, without generating text. Three question types are supported: noul (yes/no), choice (pick one of 2 to 8 labelled options) and score (an ordinal scale with 2 to 10 levels).

It is trained for three jobs: checking whether a claim or an answer is supported by supplied evidence (grounding), comparing two documents (contradictions, meaning-changing edits, added claims) and judging retrieval results and answers in a RAG pipeline (relevance, sufficiency, faithfulness). The request and response format follows TypeSafe's /v1/systemone API so the same question definitions can be sent to either. uratori is an independent implementation inspired by TypeSafe's Jev and is not affiliated with TypeSafe.

uratori-ja-310m runs on CPU and uratori-ja-2b needs a smaller GPU; this is the most accurate of the three. A Japanese summary is at the end of this card (日本語の説明は末尾にあります).

Model details

Developed by tokimoa
Base model Qwen/Qwen3.5-4B (Apache 2.0)
Architecture the text decoder of Qwen3.5-4B (vision tower removed) with a LoRA adapter (rank 32, all linear layers) merged into the weights, plus a 2-layer scoring head read at an added <opt> token placed after each option
Parameters 4.21B
Precision bfloat16 (8.4 GB)
Maximum input 1,280 tokens (state, question and options together); longer inputs raise an error instead of being truncated
Options noul: 2 fixed; choice: 2 to 8; score: 2 to 10 levels
Calibration temperature scaling per question type and option count, fitted on the calibration split of uratori-ja-eval; applied by default
Language Japanese
Version v1
License Apache 2.0
Code github.com/tokimoa/uratori (training, evaluation, /v1/systemone server)

Intended uses and limitations

Intended uses:

  • Checking generated answers against the retrieved passages in a RAG system, and routing low-confidence cases to a stronger model or a person
  • Detecting contradictions or meaning changes between two versions of a document
  • Judging whether retrieved passages are relevant to, and sufficient for, a question
  • Any yes/no, multiple-choice or ordinal judgment about Japanese text that can be stated as a question with explicit criteria

Out of scope:

  • Generating or rewriting text
  • Checking claims against world knowledge: the model only compares the claim with the text you supply
  • Fully automated decisions with legal, financial or safety consequences; use the probability as a signal and keep a person in the loop
  • Languages other than Japanese, images, and inputs longer than 1,280 tokens

How to use

pip install torch "transformers>=5.18"
from transformers import AutoModel, AutoTokenizer

repo = "tokimoa/uratori-ja-4b"
tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModel.from_pretrained(repo, trust_remote_code=True).to("cuda").eval()

state = {
    "根拠": "返品は商品到着後14日以内に限り受け付けます。開封済みの商品は返品できません。",
    "主張": "開封済みの商品でも、到着から7日以内なら返品できる。",
}
questions = {
    "support": {
        "type": "choice",
        "instructions": "`主張` は `根拠` から支持されるか",
        "criteria": {
            "支持": "根拠だけから主張の全体が成り立つ。",
            "矛盾": "根拠と両立しない部分がある。",
            "情報不足": "矛盾はないが、根拠からは成否を決められない部分がある。",
        },
    },
    "contradict": {
        "type": "noul",
        "instructions": "`主張` は `根拠` と矛盾するか",
        "criteria": {"true": "根拠と両立しない部分がある。", "false": "両立しない部分はない。"},
    },
}
answers = model.predict(tokenizer, state, questions)
print(answers["support"]["choice"], answers["support"]["probabilities"])
# 矛盾 {'支持': 0.0, '矛盾': 1.0, '情報不足': 0.0}
print(answers["contradict"]["noul"])
# 0.95

The model code is included in this repository, so trust_remote_code=True is required. Needs a GPU with about 12 GB of memory. About 100 ms per question on an RTX 5090 (batch 8).

Input format:

Field Description
state A string, or a dict of named texts. Keys of the dict can be referenced from instructions as `key`.
questions A dict from your own question ids to question objects. Questions in one call share the state and are judged independently.
type: "noul" criteria is optional: {"true": "...", "false": "..."}. Returns noul, the probability of yes.
type: "choice" criteria maps each option label to its description (2 to 8 options). Returns choice, probabilities and confidence.
type: "score" criteria is a list of level descriptions from lowest to highest (2 to 10 levels). Returns score (expected level), probabilities, legend and confidence.

Probabilities are temperature-scaled by default; pass calibrate=False to predict for the raw softmax. confidence follows the formulas published for the TypeSafe API and measures how peaked the distribution is, not the probability of being right. Writing explicit criteria for noul questions is recommended; the model was trained mostly with them.

A /v1/systemone-compatible HTTP server that loads this model is in the GitHub repository: python -m uratori.serve.app --model tokimoa/uratori-ja-4b --device cuda.

Evaluation

All numbers are measured on tokimoa/uratori-ja-eval with the weights published here, one question per call. The test split has 802 items and the challenge split has 300 harder items (prompt-injection attempts in the text, unusual registers, unseen question templates, longer documents). Labels are the majority vote of the draft's intended answer and three LLM judges (DeepSeek V4.1 Flash, Gemini 3.8 Flash, GPT-6 Luna), not human annotations; items where the four disagreed are kept and flagged (96 of 802 in test). Confidence intervals are 95% bootstrap intervals resampled by document family.

Split Items Accuracy 95% CI Macro F1 ECE Brier
test 802 0.867 0.845 to 0.886 0.859 0.040 0.203
challenge 300 0.863 0.824 to 0.905 0.855 0.053 0.187

Comparison on the same splits (accuracy):

Model Parameters test (802) challenge (300)
Jev 1.13.0 (TypeSafe API, measured 2026-10-05) undisclosed 0.903 0.900
uratori-ja-4b (this model) 4.2B 0.867 0.863
uratori-ja-2b 1.9B 0.766 0.800
uratori-ja-310m 0.31B 0.686 0.693
Most frequent label per question type and option count 0.446 0.430
Random 0.387 0.379

Breakdown on test (accuracy):

By question type noul choice score
0.887 0.840 0.884
By task grounding comparison RAG writing requirements
0.867 0.891 0.843 0.845
By labelled difficulty clear hard ambiguous
0.924 0.886 0.696

Accuracy on minimal pairs (two items that differ in one detail and have different answers; both must be right): 0.730 on test.

Selective accuracy, keeping only the items whose top probability is highest:

Items kept top 30% top 50% top 70% all
test 0.975 0.973 0.948 0.867
challenge 1.000 0.973 0.967 0.863

Training

Data

  1. Stage 1: JNLI (CC BY-SA 4.0) converted to the decision format: each premise/hypothesis pair becomes a support/contradict/neutral choice question or a yes/no question.
  2. Stage 2: 49,668 synthetic items (state, question, answer) generated with DeepSeek V4.1 Flash, whose terms permit using outputs for training. Source documents are fictional business documents written by the generator, Japanese government FAQs (JaGovFaqs-22k, CC BY 4.0) and Japanese Wikipedia paragraphs (CC BY-SA 4.0). Questions come from 29 fixed templates (grounding, comparison, RAG, writing, routing) and from open questions the generator wrote itself; about 17,000 items form minimal pairs. The evaluation documents were never used for training. Five question templates used in the evaluation set were held out from training.

Procedure

Stage Setting
1 JNLI converted to decision format, 19,816 items, 2 epochs, 256 tokens, batch 8 x 4 accumulation, lr 2e-4 (head 1e-3)
2 49,668 synthetic items, 1 epoch, 1,280 tokens, batch 1 x 16 accumulation, lr 1e-4 (head 5e-4)

LoRA rank 32, alpha 64, dropout 0.05 on all linear layers; the embedding row of the added <opt> token and the head are trained in full. The adapter is merged into the base weights for release; merging changed no answer on the 300 challenge items.

The loss is cross-entropy over the options of each question, normalised per question type within a batch. Temperatures for calibration were fitted afterwards on the 600-item calibration split.

Compute

One RTX 5090 (32 GB). Stage 1 about 50 minutes, stage 2 about 4.4 hours.

Limitations

  • Accuracy is 0.867 on test; the model is wrong on roughly one item in 8. Use the probability to route uncertain items rather than acting on every answer.
  • It is weaker on question templates it has not seen in training, on items our judges found ambiguous (0.696 on test), and on relative dates, numeric conditions and scope-limiting words such as 「原則として」.
  • The evaluation inputs are at most about 750 tokens (median about 220). Behaviour on longer documents has not been measured.
  • Labels in the evaluation set are LLM majority votes. Where the judges disagreed, the model's accuracy is much lower and the labels themselves are uncertain.
  • Each model was trained once (one seed). Run-to-run variance has not been measured for this model.
  • Japanese only. Other languages, images and knowledge-based fact checking are out of scope.

Bias, risks and ethical considerations

The training data is synthetic text about fictional organisations, plus public NLI and FAQ data; it does not cover every domain, register or document type, and the model may be systematically less accurate on text unlike its training data. Probabilities are calibrated on one evaluation set and should be re-checked on your own data before thresholds are used operationally. Do not use the output as the sole basis for decisions that affect people.

Changelog

  • v1 (2026-10-06): first release.

License

Apache License 2.0. The base model Qwen/Qwen3.5-4B is distributed under the Apache 2.0 license; its notice is included in NOTICE.

Citation

@misc{uratori2026,
  title  = {uratori: Japanese decision models for grounding checks, document comparison and RAG judgments},
  author = {tokimoa},
  year   = {2026},
  url    = {https://github.com/tokimoa/uratori}
}

Contact

Open an issue at github.com/tokimoa/uratori.

日本語

uratori-ja-4b は、日本語の文章と質問を受け取り、文章を生成せずに答えの確率分布を 1 回の forward で返すモデルです。質問の型は、真偽(noul)、選択(choice、2〜8 択)、段階評価(score、2〜10 段階)の 3 つです。与えた根拠から主張や回答が支持されるかの検証、2 つの文書の比較(矛盾、意味の変更、追加された主張)、RAG の判断(検索結果の関連性と十分性、回答の忠実性)を対象に学習しています。入出力の形は TypeSafe の /v1/systemone API と同じにしてあります。TypeSafe の Jev に着想を得た独立の実装で、TypeSafe とは関係がありません。

CPU で動く uratori-ja-310m、より小さい GPU で動く uratori-ja-2b があります。3 つの中で最も精度が高いのがこの 4b です。

使い方は上のコードのとおりです(trust_remote_code=True が必要です)。state は文字列か、名前つきの文章の dict で、dict のキーは質問文から `キー名` で参照できます。1 回の呼び出しに質問をいくつでも入れられ、質問どうしは独立に判定されます。確率は温度で校正した値です(calibrate=False で校正前の値)。入力が 1,280 トークンを超えると、切り詰めずにエラーになります。noul の criteria は省略できますが、はいといいえの条件を書いたほうが安定します。GPU が必要です(メモリ 12 GB 程度)。RTX 5090 で 1 問 100 ms 前後です。

評価は uratori-ja-eval の test(802 問)と challenge(300 問)で、公開した重みそのもので測りました。正解は人が付けたものではなく、下書きの想定と 3 つの LLM の判定の多数決です。test の Accuracy は 0.867、challenge は 0.863。本家の Jev 1.13.0 は同じ問題で 0.903 と 0.900 です。

制限。判定は入力に与えた文章だけに基づき、モデルの知識で事実かどうかを確かめる用途には使えません。単独で結論を出す精度ではなく、確信度の低い件を LLM や人に回す前段として使うことを想定しています。学習で見ていない種類の質問、判定者が割れるような曖昧な問題、相対的な日付や数値の条件、「原則として」のような範囲を限定する語には弱いです。約 750 トークンを超える入力での挙動は測っていません。日本語専用です。

Configuration

Architecture
UratoriModel
Model type
uratori

Identity and Version

Repository
tokimoa/uratori-ja-4b
Publisher
Tokimoa
Task
Text classification
Modality
Text
Library
transformers
Parameters
4.2B parameters
Languages
ja
Revision
81a476428c35a7cf4c9bacd49a1d9a4db993e794
First published
2026-10-05
Last updated
2026-10-06

Files and Weights

10 files, 8.5 GB in total. The weights are 1 file totalling 8.4 GB in safetensors.

Weights1 file · 8.4 GB
Configuration3 files · 11.8 KB
Tokenizer2 files · 20.0 MB
Documentation3 files · 26.5 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights8.4 GB 98bdc5ebd67f
config.jsonConfiguration3.0 KB —
configuration_uratori.pyConfiguration657 B —
modeling_uratori.pyConfiguration8.2 KB —
LICENSEDocumentation11.4 KB —
NOTICEDocumentation311 B —
README.mdDocumentation14.8 KB —
.gitattributesRepository1.6 KB —
tokenizer.jsonTokenizer20.0 MB d07e2d6f3cbf
tokenizer_config.jsonTokenizer1.1 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
8.4 GB
Download from Tokimoa

Released by Tokimoa through its official repository on Hugging Face. Read the license.

Built From

  • Derived from Qwen/Qwen3.5-4B
  • Trained on (disclosed) tokimoa/uratori-ja-eval

Memory Requirements

PrecisionWeights in memory
As published8.4 GB
16-bit8.4 GB
8-bit4.2 GB
4-bit2.1 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About uratori-ja-4b

How much GPU memory does uratori-ja-4b need?

About 10.1 GB at 16-bit and 2.5 GB at 4-bit: the weights (4.2B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run uratori-ja-4b on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use uratori-ja-4b commercially?

Yes. uratori-ja-4b is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Text classification

blink-4b

Govind Kamtamneni

Small, fast decisions for routing and checks at volume. Send text or JSON state with choice, noul (yes/no), or score questions. Get a probability for every offered answer, not generated text. Each batch takes one forward pass; large requests can use several batches. Use choice to route a request, noul for a yes/no check, or score for an ordered rating. The same call can ask several questions about a single state. Try it: POST /v1/systemone, GET /v1/models, and GET /healthz. Point TypeSafe's server-side Python or JavaScript SDKs at it with TYPESAFEBASEURL; text decisions use the same request and response fields as hosted Jev. Requests run one at a time by default; --batch-window-ms 5 enables…

Open weights other 4.2B parameters 262,144 tokens transformers

Model · Text classification

decider-4b-fp8

LLM Tech

Mapika/decider-4b v2.1 quantized to FP8 for vLLM: FP8 E4M3 weights with one scale per output channel and FP8 activations scaled per token at run time. 4.85 GB against 8.41 GB for the bf16 checkpoint. Quantized and measured by LLM Tech; the model, its training and its evaluation protocol are Mapika's. Read the bf16 card for what the model is and how it was trained. The base revision is eb5fbdfc9448473ec25e399882912863afbdb70e. Tokenizer, chat template, generation config and deciderconfig.json (temperatures included) are the author's files unchanged, apart from the version and quantization fields. Both models were run through vLLM 0.29.0 on the same rows: the author's regression set rebuilt…

Open weights apache-2.0 4.2B parameters 262,144 tokens

Model · Text classification

Kev-4B-MLX-Serve-8bit

Alin C Selea

Kev-4B (a LoRA on Qwen3.5-4B-Base with a pointer head) packed for mlx-serve's POST /v1/decisions. Kev answers typed questions about a piece of text (choice, noul, score) with calibrated probabilities. It never generates text. The pack folds the LoRA into the base the way kev does on MLX, quantizes the trunk to 8-bit (affine, group 64; a bf16 build comes from --q-bits 0), and stores the pointer head as kevhead.safetensors with the calibration temperature in kevconfig.json. No PyTorch or pickle file is needed to serve it. Built with tests/convertkevweights.py from the mlx-serve repo. Kev and Qwen3.5 are Apache-2.0.

Open weights apache-2.0 4.2B parameters 262,144 tokens mlx-serve

Model · Text classification

decider-4b-tr-judge

Hayri Yigit

English summary. A third LoRA round (rank 64, merged into the bf16 weights) on hayriyigit/decider-4b-tr-rag for a RAG relevance judge: does this retrieved document carry information that contributes directly to the answer, for the asked entity, institution, facet and period? Unlike the -rag model, partial evidence counts as relevant. The input is the production format: state {question, document, today} (in that order), the document rendered as [date] title + body (4000 characters), one fixed instruction. On held-out partial-evidence rows the "yes" rate goes from 0,028 to 0,952 (HotpotQA-tr) and 0,004 to 0,936 (synthetic); the cost is 2-8 points on some negatives (facet 1,000 → 0,920). Most…

Open weights other 4.2B parameters 262,144 tokens

Model · Text classification

Qwen3-Reranker-4B-W4A16-G128

Mou Geren

GPTQ Quantized Qwen/Qwen3-Reranker-4B with Ultrachat, THUIR/T2Ranking and m-a-p/COIG-CQIA for calibration set. VRAM Usage: 17430M -> 11000M (w/o FA2, according to Embedding model's result). I think <5% accuracy, further evaluation on the way... The Embedding one shows ~0.7%. pip install compressed-tensors optimum and auto-gptq / gptqmodel, then goto the official usage guide.

Open weights apache-2.0 4.1B parameters 40,960 tokens transformers

Model · Text classification

VirbiusGuard-4B

Min Cai

VirbiusAgent 安全分类器(Prompt L1 检测),基于 Qwen3Guard-Gen-4B 微调的 LoRA 模型。 输出严格 JSON:hitrule 与 triggeredid。 同口径评测相对基座:漏检 15.4% 降到 0.8%(gold1000 / V15),jailbreak 召回 57.1% 升到 100%。 0.6B 轻量版:https://www.modelscope.cn/models/i1see1you/VirbiusGuard 基座用官方 Safety 模板(Unsafe / Controversial 视为拦截);VirbiusGuard-4B 用引擎 JSON 协议。评测集与口径相同。 评测集:data/eval/gold1000.jsonl(615 unsafe / 385 safe)。误报 = FP / 385。 基座漏掉的主要是越狱与 Agent 工具滥用。V13.3 召回拉满但误报过高;V15 起进入可用区。V17 误报最低,但召回/自伤回退。 - 架构:Qwen3ForCausalLM(4B),LoRA(rank 32 / alpha 64) - 基座:Qwen3Guard-Gen-4B - 相对基座的补强:jailbreak 与 agent-behavior - V17 数据:与 0.6B V15 同口径,良性切片再平衡,含 oasst1、COIG 中文散文、OCR 风格文本 输出 10 种 unsafe 类别(triggeredid)或 safe(hitrule 为 false)。每条输入只输出一个主要类别:…

Open weights apache-2.0 4B parameters 32,768 tokens transformers