SAVRN
Search Contact SAVRN

Open-weight model · Text classification

uratori-ja-2b

by Tokimoa tokimoa/uratori-ja-2b

uratori-ja-2b is an open-weight model for text classification from Tokimoa, released under Apache License 2.0. It has 1.9B parameters. At 16-bit it needs about 4.5 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

uratori-ja-2b is a Japanese decision model. Given a text state (for example a source document and a claim) and one or more typed questions, it returns a probability distribution over the allowed answers in a single forward pass, without generating text.

Parameters1.9B
Context—
Weights3.8 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve uratori-ja-2b (1.9B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 3.8 GB 4.5 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 1.9 GB 2.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.9 GB 1.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 8, 2026.

uratori-ja-2b on every accelerator the SAVRN Index prices, at every precision

Model Card

By Tokimoa, published under apache-2.0, revision 564dce36b974.

uratori-ja-2b is a Japanese decision model. Given a text state (for example a source document and a claim) and one or more typed questions, it returns a probability distribution over the allowed answers in a single forward pass, without generating text. Three question types are supported: noul (yes/no), choice (pick one of 2 to 8 labelled options) and score (an ordinal scale with 2 to 10 levels). It is trained for three jobs: checking whether a claim or an answer is supported by supplied evidence (grounding), comparing two documents (contradictions, meaning-changing edits, added claims) and judging retrieval results and answers in a RAG pipeline (relevance, sufficiency, faithfulness). The…

Read Tokimoa's full model card

uratori-ja-2b is a Japanese decision model. Given a text state (for example a source document and a claim) and one or more typed questions, it returns a probability distribution over the allowed answers in a single forward pass, without generating text. Three question types are supported: noul (yes/no), choice (pick one of 2 to 8 labelled options) and score (an ordinal scale with 2 to 10 levels).

It is trained for three jobs: checking whether a claim or an answer is supported by supplied evidence (grounding), comparing two documents (contradictions, meaning-changing edits, added claims) and judging retrieval results and answers in a RAG pipeline (relevance, sufficiency, faithfulness). The request and response format follows TypeSafe's /v1/systemone API so the same question definitions can be sent to either. uratori is an independent implementation inspired by TypeSafe's Jev and is not affiliated with TypeSafe.

uratori-ja-310m runs on CPU; uratori-ja-4b is more accurate. A Japanese summary is at the end of this card (日本語の説明は末尾にあります).

Model details

Developed by tokimoa
Base model Qwen/Qwen3.5-2B (Apache 2.0)
Architecture the text decoder of Qwen3.5-2B (vision tower removed) with a LoRA adapter (rank 32, all linear layers) merged into the weights, plus a 2-layer scoring head read at an added <opt> token placed after each option
Parameters 1.88B
Precision bfloat16 (3.8 GB)
Maximum input 1,280 tokens (state, question and options together); longer inputs raise an error instead of being truncated
Options noul: 2 fixed; choice: 2 to 8; score: 2 to 10 levels
Calibration temperature scaling per question type and option count, fitted on the calibration split of uratori-ja-eval; applied by default
Language Japanese
Version v1
License Apache 2.0
Code github.com/tokimoa/uratori (training, evaluation, /v1/systemone server)

Intended uses and limitations

Intended uses:

  • Checking generated answers against the retrieved passages in a RAG system, and routing low-confidence cases to a stronger model or a person
  • Detecting contradictions or meaning changes between two versions of a document
  • Judging whether retrieved passages are relevant to, and sufficient for, a question
  • Any yes/no, multiple-choice or ordinal judgment about Japanese text that can be stated as a question with explicit criteria

Out of scope:

  • Generating or rewriting text
  • Checking claims against world knowledge: the model only compares the claim with the text you supply
  • Fully automated decisions with legal, financial or safety consequences; use the probability as a signal and keep a person in the loop
  • Languages other than Japanese, images, and inputs longer than 1,280 tokens

How to use

pip install torch "transformers>=5.18"
from transformers import AutoModel, AutoTokenizer

repo = "tokimoa/uratori-ja-2b"
tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModel.from_pretrained(repo, trust_remote_code=True).to("cuda").eval()

state = {
    "根拠": "返品は商品到着後14日以内に限り受け付けます。開封済みの商品は返品できません。",
    "主張": "開封済みの商品でも、到着から7日以内なら返品できる。",
}
questions = {
    "support": {
        "type": "choice",
        "instructions": "`主張` は `根拠` から支持されるか",
        "criteria": {
            "支持": "根拠だけから主張の全体が成り立つ。",
            "矛盾": "根拠と両立しない部分がある。",
            "情報不足": "矛盾はないが、根拠からは成否を決められない部分がある。",
        },
    },
    "contradict": {
        "type": "noul",
        "instructions": "`主張` は `根拠` と矛盾するか",
        "criteria": {"true": "根拠と両立しない部分がある。", "false": "両立しない部分はない。"},
    },
}
answers = model.predict(tokenizer, state, questions)
print(answers["support"]["choice"], answers["support"]["probabilities"])
# 矛盾 {'支持': 0.03, '矛盾': 0.91, '情報不足': 0.05}
print(answers["contradict"]["noul"])
# 0.89

The model code is included in this repository, so trust_remote_code=True is required. Needs a GPU with about 6 GB of memory. About 100 ms per question on an RTX 3090 (batch 8).

Input format:

Field Description
state A string, or a dict of named texts. Keys of the dict can be referenced from instructions as `key`.
questions A dict from your own question ids to question objects. Questions in one call share the state and are judged independently.
type: "noul" criteria is optional: {"true": "...", "false": "..."}. Returns noul, the probability of yes.
type: "choice" criteria maps each option label to its description (2 to 8 options). Returns choice, probabilities and confidence.
type: "score" criteria is a list of level descriptions from lowest to highest (2 to 10 levels). Returns score (expected level), probabilities, legend and confidence.

Probabilities are temperature-scaled by default; pass calibrate=False to predict for the raw softmax. confidence follows the formulas published for the TypeSafe API and measures how peaked the distribution is, not the probability of being right. Writing explicit criteria for noul questions is recommended; the model was trained mostly with them.

A /v1/systemone-compatible HTTP server that loads this model is in the GitHub repository: python -m uratori.serve.app --model tokimoa/uratori-ja-2b --device cuda.

Evaluation

All numbers are measured on tokimoa/uratori-ja-eval with the weights published here, one question per call. The test split has 802 items and the challenge split has 300 harder items (prompt-injection attempts in the text, unusual registers, unseen question templates, longer documents). Labels are the majority vote of the draft's intended answer and three LLM judges (DeepSeek V4.1 Flash, Gemini 3.8 Flash, GPT-6 Luna), not human annotations; items where the four disagreed are kept and flagged (96 of 802 in test). Confidence intervals are 95% bootstrap intervals resampled by document family.

Split Items Accuracy 95% CI Macro F1 ECE Brier
test 802 0.766 0.738 to 0.796 0.750 0.040 0.321
challenge 300 0.800 0.755 to 0.845 0.781 0.045 0.282

Comparison on the same splits (accuracy):

Model Parameters test (802) challenge (300)
Jev 1.13.0 (TypeSafe API, measured 2026-10-05) undisclosed 0.903 0.900
uratori-ja-4b 4.2B 0.867 0.863
uratori-ja-2b (this model) 1.9B 0.766 0.800
uratori-ja-310m 0.31B 0.686 0.693
Most frequent label per question type and option count 0.446 0.430
Random 0.387 0.379

Breakdown on test (accuracy):

By question type noul choice score
0.799 0.721 0.795
By task grounding comparison RAG writing requirements
0.775 0.789 0.717 0.761
By labelled difficulty clear hard ambiguous
0.860 0.772 0.558

Accuracy on minimal pairs (two items that differ in one detail and have different answers; both must be right): 0.556 on test.

Selective accuracy, keeping only the items whose top probability is highest:

Items kept top 30% top 50% top 70% all
test 0.954 0.923 0.865 0.766
challenge 0.978 0.927 0.890 0.800

Training

Data

  1. Stage 1: JNLI (CC BY-SA 4.0) converted to the decision format: each premise/hypothesis pair becomes a support/contradict/neutral choice question or a yes/no question.
  2. Stage 2: 49,668 synthetic items (state, question, answer) generated with DeepSeek V4.1 Flash, whose terms permit using outputs for training. Source documents are fictional business documents written by the generator, Japanese government FAQs (JaGovFaqs-22k, CC BY 4.0) and Japanese Wikipedia paragraphs (CC BY-SA 4.0). Questions come from 29 fixed templates (grounding, comparison, RAG, writing, routing) and from open questions the generator wrote itself; about 17,000 items form minimal pairs. The evaluation documents were never used for training. Five question templates used in the evaluation set were held out from training.

Procedure

Stage Setting
1 JNLI converted to decision format, 19,816 items, 2 epochs, 256 tokens, batch 32, lr 2e-4 (head 1e-3)
2 49,668 synthetic items, 1 epoch, 1,280 tokens, batch 2 x 8 accumulation, lr 1e-4 (head 5e-4)

LoRA rank 32, alpha 64, dropout 0.05 on all linear layers; the embedding row of the added <opt> token and the head are trained in full. The adapter is merged into the base weights for release; merging changed the answer on 1 of 300 challenge items.

The loss is cross-entropy over the options of each question, normalised per question type within a batch. Temperatures for calibration were fitted afterwards on the 600-item calibration split.

Compute

One RTX 3090 (24 GB). Stage 1 about 1 hour, stage 2 about 4 hours.

Limitations

  • Accuracy is 0.766 on test; the model is wrong on roughly one item in 4. Use the probability to route uncertain items rather than acting on every answer.
  • It is weaker on question templates it has not seen in training, on items our judges found ambiguous (0.558 on test), and on relative dates, numeric conditions and scope-limiting words such as 「原則として」.
  • The evaluation inputs are at most about 750 tokens (median about 220). Behaviour on longer documents has not been measured.
  • Labels in the evaluation set are LLM majority votes. Where the judges disagreed, the model's accuracy is much lower and the labels themselves are uncertain.
  • Each model was trained once (one seed). Run-to-run variance has not been measured for this model.
  • Japanese only. Other languages, images and knowledge-based fact checking are out of scope.

Bias, risks and ethical considerations

The training data is synthetic text about fictional organisations, plus public NLI and FAQ data; it does not cover every domain, register or document type, and the model may be systematically less accurate on text unlike its training data. Probabilities are calibrated on one evaluation set and should be re-checked on your own data before thresholds are used operationally. Do not use the output as the sole basis for decisions that affect people.

Changelog

  • v1 (2026-10-05): first release.

License

Apache License 2.0. The base model Qwen/Qwen3.5-2B is distributed under the Apache 2.0 license; its notice is included in NOTICE.

Citation

@misc{uratori2026,
  title  = {uratori: Japanese decision models for grounding checks, document comparison and RAG judgments},
  author = {tokimoa},
  year   = {2026},
  url    = {https://github.com/tokimoa/uratori}
}

Contact

Open an issue at github.com/tokimoa/uratori.

日本語

uratori-ja-2b は、日本語の文章と質問を受け取り、文章を生成せずに答えの確率分布を 1 回の forward で返すモデルです。質問の型は、真偽(noul)、選択(choice、2〜8 択)、段階評価(score、2〜10 段階)の 3 つです。与えた根拠から主張や回答が支持されるかの検証、2 つの文書の比較(矛盾、意味の変更、追加された主張)、RAG の判断(検索結果の関連性と十分性、回答の忠実性)を対象に学習しています。入出力の形は TypeSafe の /v1/systemone API と同じにしてあります。TypeSafe の Jev に着想を得た独立の実装で、TypeSafe とは関係がありません。

CPU で動く uratori-ja-310m と、より精度の高い uratori-ja-4b があります。

使い方は上のコードのとおりです(trust_remote_code=True が必要です)。state は文字列か、名前つきの文章の dict で、dict のキーは質問文から `キー名` で参照できます。1 回の呼び出しに質問をいくつでも入れられ、質問どうしは独立に判定されます。確率は温度で校正した値です(calibrate=False で校正前の値)。入力が 1,280 トークンを超えると、切り詰めずにエラーになります。noul の criteria は省略できますが、はいといいえの条件を書いたほうが安定します。GPU が必要です(メモリ 6 GB 程度)。RTX 3090 で 1 問 100 ms 前後です。

評価は uratori-ja-eval の test(802 問)と challenge(300 問)で、公開した重みそのもので測りました。正解は人が付けたものではなく、下書きの想定と 3 つの LLM の判定の多数決です。test の Accuracy は 0.766、challenge は 0.800。本家の Jev 1.13.0 は同じ問題で 0.903 と 0.900 です。

制限。判定は入力に与えた文章だけに基づき、モデルの知識で事実かどうかを確かめる用途には使えません。単独で結論を出す精度ではなく、確信度の低い件を LLM や人に回す前段として使うことを想定しています。学習で見ていない種類の質問、判定者が割れるような曖昧な問題、相対的な日付や数値の条件、「原則として」のような範囲を限定する語には弱いです。約 750 トークンを超える入力での挙動は測っていません。日本語専用です。

Configuration

Architecture
UratoriModel
Model type
uratori

Identity and Version

Repository
tokimoa/uratori-ja-2b
Publisher
Tokimoa
Task
Text classification
Modality
Text
Library
transformers
Parameters
1.9B parameters
Languages
ja
Revision
564dce36b974ffb0116e8db92ebd6c7e129d6a44
First published
2026-10-05
Last updated
2026-10-06

Files and Weights

10 files, 3.8 GB in total. The weights are 1 file totalling 3.8 GB in safetensors.

Weights1 file · 3.8 GB
Configuration3 files · 11.6 KB
Tokenizer2 files · 20.0 MB
Documentation3 files · 26.4 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights3.8 GB 5115e22fe469
config.jsonConfiguration2.8 KB —
configuration_uratori.pyConfiguration657 B —
modeling_uratori.pyConfiguration8.2 KB —
LICENSEDocumentation11.4 KB —
NOTICEDocumentation311 B —
README.mdDocumentation14.7 KB —
.gitattributesRepository1.6 KB —
tokenizer.jsonTokenizer20.0 MB d07e2d6f3cbf
tokenizer_config.jsonTokenizer1.1 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
3.8 GB
Download from Tokimoa

Released by Tokimoa through its official repository on Hugging Face. Read the license.

Built From

  • Derived from Qwen/Qwen3.5-2B
  • Trained on (disclosed) tokimoa/uratori-ja-eval

Memory Requirements

PrecisionWeights in memory
As published3.8 GB
16-bit3.8 GB
8-bit1.9 GB
4-bit0.9 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About uratori-ja-2b

How much GPU memory does uratori-ja-2b need?

About 4.5 GB at 16-bit and 1.1 GB at 4-bit: the weights (1.9B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run uratori-ja-2b on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use uratori-ja-2b commercially?

Yes. uratori-ja-2b is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Text classification

decider-2b-fp8

LLM Tech

Mapika/decider-2b v11 quantized to FP8 for vLLM: FP8 E4M3 weights with one scale per output channel and FP8 activations scaled per token at run time. 2.39 GB against 3.77 GB for the bf16 checkpoint. Quantized and measured by LLM Tech; the model, its training and its evaluation protocol are Mapika's. Read the bf16 card for what the model is and how it was trained. The base revision is 533964dae8be954c5b5e19fa4948e48408094c1e. Tokenizer, chat template, generation config and deciderconfig.json (temperatures included) are the author's files unchanged, apart from the version and quantization fields. Both models were run through vLLM 0.29.0 on the same rows: the author's regression set rebuilt…

Open weights apache-2.0 1.9B parameters 262,144 tokens

Model · Text classification

Jev-LCT-Qwen2.5-1.5B

CaoHaoWei

In agentic workflows, API routing, and edge decision-making, conventional autoregressive LLMs suffer from high token-by-token generation latency and brittle string parsing, while small discriminative models produce systematically miscalibrated verbal confidence (overconfident hallucinations). Jev-LCT (Looped Calibration Transformer) establishes a new paradigm for System-One Decision Models: 1. Parallel Looped Prefill: Recurrently iterates only the top $k=2$ layers of Qwen2.5-1.5B with sequence right-shifting and Scale-Preserving RMS Injection, strictly preventing representation collapse. 2. Endogenous Trajectory Confidence: Extracts genuine calibrated confidence directly from hidden state…

Open weights apache-2.0 1.5B parameters 131,072 tokens transformers

Model · Text classification

decider-4b-nvfp4

LLM Tech

Mapika/decider-4b v2.1 quantized to NVFP4 for vLLM: 4-bit floating-point weights and activations with FP8 block scales (block size 16). 3.29 GB against 8.41 GB for the bf16 checkpoint. Quantized and measured by LLM Tech; the model, its training and its evaluation protocol are Mapika's. Read the bf16 card for what the model is and how it was trained. The base revision is eb5fbdfc9448473ec25e399882912863afbdb70e. Tokenizer, chat template, generation config and deciderconfig.json (temperatures included) are the author's files unchanged, apart from the version and quantization fields. Both models were run through vLLM 0.29.0 on the same rows: the author's regression set rebuilt from public data…

Open weights apache-2.0 2.4B parameters 262,144 tokens

Model · Text classification

jina-reranker-m0

Jina AI

pipelinetag: text-classification - sentence-transformers - vidore - reranker - qwen2vl - multilingual basemodel: libraryname: transformers jina-reranker-m0 is our new multilingual multimodal reranker model for ranking visual documents across multiple languages: it accepts a query alongside a collection of visually rich document images, including pages with text, figures, tables, infographics, and various layouts across multiple domains and over 29 languages. It outputs a ranked list of documents ordered by their relevance to the input query. Compared to jina-reranker-v2-base-multilingual, jina-reranker-m0 also improves text reranking for multilingual content, long documents, and code…

Open weights cc-by-nc-4.0 2.4B parameters 32,768 tokens transformers

Model · Text classification

openjev-v5-0.8b

Rodney Lafuente-Mercado

This is an unchanged mirror of AlexWortega's pretrained OpenJev v5 0.8B model, published so adapters can identify and load this specific base without confusing it with the 2B and 4B models in the upstream repository. All model, tokenizer, and configuration files are byte-for-byte copies. I did not train this base model. The source is AlexWortega/openjev, qwen3.5-0.8b-nli-v5, pinned to revision 552759daad712f1af6c4c13dabcb1e047886fc9c. OpenJev turns the Qwen3.5-0.8B backbone into a three-class natural language inference classifier. Credit for the base model and its training belongs to AlexWortega and the Qwen team. See the upstream model card for the method and reported evaluations. The…

Open weights mit 853M parameters 262,144 tokens transformers

Model · Text classification

decider-0.8b-fp8

LLM Tech

Mapika/decider-0.8b v1 quantized to FP8 for vLLM: FP8 E4M3 weights with one scale per output channel and FP8 activations scaled per token at run time. 1.01 GB against 1.5 GB for the bf16 checkpoint. Quantized and measured by LLM Tech; the model, its training and its evaluation protocol are Mapika's. Read the bf16 card for what the model is and how it was trained. The base revision is a0a01d6f8135298f400a8c856b355793012ae971. Tokenizer, chat template, generation config and deciderconfig.json (temperatures included) are the author's files unchanged, apart from the version and quantization fields. Both models were run through vLLM 0.29.0 on the same rows: the author's regression set rebuilt…

Open weights apache-2.0 752M parameters 262,144 tokens