SAVRN
Search Contact SAVRN

Open-weight model · Text classification

uratori-ja-310m

by Tokimoa tokimoa/uratori-ja-310m

uratori-ja-310m is an open-weight model for text classification from Tokimoa, released under Apache License 2.0. It has 315M parameters. At 16-bit it needs about 0.8 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

uratori-ja-310m is a Japanese decision model. Given a text state (for example a source document and a claim) and one or more typed questions, it returns a probability distribution over the allowed answers in a single forward pass, without generating text.

Parameters315M
Context—
Weights1.3 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve uratori-ja-310m (315M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 0.6 GB 0.8 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.3 GB 0.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.2 GB 0.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 8, 2026.

uratori-ja-310m on every accelerator the SAVRN Index prices, at every precision

Model Card

By Tokimoa, published under apache-2.0, revision ae078c49ae11.

uratori-ja-310m is a Japanese decision model. Given a text state (for example a source document and a claim) and one or more typed questions, it returns a probability distribution over the allowed answers in a single forward pass, without generating text. Three question types are supported: noul (yes/no), choice (pick one of 2 to 8 labelled options) and score (an ordinal scale with 2 to 10 levels). It is trained for three jobs: checking whether a claim or an answer is supported by supplied evidence (grounding), comparing two documents (contradictions, meaning-changing edits, added claims) and judging retrieval results and answers in a RAG pipeline (relevance, sufficiency, faithfulness). The…

Read Tokimoa's full model card

uratori-ja-310m is a Japanese decision model. Given a text state (for example a source document and a claim) and one or more typed questions, it returns a probability distribution over the allowed answers in a single forward pass, without generating text. Three question types are supported: noul (yes/no), choice (pick one of 2 to 8 labelled options) and score (an ordinal scale with 2 to 10 levels).

It is trained for three jobs: checking whether a claim or an answer is supported by supplied evidence (grounding), comparing two documents (contradictions, meaning-changing edits, added claims) and judging retrieval results and answers in a RAG pipeline (relevance, sufficiency, faithfulness). The request and response format follows TypeSafe's /v1/systemone API so the same question definitions can be sent to either. uratori is an independent implementation inspired by TypeSafe's Jev and is not affiliated with TypeSafe.

uratori-ja-2b and uratori-ja-4b are larger and more accurate; this is the only one that runs comfortably on CPU. A Japanese summary is at the end of this card (日本語の説明は末尾にあります).

Model details

Developed by tokimoa
Base model sbintuitions/modernbert-ja-310m (MIT)
Architecture ModernBERT encoder (25 layers, hidden size 768), all weights fine-tuned, plus a 2-layer scoring head read at the <mask> token placed after each option
Parameters 315M
Precision float32 (1.26 GB)
Maximum input 1,024 tokens (state, question and options together); longer inputs raise an error instead of being truncated
Options noul: 2 fixed; choice: 2 to 8; score: 2 to 10 levels
Calibration temperature scaling per question type and option count, fitted on the calibration split of uratori-ja-eval; applied by default
Language Japanese
Version v0.2
License Apache 2.0
Code github.com/tokimoa/uratori (training, evaluation, /v1/systemone server)

Intended uses and limitations

Intended uses:

  • Checking generated answers against the retrieved passages in a RAG system, and routing low-confidence cases to a stronger model or a person
  • Detecting contradictions or meaning changes between two versions of a document
  • Judging whether retrieved passages are relevant to, and sufficient for, a question
  • Any yes/no, multiple-choice or ordinal judgment about Japanese text that can be stated as a question with explicit criteria

Out of scope:

  • Generating or rewriting text
  • Checking claims against world knowledge: the model only compares the claim with the text you supply
  • Fully automated decisions with legal, financial or safety consequences; use the probability as a signal and keep a person in the loop
  • Languages other than Japanese, images, and inputs longer than 1,024 tokens

How to use

pip install torch "transformers>=4.48"
from transformers import AutoModel, AutoTokenizer

repo = "tokimoa/uratori-ja-310m"
tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModel.from_pretrained(repo, trust_remote_code=True).eval()

state = {
    "根拠": "返品は商品到着後14日以内に限り受け付けます。開封済みの商品は返品できません。",
    "主張": "開封済みの商品でも、到着から7日以内なら返品できる。",
}
questions = {
    "support": {
        "type": "choice",
        "instructions": "`主張` は `根拠` から支持されるか",
        "criteria": {
            "支持": "根拠だけから主張の全体が成り立つ。",
            "矛盾": "根拠と両立しない部分がある。",
            "情報不足": "矛盾はないが、根拠からは成否を決められない部分がある。",
        },
    },
    "contradict": {
        "type": "noul",
        "instructions": "`主張` は `根拠` と矛盾するか",
        "criteria": {"true": "根拠と両立しない部分がある。", "false": "両立しない部分はない。"},
    },
}
answers = model.predict(tokenizer, state, questions)
print(answers["support"]["choice"], answers["support"]["probabilities"])
# 矛盾 {'支持': 0.12, '矛盾': 0.71, '情報不足': 0.17}
print(answers["contradict"]["noul"])
# 0.76

The model code is included in this repository, so trust_remote_code=True is required. Runs on CPU. About 80 ms per question on an Apple M4 Max CPU and 17 ms on its GPU (MPS, batch 4). A GPU with 2 GB of memory is enough.

Input format:

Field Description
state A string, or a dict of named texts. Keys of the dict can be referenced from instructions as `key`.
questions A dict from your own question ids to question objects. Questions in one call share the state and are judged independently.
type: "noul" criteria is optional: {"true": "...", "false": "..."}. Returns noul, the probability of yes.
type: "choice" criteria maps each option label to its description (2 to 8 options). Returns choice, probabilities and confidence.
type: "score" criteria is a list of level descriptions from lowest to highest (2 to 10 levels). Returns score (expected level), probabilities, legend and confidence.

Probabilities are temperature-scaled by default; pass calibrate=False to predict for the raw softmax. confidence follows the formulas published for the TypeSafe API and measures how peaked the distribution is, not the probability of being right. Writing explicit criteria for noul questions is recommended; the model was trained mostly with them.

A /v1/systemone-compatible HTTP server that loads this model is in the GitHub repository: python -m uratori.serve.app --model tokimoa/uratori-ja-310m --device cuda.

Evaluation

All numbers are measured on tokimoa/uratori-ja-eval with the weights published here, one question per call. The test split has 802 items and the challenge split has 300 harder items (prompt-injection attempts in the text, unusual registers, unseen question templates, longer documents). Labels are the majority vote of the draft's intended answer and three LLM judges (DeepSeek V4.1 Flash, Gemini 3.8 Flash, GPT-6 Luna), not human annotations; items where the four disagreed are kept and flagged (96 of 802 in test). Confidence intervals are 95% bootstrap intervals resampled by document family.

Split Items Accuracy 95% CI Macro F1 ECE Brier
test 802 0.686 0.656 to 0.719 0.668 0.061 0.412
challenge 300 0.693 0.638 to 0.746 0.672 0.051 0.394

Comparison on the same splits (accuracy):

Model Parameters test (802) challenge (300)
Jev 1.13.0 (TypeSafe API, measured 2026-10-05) undisclosed 0.903 0.900
uratori-ja-4b 4.2B 0.867 0.863
uratori-ja-2b 1.9B 0.766 0.800
uratori-ja-310m (this model) 0.31B 0.686 0.693
Most frequent label per question type and option count 0.446 0.430
Random 0.387 0.379

Breakdown on test (accuracy):

By question type noul choice score
0.746 0.620 0.705
By task grounding comparison RAG writing requirements
0.734 0.583 0.697 0.634
By labelled difficulty clear hard ambiguous
0.766 0.687 0.522

Accuracy on minimal pairs (two items that differ in one detail and have different answers; both must be right): 0.423 on test.

Selective accuracy, keeping only the items whose top probability is highest:

Items kept top 30% top 50% top 70% all
test 0.917 0.838 0.774 0.686
challenge 0.933 0.853 0.795 0.693

Training

Data

  1. Stage 1: JNLI (CC BY-SA 4.0) converted to the decision format: each premise/hypothesis pair becomes a support/contradict/neutral choice question or a yes/no question.
  2. Stage 2: 49,668 synthetic items (state, question, answer) generated with DeepSeek V4.1 Flash, whose terms permit using outputs for training. Source documents are fictional business documents written by the generator, Japanese government FAQs (JaGovFaqs-22k, CC BY 4.0) and Japanese Wikipedia paragraphs (CC BY-SA 4.0). Questions come from 29 fixed templates (grounding, comparison, RAG, writing, routing) and from open questions the generator wrote itself; about 17,000 items form minimal pairs. The evaluation documents were never used for training. Five question templates used in the evaluation set were held out from training.

Procedure

Stage Setting
1 JNLI converted to decision format, 19,816 items, 3 epochs, 256 tokens, batch 32, lr 2e-5 (head 5e-4), full fine-tuning
2 49,668 synthetic items, 1 epoch, 1,024 tokens, batch 4 x 4 accumulation, lr 2e-5 (head 5e-4), full fine-tuning; 30% of noul items shown without criteria and 50% of choice items with shuffled options

The loss is cross-entropy over the options of each question, normalised per question type within a batch. Temperatures for calibration were fitted afterwards on the 600-item calibration split.

Results vary between training runs. Three stage-2 runs with different seeds from the same stage-1 checkpoint spanned 1.1 points of accuracy on the full 1,200-item test set; re-running stage 1 as well, final accuracy spanned 0.62 to 0.71 across five runs, and the runs that ended low had a noticeably higher validation NLL after stage 1 (0.34 against 0.24 to 0.29). If you retrain, train stage 1 more than once and keep the checkpoint with the lowest validation NLL. The published weights come from a run that ended at the top of that range.

Compute

Apple M4 Max (36 GB). Stage 1 about 35 minutes, stage 2 about 1.5 hours.

Limitations

  • Accuracy is 0.686 on test; the model is wrong on roughly one item in 3. Use the probability to route uncertain items rather than acting on every answer.
  • It is weaker on question templates it has not seen in training, on items our judges found ambiguous (0.522 on test), and on relative dates, numeric conditions and scope-limiting words such as 「原則として」.
  • The evaluation inputs are at most about 750 tokens (median about 220). Behaviour on longer documents has not been measured.
  • Labels in the evaluation set are LLM majority votes. Where the judges disagreed, the model's accuracy is much lower and the labels themselves are uncertain.
  • Each model was trained once (one seed). See the variance note under Training.
  • Japanese only. Other languages, images and knowledge-based fact checking are out of scope.

Bias, risks and ethical considerations

The training data is synthetic text about fictional organisations, plus public NLI and FAQ data; it does not cover every domain, register or document type, and the model may be systematically less accurate on text unlike its training data. Probabilities are calibrated on one evaluation set and should be re-checked on your own data before thresholds are used operationally. Do not use the output as the sole basis for decisions that affect people.

Changelog

  • v0.2 (2026-10-06): retrained with input augmentation. Noul questions without criteria now work (accuracy on test noul items without criteria: 0.551 → 0.757) and choice answers no longer depend on option order (accuracy with reversed options: 0.584 → 0.623). Overall accuracy is unchanged within run-to-run variance (v0.1 was 0.704 on test, 0.697 on challenge).
  • v0.1 (2026-10-05): first release.

License

Apache License 2.0. The base model sbintuitions/modernbert-ja-310m is distributed under the MIT license; its notice is included in NOTICE.

Citation

@misc{uratori2026,
  title  = {uratori: Japanese decision models for grounding checks, document comparison and RAG judgments},
  author = {tokimoa},
  year   = {2026},
  url    = {https://github.com/tokimoa/uratori}
}

Contact

Open an issue at github.com/tokimoa/uratori.

日本語

uratori-ja-310m は、日本語の文章と質問を受け取り、文章を生成せずに答えの確率分布を 1 回の forward で返すモデルです。質問の型は、真偽(noul)、選択(choice、2〜8 択)、段階評価(score、2〜10 段階)の 3 つです。与えた根拠から主張や回答が支持されるかの検証、2 つの文書の比較(矛盾、意味の変更、追加された主張)、RAG の判断(検索結果の関連性と十分性、回答の忠実性)を対象に学習しています。入出力の形は TypeSafe の /v1/systemone API と同じにしてあります。TypeSafe の Jev に着想を得た独立の実装で、TypeSafe とは関係がありません。

精度の高い版に uratori-ja-2b と uratori-ja-4b があります。CPU で実用的に動くのはこの 310m だけです。

使い方は上のコードのとおりです(trust_remote_code=True が必要です)。state は文字列か、名前つきの文章の dict で、dict のキーは質問文から `キー名` で参照できます。1 回の呼び出しに質問をいくつでも入れられ、質問どうしは独立に判定されます。確率は温度で校正した値です(calibrate=False で校正前の値)。入力が 1,024 トークンを超えると、切り詰めずにエラーになります。noul の criteria は省略できますが、はいといいえの条件を書いたほうが安定します。CPU で動きます。Apple M4 Max の CPU で 1 問 80 ms 前後、GPU(MPS、batch 4)で 17 ms です。

評価は uratori-ja-eval の test(802 問)と challenge(300 問)で、公開した重みそのもので測りました。正解は人が付けたものではなく、下書きの想定と 3 つの LLM の判定の多数決です。test の Accuracy は 0.686、challenge は 0.693。本家の Jev 1.13.0 は同じ問題で 0.903 と 0.900 です。

制限。判定は入力に与えた文章だけに基づき、モデルの知識で事実かどうかを確かめる用途には使えません。単独で結論を出す精度ではなく、確信度の低い件を LLM や人に回す前段として使うことを想定しています。学習で見ていない種類の質問、判定者が割れるような曖昧な問題、相対的な日付や数値の条件、「原則として」のような範囲を限定する語には弱いです。約 750 トークンを超える入力での挙動は測っていません。日本語専用です。

Configuration

Architecture
UratoriModel
Model type
uratori

Identity and Version

Repository
tokimoa/uratori-ja-310m
Publisher
Tokimoa
Task
Text classification
Modality
Text
Library
transformers
Parameters
315M parameters
Languages
ja
Revision
ae078c49ae11e9d8b62bccf568e01bc776408d84
First published
2026-10-05
Last updated
2026-10-06

Files and Weights

10 files, 1.3 GB in total. The weights are 1 file totalling 1.3 GB in safetensors.

Weights1 file · 1.3 GB
Configuration3 files · 11.8 KB
Tokenizer2 files · 6.7 MB
Documentation3 files · 28.3 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights1.3 GB 7ea6fc599f49
config.jsonConfiguration3.0 KB —
configuration_uratori.pyConfiguration657 B —
modeling_uratori.pyConfiguration8.2 KB —
LICENSEDocumentation11.4 KB —
NOTICEDocumentation1.3 KB —
README.mdDocumentation15.6 KB —
.gitattributesRepository1.5 KB —
tokenizer.jsonTokenizer6.7 MB —
tokenizer_config.jsonTokenizer597 B —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
1.3 GB
Download from Tokimoa

Released by Tokimoa through its official repository on Hugging Face. Read the license.

Built From

  • Derived from sbintuitions/modernbert-ja-310m
  • Trained on (disclosed) tokimoa/uratori-ja-eval

Memory Requirements

PrecisionWeights in memory
As published1.3 GB
16-bit0.6 GB
8-bit0.3 GB
4-bit0.2 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About uratori-ja-310m

How much GPU memory does uratori-ja-310m need?

About 0.8 GB at 16-bit and 0.2 GB at 4-bit: the weights (315M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run uratori-ja-310m on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use uratori-ja-310m commercially?

Yes. uratori-ja-310m is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Text classification

laya-multilingual

Convai Innovations

Non-autoregressive System 1 decision model covering 100+ languages. Give it a state (text, email, ticket, or JSON) and typed questions; it returns typed answers with probabilities in a single forward pass. No text generation, so nothing to parse and nothing to hallucinate. Part of the Laya family — use this checkpoint for anything that is not English. Since laya 0.3.13 the default Router() keeps both english and this checkpoint resident, so a mixed workload no longer swaps checkpoints on every language change. For a server, load them up front so even the first request of each language is just a forward pass: router.attach("multilingual", agent) registers an Agent you already built, so a…

Open weights apache-2.0 322M parameters transformers

Model · Text classification

laya-coreai

Andrey Babikov

Laya typed decisions on Apple Silicon, running on the Core AI runtime — the successor to Core ML. This is a.aimodel asset exported from via Apple's coreai-torch bridge. It outputs choice / score / noul probabilities (and RL action logits) with zero generated tokens and no PyTorch, Core ML, Transformers, or cloud API at inference time. macOS 27+ (Core AI runtime), Python 3.10+. Validated on M3 Max / macOS 27.2. Validate the download end-to-end (all three specializations, timing, contract checks): Snake demo with the model (terminal game, reuses the laya-coreml UI + safety shield; automatically uses the B3 asset when present for ~2x game throughput): ~3× faster per pass than the fastest Core…

Open weights apache-2.0 322M parameters coreai

Model · Text classification

laya-pt-es-typed

Telepatia

This checkpoint fine-tunes convaiinnovations/laya-multilingual for native choice, score, and noul decisions in Portuguese and Spanish. It keeps the original 322M-parameter mmBERT architecture. It adds no inference component and does not generate text. It returns typed answers and probabilities in one forward pass. This is a text model. Inference takes a textual state plus typed questions. The second training stage used text decisions derived from public speech corpora, but this checkpoint does not accept audio by itself. The separate audio projector is not included. The official Laya SDK defines these primitives as follows: - choice: selects one key from a runtime-defined criteria object.…

Open weights apache-2.0 322M parameters laya

Model · Text classification

laya-multilingual

Scott Lamkin

Non-autoregressive System 1 decision model covering 100+ languages. Give it a state (text, email, ticket, or JSON) and typed questions; it returns typed answers with probabilities in a single forward pass. No text generation, so nothing to parse and nothing to hallucinate. Part of the Laya family — use this checkpoint for anything that is not English. The default Router() keeps both english and this checkpoint resident, so a mixed workload no longer swaps checkpoints on every language change. For a server, load them up front so even the first request of each language is just a forward pass: router.attach("multilingual", agent) registers an Agent you already built, so a process that loaded…

Open weights apache-2.0 322M parameters transformers

Classifies GitHub issues written in any language as bug, feature, question or docs. A fine-tune of Laya multilingual (mmBERT-base) used by the laya-triage GitHub Action for non-English issues, next to the English model laya-triage-en. The same 500 NLBSE'23 validation issues, machine-translated with NLLB-200 into 13 languages. Accuracy (±3 points per language): laya-triage and Jev are within noise of each other across languages; both are far ahead of the untuned base. Translations can flatter a model trained on translations, so we also checked real issues: on 367 non-English issues opened in 2026 (never seen, written by people, not translated) accuracy went from 47.1% to 65.7%. Use it…

Open weights apache-2.0 322M parameters

Model · Text classification

rex

NguyenThanhDat

System 1 calibrated decision model for zero-latency threat triage

Open weights apache-2.0 322M parameters