The custom human-supervised typed-judgment corpus used for OpenJudgement-4B-Preview, developed by Kitani. It adapts documented 2025–2026 source releases into Noul, Choice, and Score decisions.
Dataset Card
The custom human-supervised typed-judgment corpus used for OpenJudgement-4B-Preview, developed by Kitani. It adapts documented 2025–2026 source releases into Noul, Choice, and Score decisions. No Mix-v3 rows or local synthetic generators were reused in this corpus. This is an experimental training dataset, not evidence of Jev parity. Coverage is incomplete, source labels can be disputed, and the model remains unfinished. The source-specific terms below continue to apply. The selected model checkpoint was trained for 700 optimizer steps, seeing 44,741 decisions and 49,046,168 input tokens. Those are actual training exposures; the full corpus and split sizes below describe the available data…
Excerpt from the card by Kitani, licensed other.
Structure
default 111,049 rows
| Split | Rows | Size |
|---|---|---|
| train | 87,621 | 400.8 MB |
| validation | 11,278 | 55.2 MB |
| calibration | 2,968 | 13.4 MB |
| test | 7,446 | 17.2 MB |
| challenge | 1,736 | 2.1 MB |
Details
- Repository
- kitaniai/OpenJudgment-4B-Preview
- Publisher
- Kitani
- Task category
- Text classification
- Tags
- Not stated by the source
- Size category
- Not stated by the source
- Languages
- en, uk
- Revision
- 7a1d5b4692cd9a44183553ba92762d045664a4f3
- Last updated
- 2026-09-24
Files
22 files, 149.7 MB in total.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| calibration.parquet | Data | 4.1 MB | 474596992c10 |
| challenge.parquet | Data | 381.4 KB | 08a5d3c18ce0 |
| checksums.json | Data | 663 B | — |
| manifest.json | Data | 6.1 KB | — |
| reports/audit.json | Data | 726 B | — |
| reports/quality_review.json | Data | 2.2 KB | — |
| reports/tokenization.json | Data | 1.2 KB | — |
| source_manifest.json | Data | 3.2 KB | — |
| test.parquet | Data | 5.5 MB | 17c50c5a1a6b |
| train.parquet | Data | 123.0 MB | dfae7d238e60 |
| validation.parquet | Data | 16.7 MB | 2c317ab02e5f |
| README.md | Documentation | 7.4 KB | — |
| code/test_v4_helpsteer.py | Other | 3.3 KB | — |
| code/test_v4_intent.py | Other | 1.7 KB | — |
| code/v4_audit.py | Other | 5.4 KB | — |
| code/v4_build.py | Other | 12.3 KB | — |
| code/v4_fetch.py | Other | 1.9 KB | — |
| code/v4_helpsteer.py | Other | 8.4 KB | — |
| code/v4_intent.py | Other | 4.0 KB | — |
| code/v4_publish.py | Other | 9.4 KB | — |
| code/v4_token_stats.py | Other | 3.0 KB | — |
| .gitattributes | Repository | 2.5 KB | — |
License and Download
- License
- other
- Access
- No access gate
Released by Kitani through its official repository on Hugging Face. Read the license.
Models Trained on This Dataset
- Trained on (disclosed)OpenJudgement-4B-Preview