This model is a fine-tuned version of microsoft/deberta-v3-large on an unknown dataset. It achieves the following results on the evaluation set: The following hyperparameters were used during training: - learningrate: 3e-05 - trainbatchsize: 16 - evalbatchsize: 16 - gradientaccumulationsteps: 2 - totaltrainbatchsize: 32 - lrschedulertype: cosine - lrschedulerwarmupsteps: 0.1 - numepochs: 1000 - Transformers 5.12.1 - Pytorch 2.11.0+cu128 - Datasets 5.0.1 - Tokenizers 0.22.2
Open-weight model · Text classification
DEBATE-kor-large
by Jong Rock Jeong jongrock17/DEBATE-kor-large
DEBATE-kor-large is an open-weight model for text classification from Jong Rock Jeong. It has 435M parameters and a 512-token context. At 16-bit it needs about 1 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.
DEBATE-kor-large is a Korean-adapted Political DEBATE model for binary natural language inference (NLI) on political text.
Runs On
What it takes to serve DEBATE-kor-large (435M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 0.9 GB | 1.0 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 0.4 GB | 0.5 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 0.2 GB | 0.3 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.
DEBATE-kor-large on every accelerator the SAVRN Index prices, at every precision
Model Card
DEBATE-kor-large is a Korean-adapted Political DEBATE model for binary natural language inference (NLI) on political text. The model is initialized from mlburnham/PoliticalDEBATEDeBERTalargev1.1, the original DeBERTa-based Political DEBATE checkpoint, and subsequently fine-tuned on jongrock17/PolNLI-kor, a Korean translation and adaptation of PolNLI. Political DEBATE DeBERTa-large → PolNLI-kor → DEBATE-kor-large Unlike the PolNLI-kor-RoBERTa model family, which starts from Korean-pretrained KLUE-RoBERTa encoders, DEBATE-kor directly adapts the original Political DEBATE checkpoint to Korean political NLI. DEBATE-kor formulates NLI as a binary classification problem. notentailment combines…
Excerpt from the card by Jong Rock Jeong.
Configuration
- Architecture
- DebertaV2ForSequenceClassification
- Context length (tokens)
- 512
- Layers
- 24
- Hidden size
- 1,024
- Feed-forward size
- 4,096
- Attention heads
- 16
- Vocabulary size
- 128,100
- Model type
- deberta-v2
Identity and Version
- Repository
- jongrock17/DEBATE-kor-large
- Publisher
- Jong Rock Jeong
- Task
- Text classification
- Modality
- Text
- Library
- transformers
- Parameters
- 435M parameters
- Languages
- ko
- Revision
- 49148c3cc68dc484f5f5b81c568cc500a0173fa9
- First published
- 2026-09-27
- Last updated
- 2026-09-27
Files and Weights
41 files, 12.2 GB in total. The weights are 14 files totalling 12.2 GB in bin, pt, pth, safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| checkpoint-32118/model.safetensors | Weights | 1.7 GB | 0ccd225f7c3d |
| checkpoint-32118/optimizer.pt | Weights | 3.5 GB | 4fedd4f99582 |
| checkpoint-32118/rng_state.pth | Weights | 14.6 KB | a3dc290565b5 |
| checkpoint-32118/scaler.pt | Weights | 1.4 KB | d5e6141d1bae |
| checkpoint-32118/scheduler.pt | Weights | 1.5 KB | ce729b13bed9 |
| checkpoint-32118/training_args.bin | Weights | 5.3 KB | 39a1a9695ee3 |
| checkpoint-5000/model.safetensors | Weights | 1.7 GB | 2c9c86edac54 |
| checkpoint-5000/optimizer.pt | Weights | 3.5 GB | 37191551849a |
| checkpoint-5000/rng_state.pth | Weights | 14.6 KB | c363ba9999fc |
| checkpoint-5000/scaler.pt | Weights | 1.4 KB | 272e402d1b17 |
| checkpoint-5000/scheduler.pt | Weights | 1.5 KB | a519864afda9 |
| checkpoint-5000/training_args.bin | Weights | 5.3 KB | 39a1a9695ee3 |
| model.safetensors | Weights | 1.7 GB | 2c9c86edac54 |
| training_args.bin | Weights | 5.3 KB | 39a1a9695ee3 |
| checkpoint-32118/config.json | Configuration | 1.1 KB | — |
| checkpoint-32118/trainer_state.json | Configuration | 147.3 KB | — |
| checkpoint-5000/config.json | Configuration | 1.1 KB | — |
| checkpoint-5000/trainer_state.json | Configuration | 23.5 KB | — |
| config.json | Configuration | 1.1 KB | — |
| metrics.json | Configuration | 1.4 KB | — |
| README.md | Documentation | 8.1 KB | — |
| log_history.csv | Other | 92.8 KB | — |
| runs/Jul23_09-41-29_883cefcad1d3/events.out.tfevents.1784799693.883cefcad1d3.5503.2 | Other | 174.2 KB | 11254f5e4d36 |
| runs/Jul23_09-41-29_883cefcad1d3/events.out.tfevents.1784808236.883cefcad1d3.5503.3 | Other | 2.1 KB | 8e8720abe7ea |
| test_confusion_matrix.png | Other | 64.1 KB | — |
| test_labels.npy | Other | 123.1 KB | abff1f69fcd8 |
| test_logits.npy | Other | 123.1 KB | 9b9f8ba12868 |
| test_preds.csv | Other | 61.5 KB | — |
| test_preds.npy | Other | 123.1 KB | b5e814053423 |
| val_confusion_matrix.png | Other | 65.6 KB | — |
| val_labels.npy | Other | 120.4 KB | 008be99f5f5f |
| val_logits.npy | Other | 120.4 KB | 245fee6a2011 |
| val_preds.csv | Other | 60.2 KB | — |
| val_preds.npy | Other | 120.4 KB | 275549b60589 |
| .gitattributes | Repository | 1.5 KB | — |
| checkpoint-32118/tokenizer.json | Tokenizer | 8.3 MB | — |
| checkpoint-32118/tokenizer_config.json | Tokenizer | 695 B | — |
| checkpoint-5000/tokenizer.json | Tokenizer | 8.3 MB | — |
| checkpoint-5000/tokenizer_config.json | Tokenizer | 695 B | — |
| tokenizer.json | Tokenizer | 8.3 MB | — |
| tokenizer_config.json | Tokenizer | 695 B | — |
License and Download
- License
- Not stated by the source
- Access
- Open weights, no gate
- Download size
- 12.2 GB
Released by Jong Rock Jeong through its official repository on Hugging Face.
Built From
- Derived from mlburnham/Political_DEBATE_DeBERTa_large_v1.1
- Trained on (disclosed) jongrock17/PolNLI-kor
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 12.2 GB |
| 16-bit | 0.9 GB |
| 8-bit | 0.4 GB |
| 4-bit | 0.2 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About DEBATE-kor-large
How much GPU memory does DEBATE-kor-large need?
About 1 GB at 16-bit and 0.3 GB at 4-bit: the weights (435M parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run DEBATE-kor-large on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
What is DEBATE-kor-large's context length?
512 tokens, from the maximum position embeddings in its published configuration.
Similar Models
Opir-multitask-large is the English, highest-accuracy multi-task checkpoint in the Opir family: an encoder-based GLiClass guardrail model for real-time LLM safety filtering. It supports binary safe/unsafe classification, toxicity detection, jailbreak and prompt-injection detection, and zero-shot harmful-content categorization over a hierarchical safety taxonomy. This card is for knowledgator/opir-multitask-large. The model is used through GLiClass zero-shot classification: pass text plus the candidate labels you want scored. Use single-label mode for binary safe/unsafe decisions and multi-label mode for taxonomy, toxicity, jailbreak, or custom policy labels. Use multi-label mode when you…
This is an experimental, uncalibrated choice-ranking model. It is not an official Jev model, a validated general-purpose reasoner, or an automatic decision-maker. The model ranks 2–16 user-supplied candidate texts for a natural-language context and question and returns all candidate probabilities through the ERABI code. Decisions should be reviewed by a person. - Practical V1 data consists of original synthetic Japanese, English, and Simplified Chinese examples in six task families, generated and answer-blind rejudged with DeepSeek V4.1 Flash. - The Exam-QA source was filtered and transformed with the same DeepSeek model. Symbolic answer labels were mapped to source choice text. Ambiguous…
Non-autoregressive System 1 decision model, fine-tuned on the typed-decisions workflows: agent-trace observability, customer service, invoice processing and security incidents. Part of the Laya family. 400 test cases, 2,000 decisions, measured on the official test split. +3.9 points over Jev's published 0.727, above the 0.735 teacher ceiling, with 2.4x better Brier and 1.6x better score MAE. Jev figures are third-party published, not measured here — there is no TypeSafe API access in this project, and sample sizes and prompts differ. Treat the comparison as indicative. Router will not select this checkpoint automatically unless you construct it with autotaskdetection=True — it is…
Multilingual, non-autoregressive System 1 decision model. Give it a state (text, email, ticket, or JSON) and typed questions; it returns typed answers with mathematically calibrated probabilities in a single forward pass (~33 ms) across 100+ languages. Trained with reinforcement learning against strictly proper scoring rules (RLCD), so reporting honest probabilities is the only way to maximise reward. It never generates text, so there is nothing to parse and nothing to hallucinate. This repo holds all three checkpoints and is the hub for the family. The English checkpoint is at the repo root; the other two are bundled subfolders, and only the one you request is downloaded: pip install -U…
A fine-tuned Laya model (Convai Innovations, 421M, ModernBERT-large backbone) that answers one typed noul question: does a value correctly match its key label? (e.g. firstname = John → true, firstname = 1992 → false). Fine-tuned on the IBM KVP-10K dataset via the pre-parsed community mirror (OCR pre-extracted; no OCR step needed). This is the v2 (remediated) release: document-level 80/10/10 data partitioning, mixed-class frozen test set, calib-split temperature fitting, and a frozen category-stratified release gate. It supersedes the v1 release pre-remediation, historical artifact). Laya is a non-autoregressive decision model: you give it a state (text/dict) plus typed questions (choice…