SAVRN
Search Contact SAVRN

Open-weight model · Text classification

DEBATE-kor-base

by Jong Rock Jeong jongrock17/DEBATE-kor-base

DEBATE-kor-base is an open-weight model for text classification from Jong Rock Jeong. It has 184M parameters and a 512-token context. At 16-bit it needs about 0.4 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

DEBATE-kor-base is a Korean-adapted Political DEBATE model for binary natural language inference (NLI) on political text.

Parameters184M
Context512
Weights5.2 GB
License—
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve DEBATE-kor-base (184M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 0.4 GB 0.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.2 GB 0.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.1 GB 0.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.

DEBATE-kor-base on every accelerator the SAVRN Index prices, at every precision

Model Card

DEBATE-kor-base is a Korean-adapted Political DEBATE model for binary natural language inference (NLI) on political text. The model is initialized from mlburnham/PoliticalDEBATEDeBERTabasev1.1, the original DeBERTa-based Political DEBATE checkpoint, and subsequently fine-tuned on jongrock17/PolNLI-kor, a Korean translation and adaptation of PolNLI. Political DEBATE DeBERTa-base → PolNLI-kor → DEBATE-kor-base Unlike the PolNLI-kor-RoBERTa model family, which starts from Korean-pretrained KLUE-RoBERTa encoders, DEBATE-kor directly adapts the original Political DEBATE checkpoint to Korean political NLI. DEBATE-kor formulates NLI as a binary classification problem. notentailment combines…

Excerpt from the card by Jong Rock Jeong.

Configuration

Architecture
DebertaV2ForSequenceClassification
Context length (tokens)
512
Layers
12
Hidden size
768
Feed-forward size
3,072
Attention heads
12
Vocabulary size
128,100
Model type
deberta-v2

Identity and Version

Repository
jongrock17/DEBATE-kor-base
Publisher
Jong Rock Jeong
Task
Text classification
Modality
Text
Library
transformers
Parameters
184M parameters
Languages
ko
Revision
554f705473156e4fd229f0c85494fa1793817b2a
First published
2026-09-27
Last updated
2026-09-27

Files and Weights

41 files, 5.2 GB in total. The weights are 14 files totalling 5.2 GB in bin, pt, pth, safetensors.

Weights14 files · 5.2 GB
Configuration6 files · 288.6 KB
Tokenizer6 files · 25.0 MB
Documentation1 file · 8.5 KB
Other13 files · 1.3 MB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
checkpoint-30000/model.safetensorsWeights737.7 MB 0ab7283b97e1
checkpoint-30000/optimizer.ptWeights1.5 GB 614f50042068
checkpoint-30000/rng_state.pthWeights14.7 KB 95bbd2f42e2a
checkpoint-30000/scaler.ptWeights1.4 KB 5ba887720622
checkpoint-30000/scheduler.ptWeights1.5 KB ff2877013da3
checkpoint-30000/training_args.binWeights5.3 KB 0f3128f1ad2e
checkpoint-32118/model.safetensorsWeights737.7 MB f90fce3e4b2c
checkpoint-32118/optimizer.ptWeights1.5 GB 34a9d38b27e8
checkpoint-32118/rng_state.pthWeights14.7 KB 59ce39be35f8
checkpoint-32118/scaler.ptWeights1.4 KB 31ba55ba6c4f
checkpoint-32118/scheduler.ptWeights1.5 KB ce729b13bed9
checkpoint-32118/training_args.binWeights5.3 KB 0f3128f1ad2e
model.safetensorsWeights737.7 MB 0ab7283b97e1
training_args.binWeights5.3 KB 0f3128f1ad2e
checkpoint-30000/config.jsonConfiguration1.1 KB —
checkpoint-30000/trainer_state.jsonConfiguration136.9 KB —
checkpoint-32118/config.jsonConfiguration1.1 KB —
checkpoint-32118/trainer_state.jsonConfiguration147.1 KB —
config.jsonConfiguration1.1 KB —
metrics.jsonConfiguration1.4 KB —
README.mdDocumentation8.5 KB —
log_history.csvOther92.7 KB —
runs/Jul23_08-29-16_883cefcad1d3/events.out.tfevents.1784795358.883cefcad1d3.5503.0Other174.2 KB e7fbdb3454c0
runs/Jul23_08-29-16_883cefcad1d3/events.out.tfevents.1784799523.883cefcad1d3.5503.1Other2.1 KB 497432a97822
test_confusion_matrix.pngOther64.1 KB —
test_labels.npyOther123.1 KB abff1f69fcd8
test_logits.npyOther123.1 KB 72621211a2ca
test_preds.csvOther61.5 KB —
test_preds.npyOther123.1 KB 0b5e8ea17eec
val_confusion_matrix.pngOther66.0 KB —
val_labels.npyOther120.4 KB 008be99f5f5f
val_logits.npyOther120.4 KB bcd1fb8feacf
val_preds.csvOther60.2 KB —
val_preds.npyOther120.4 KB b54c75553847
.gitattributesRepository1.5 KB —
checkpoint-30000/tokenizer.jsonTokenizer8.3 MB —
checkpoint-30000/tokenizer_config.jsonTokenizer695 B —
checkpoint-32118/tokenizer.jsonTokenizer8.3 MB —
checkpoint-32118/tokenizer_config.jsonTokenizer695 B —
tokenizer.jsonTokenizer8.3 MB —
tokenizer_config.jsonTokenizer695 B —

License and Download

License
Not stated by the source
Access
Open weights, no gate
Download size
5.2 GB
Download from Jong Rock Jeong

Released by Jong Rock Jeong through its official repository on Hugging Face.

Built From

  • Derived from mlburnham/Political_DEBATE_DeBERTa_base_v1.1
  • Trained on (disclosed) jongrock17/PolNLI-kor

Memory Requirements

PrecisionWeights in memory
As published5.2 GB
16-bit0.4 GB
8-bit0.2 GB
4-bit0.1 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About DEBATE-kor-base

How much GPU memory does DEBATE-kor-base need?

About 0.4 GB at 16-bit and 0.1 GB at 4-bit: the weights (184M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run DEBATE-kor-base on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

What is DEBATE-kor-base's context length?

512 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text classification

deberta-v3-base-prompt-injection-v2

Protect AI

This model is a fine-tuned version of microsoft/deberta-v3-base specifically developed to detect and classify prompt injection attacks which can manipulate language models into producing unintended outputs. Prompt injection attacks manipulate language models by inserting or altering prompts to trigger harmful or unintended responses. The deberta-v3-base-prompt-injection-v2 model is designed to enhance security in language model applications by detecting these malicious interventions. This model classifies inputs into benign (0) and injection-detected (1). deberta-v3-base-prompt-injection-v2 is highly accurate in identifying prompt injections in English. It does not detect jailbreak attacks…

Open weights apache-2.0 184M parameters 512 tokens transformers

This model was trained on 1.279.665 hypothesis-premise pairs from 8 NLI datasets: MultiNLI, Fever-NLI, LingNLI and DocNLI (which includes ANLI, QNLI, DUC, CNN/DailyMail, Curation). It is the only model in the model hub trained on 8 NLI datasets, including DocNLI with very long texts to learn long range reasoning. Note that the model was trained on binary NLI to predict either "entailment" or "not-entailment". The DocNLI merges the classes "neural" and "contradiction" into "not-entailment" to enable the inclusion of the DocNLI dataset. The base model is DeBERTa-v3-base from Microsoft. The v3 variant of DeBERTa substantially outperforms previous versions of the model by including a different…

Open weights mit 184M parameters 512 tokens transformers

Model · Text classification

ai-agent-sme-sentiment-mbert

Atsushi Hatakeyama

This is a three-class Japanese/English sentiment classifier for short hospitality reviews, fine-tuned from google-bert/bert-base-multilingual-cased. It categorizes each review as negative, neutral, or positive. The model predicts one overall label for a short Japanese or English hospitality It is intended for research, prototyping, and human-reviewed analytics. It must not be used as the sole basis for consequential business, employment, moderation, or customer-service decisions. Applications using this model should allow operators to review and correct its predictions. The model was fine-tuned on 480 synthetic bilingual reviews: 240 Japanese and 240 English, balanced across the three…

Open weights apache-2.0 178M parameters 512 tokens transformers

Model · Text classification

bert-base-multilingual-uncased-sentiment

NLP Town

Visit the NLP Town website for an updated version of this model, with a 40% error reduction on product reviews. This is a bert-base-multilingual-uncased model finetuned for sentiment analysis on product reviews in six languages: English, Dutch, German, French, Spanish, and Italian. It predicts the sentiment of the review as a number of stars (between 1 and 5). This model is intended for direct use as a sentiment analysis model for product reviews in any of the six languages above or for further finetuning on related sentiment analysis tasks. Here is the number of product reviews we used for finetuning the model: The fine-tuned model obtained the following accuracy on 5,000 held-out product…

Open weights mit 167M parameters 512 tokens transformers

Model · Text classification

pyrrho-v2-nano-g1

Yan Fitzner

Pyrrho is a CPU-runnable co-processor for retrieval-augmented generation. It reads a question before retrieval to suggest what evidence to seek, then reads the question with retrieved passages to assess whether those passages support an answer. A surrounding RAG runtime decides whether to answer, retrieve again, or surface a conflict. The broader project is described in the The evidence verdict is SUFFICIENT, DISPUTED, or INSUFFICIENT. These are predictions about the supplied passages. The model does not retrieve sources, generate answers or citations, check external facts, or prove that a corpus has been searched completely. The two passes use the same encoder with different input…

Open weights cc-by-nc-4.0 150M parameters 8,192 tokens transformers

Model · Text classification

Firebird-ModernBERT-512-RW

Noumenon, Inc.

Firebird-ModernBERT-512-RW is an experimental post-trained variant of It is a ~149M parameter ModernBERT binary classifier for distinguishing: - 0 — HUMAN The maximum sequence length is 512 tokens. This checkpoint was produced through reward-weighted classifier post-training. The original Firebird checkpoint was kept frozen as a reference model. Training examples were scored by the original classifier, difficult examples received larger loss weights, and the post-trained model was constrained against the frozen reference using a KL penalty. L = weightedcrossentropy + beta KL(reference || policy) Hard human examples received greater weighting than ordinary examples because one goal of the…

Open weights 150M parameters 8,192 tokens transformers