SAVRN
Search Contact SAVRN

Open-weight model · Question answering

koelectra-small-v2-distilled-korquad-384

by Jangwon Park monologg/koelectra-small-v2-distilled-korquad-384

Parameters14M
Context512
Weights207.8 MB
License
AccessOpen weights
Monthly Downloads163.1k

Runs On

What it takes to serve koelectra-small-v2-distilled-korquad-384 (14M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 0.0 GB 0.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.0 GB 0.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.0 GB 0.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

Model Card

The publisher has not written a card for this model.

Configuration

Architecture
ElectraForQuestionAnswering
Context length (tokens)
512
Layers
12
Hidden size
256
Feed-forward size
1,024
Attention heads
4
Vocabulary size
32,200
Model type
electra

Identity and Version

Repository
monologg/koelectra-small-v2-distilled-korquad-384
Publisher
Jangwon Park
Task
Question answering
Modality
Text
Library
transformers
Parameters
14M parameters
Languages
Not stated by the source
Revision
256efd8763caf2d3936e141107f113e2ab51f653
First published
2022-03-02
Last updated
2023-06-12

Files and Weights

10 files, 208.0 MB in total. The weights are 5 files totalling 207.8 MB in bin, safetensors, tflite.

Weights5 files · 207.8 MB
Configuration2 files · 584 B
Tokenizer2 files · 255.2 KB
Repository1 file · 399 B
Every file
FileTypeSizeSHA-256
model.safetensorsWeights54.8 MB 484d79698351
model.tfliteWeights55.8 MB 054e017888a8
model_8bits.tfliteWeights14.5 MB 38c3c59a4c77
model_fp16.tfliteWeights27.8 MB 5c64e1cd2ebd
pytorch_model.binWeights54.8 MB 043efda1484f
config.jsonConfiguration472 B
special_tokens_map.jsonConfiguration112 B
.gitattributesRepository399 B
tokenizer_config.jsonTokenizer49 B
vocab.txtTokenizer255.2 KB

License and Download

License
Not stated by the source
Access
Open weights, no gate
Download size
207.8 MB
Download from Jangwon Park

Released by Jangwon Park through its official repository on Hugging Face.

Memory Requirements

PrecisionWeights in memory
As published207.8 MB
16-bit0.0 GB
8-bit0.0 GB
4-bit0.0 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About koelectra-small-v2-distilled-korquad-384

How much GPU memory does koelectra-small-v2-distilled-korquad-384 need?

About 0 GB at 16-bit and 0 GB at 4-bit: the weights (14M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run koelectra-small-v2-distilled-korquad-384 on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

What is koelectra-small-v2-distilled-korquad-384's context length?

512 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Question answering

mobilebert-uncased-squad-v2

Qingqing Cao

MobileBERT is a thin version of BERTLARGE, while equipped with bottleneck structures and a carefully designed balance between self-attentions and feed-forward networks. This model was fine-tuned from the HuggingFace checkpoint google/mobilebert-uncased on SQuAD2.0. CPU: Intel(R) Core(TM) i7-6800K CPU @ 3.40GHz Memory: 32 GiB GPUs: 2 GeForce GTX 1070, each with 8GiB memory GPU driver: 418.87.01, CUDA: 10.1 It took about 3.5 hours to finish. Note that the above results didn't involve any hyperparameter search.

Open weights mit 25M parameters 512 tokens transformers

Model · Question answering

mobilebert-uncased-squad-v1

Qingqing Cao

MobileBERT is a thin version of BERTLARGE, while equipped with bottleneck structures and a carefully designed balance between self-attentions and feed-forward networks. This model was fine-tuned from the HuggingFace checkpoint google/mobilebert-uncased on SQuAD1.1. CPU: Intel(R) Core(TM) i7-6800K CPU @ 3.40GHz Memory: 32 GiB GPUs: 2 GeForce GTX 1070, each with 8GiB memory GPU driver: 418.87.01, CUDA: 10.1 It took about 3 hours to finish. Note that the above results didn't involve any hyperparameter search.

Open weights mit 25M parameters 512 tokens transformers

Model · Question answering

splinter-base

Tel Aviv University

Splinter-base is the pretrained model discussed in the paper Few-Shot Question Answering by Pretraining Span Selection (at ACL 2021). Its original repository can be found here. The model is case-sensitive. Note: This model doesn't contain the pretrained weights for the QASS layer (see paper for details), and therefore the QASS layer is randomly initialized upon loading it. For the model with those weights, see tau/splinter-base-qass. Splinter is a model that is pretrained in a self-supervised fashion for few-shot question answering. This means it was pretrained on the raw texts only, with no humans labelling them in any way (which is why it can use lots of publicly available data) with an…

Open weights apache-2.0 512 tokens transformers

This model can be used for the task of question answering. The model should not be used to intentionally create hostile or alienating environments for people. Significant research has explored bias and fairness issues with language models (see, e.g., Sheng et al. (2021) and Bender et al. (2021)). Predictions generated by the model may include disturbing and harmful stereotypes across protected classes; identity characteristics; and sensitive, social, and occupational groups. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. The model creators note in the associated paper: The model creators note in the associated paper: The model…

Open weights 512 tokens transformers