SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

HyperCLOVA-X-KETI-HAECHI-32B

by Korea Electronics Technology Institute NLP KETI-NLP/HyperCLOVA-X-KETI-HAECHI-32B

HyperCLOVA-X-KETI-HAECHI-32B is an open-weight model for image and text to text from Korea Electronics Technology Institute NLP, released under other. It has 33.3B parameters and a 131,072-token context. At 16-bit it needs about 80 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 42 downloads a month.

HyperCLOVA X KETI-HAECHI-32B is a multimodal model derived from It was developed for two primary purposes: improving Korean cultural-heritage understanding and Korean OCR, and improving tool calling and multi-step, stateful agent execution.

Parameters33.3B
Context131,072
Weights66.6 GB
Licenseother
AccessOpen weights
Monthly Downloads42

Runs On

What it takes to serve HyperCLOVA-X-KETI-HAECHI-32B (33.3B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 66.6 GB 80.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 33.3 GB 40.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 16.7 GB 20.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

HyperCLOVA-X-KETI-HAECHI-32B on every accelerator the SAVRN Index prices, at every precision

Model Card

HyperCLOVA X KETI-HAECHI-32B is a multimodal model derived from It was developed for two primary purposes: improving Korean cultural-heritage understanding and Korean OCR, and improving tool calling and multi-step, stateful agent execution. Alongside these goals, the model retains broad multimodal, language, and coding capabilities from the base model. heritage objects, answers questions grounded in heritage images, and reads Korean text from signs, scenes, rendered text, and public documents. - Tool calling and long-horizon task execution: selects and calls tools, carries information across multiple turns, tracks changing state, and works toward an end-to-end goal over several steps.…

Excerpt from the card by Korea Electronics Technology Institute NLP, licensed other.

Configuration

Architecture
HCXVisionV2ForCausalLM
Context length (tokens)
131,072
Layers
72
Hidden size
5,120
Feed-forward size
24,192
Attention heads
40
Key/value heads
8
Head dimension
128
Vocabulary size
128,256
RoPE base
50,000,000
Stored precision
float32
Model type
vlm

Identity and Version

Repository
KETI-NLP/HyperCLOVA-X-KETI-HAECHI-32B
Publisher
Korea Electronics Technology Institute NLP
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
33.3B parameters
Languages
ko, en
Revision
776bf4ac7207f47f171a00b81ac4f161880f156d
First published
2026-09-04
Last updated
2026-10-05

Files and Weights

104 files, 66.7 GB in total. The weights are 29 files totalling 66.6 GB in safetensors.

Weights29 files · 66.6 GB
Configuration22 files · 618.7 KB
Tokenizer4 files · 13.6 MB
Documentation6 files · 149.0 KB
Other42 files · 11.4 MB
Repository1 file · 3.9 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00029.safetensorsWeights1.4 GB 0566432c720c
model-00002-of-00029.safetensorsWeights2.3 GB 6681c20ce949
model-00003-of-00029.safetensorsWeights2.5 GB 640566f19947
model-00004-of-00029.safetensorsWeights2.4 GB 3d96da384f4b
model-00005-of-00029.safetensorsWeights2.4 GB 68aa920f489b
model-00006-of-00029.safetensorsWeights2.4 GB 1da578313d99
model-00007-of-00029.safetensorsWeights2.5 GB 171b09a6887d
model-00008-of-00029.safetensorsWeights2.4 GB 482d223a4b56
model-00009-of-00029.safetensorsWeights2.4 GB d0743249e446
model-00010-of-00029.safetensorsWeights2.4 GB 83ef3d4cc622
model-00011-of-00029.safetensorsWeights2.5 GB 92c6cbffb89f
model-00012-of-00029.safetensorsWeights2.4 GB 8a025f073600
model-00013-of-00029.safetensorsWeights2.4 GB f6f20b5021ad
model-00014-of-00029.safetensorsWeights2.4 GB 5c75ba07fc9c
model-00015-of-00029.safetensorsWeights2.5 GB 65fcf9661939
model-00016-of-00029.safetensorsWeights2.4 GB 74a4062c9941
model-00017-of-00029.safetensorsWeights2.4 GB 1ed8a11a44f7
model-00018-of-00029.safetensorsWeights2.4 GB 63c6e813841a
model-00019-of-00029.safetensorsWeights2.5 GB 154c6657bcac
model-00020-of-00029.safetensorsWeights2.4 GB b2227ec30092
model-00021-of-00029.safetensorsWeights2.4 GB fabf52ce5acb
model-00022-of-00029.safetensorsWeights2.4 GB 3d21d3ab7c25
model-00023-of-00029.safetensorsWeights2.5 GB b8e70a839f6e
model-00024-of-00029.safetensorsWeights2.4 GB fe03e7c56cf8
model-00025-of-00029.safetensorsWeights2.4 GB 98f78cec8d0f
model-00026-of-00029.safetensorsWeights2.4 GB eadb76b513f0
model-00027-of-00029.safetensorsWeights2.5 GB 2630df8036bb
model-00028-of-00029.safetensorsWeights1.7 GB d3a6cf87eb57
model-00029-of-00029.safetensorsWeights1.4 GB 4e810281880c
BF16_CONVERSION_PROVENANCE.jsonConfiguration582 B —
PACKAGE_MANIFEST.jsonConfiguration726 B —
added_tokens.jsonConfiguration8.3 KB —
config.jsonConfiguration6.4 KB —
configuration_hyperclovax.pyConfiguration12.2 KB —
configuration_vlm.pyConfiguration4.6 KB —
generation_config.jsonConfiguration164 B —
model.safetensors.index.jsonConfiguration102.0 KB —
modeling_hyperclovax.pyConfiguration86.4 KB —
modeling_vlm.pyConfiguration91.7 KB —
preprocessor_config.jsonConfiguration654 B —
processing_vlm.pyConfiguration36.2 KB —
processor_config.jsonConfiguration128 B —
reports/TRAINING_CONFIG.yamlConfiguration9.3 KB —
reproduction/convert_hf_safetensors_to_bf16.pyConfiguration5.7 KB —
reproduction/extract_large_model_paired_examples.pyConfiguration19.2 KB —
reproduction/generate_large_model_release_docs.pyConfiguration25.7 KB —
reproduction/generate_performance_chart.pyConfiguration4.1 KB —
reproduction/package_hf_release_v2.pyConfiguration6.4 KB —
samples/paired_examples.jsonConfiguration195.6 KB —
special_tokens_map.jsonConfiguration798 B —
video_preprocessor_config.jsonConfiguration1.8 KB —
LICENSEDocumentation16.9 KB —
NOTICEDocumentation422 B —
README.mdDocumentation17.4 KB —
reports/BENCHMARK_GUIDE.mdDocumentation4.6 KB —
reports/DETAILED_SCORE_REPORT.mdDocumentation12.4 KB —
reports/PAIRED_OUTPUT_EXAMPLES.mdDocumentation97.2 KB —
HyperCLOVA_X_32B_Think.pdfOther1.6 MB c5b1aca92886
SHA256SUMSOther10.3 KB —
chat_template.jinjaOther7.6 KB —
samples/images/01_mammoth_and_ocr.jpgOther183.8 KB 4db5b69e2cf8
samples/images/02_reverse_candidates.jpgOther268.7 KB cd9f799b8a2e
samples/images/03_cultural_focus.jpgOther140.0 KB 072527d21c25
samples/images/04_h400_hs100.jpgOther473.9 KB efb5c63cb739
samples/images/cultural_focus_caption.jpgOther279.5 KB 562f08978d92
samples/images/cultural_focus_identity.jpgOther264.5 KB e9a3e13625ea
samples/images/h400_direct.jpgOther325.0 KB f2e47757ec4e
samples/images/h400_hard.jpgOther264.5 KB a50ae1be0768
samples/images/h400_knowledge_image.jpgOther401.2 KB c2ce57715548
samples/images/haechi.pngOther1.6 MB 5763215253dc
samples/images/haechi_personalization.pngOther1.8 MB da6a719c4fd6
samples/images/hs100_train.jpgOther329.3 KB eaa821540d2e
samples/images/hs100_unseen.jpgOther393.7 KB bde6f39c1ac8
samples/images/key_performance_overview.pngOther163.9 KB c0eebb3e6a93
samples/images/mammoth_heritage_multi.jpgOther222.3 KB daf02ac937a0
samples/images/mammoth_heritage_reverse_A.jpgOther60.8 KB —
samples/images/mammoth_heritage_reverse_B.jpgOther44.8 KB —
samples/images/mammoth_heritage_reverse_C.jpgOther58.7 KB —
samples/images/mammoth_heritage_reverse_D.jpgOther61.0 KB —
samples/images/mammoth_heritage_simple.jpgOther80.8 KB —
samples/images/national_treasure.jpgOther72.8 KB —
samples/images/ocr_font.jpgOther17.0 KB —
samples/images/ocr_outdoor.jpgOther77.7 KB —
samples/images/ocr_public.jpgOther19.2 KB —
samples/qualitative/01_gyeongbokgung_gyeonghoeru.jpgOther205.5 KB a68937f9e1e0
samples/qualitative/02_geumdong_yeonga_buddha.jpgOther198.0 KB 83e8f789aa29
samples/qualitative/03_white_porcelain_grape_monkey_jar.jpgOther151.7 KB eea45f66c3c6
samples/qualitative/04_muryeong_crown_ornament.jpgOther221.0 KB 123a1e1d627f
samples/qualitative/05_ocr_font_waen.pngOther12.2 KB —
samples/qualitative/06_ocr_font_jeojeollo.pngOther32.6 KB —
samples/qualitative/07_ocr_font_ipda.pngOther34.7 KB —
samples/qualitative/08_ocr_font_chap.pngOther14.2 KB —
samples/qualitative/09_ocr_outdoor_yeonse_stmary.pngOther285.3 KB 5ebcbc372251
samples/qualitative/10_ocr_outdoor_hana_karaoke.pngOther297.2 KB 26563549b12d
samples/qualitative/11_ocr_outdoor_coin_karaoke.pngOther272.1 KB e13037146e6c
samples/qualitative/12_ocr_outdoor_ddungs.pngOther264.4 KB e8a69827c683
samples/qualitative/13_ocr_public_disaster.pngOther66.1 KB —
samples/qualitative/14_ocr_public_corp_number.pngOther74.9 KB —
samples/qualitative/15_ocr_public_gimhae_market.pngOther39.5 KB —
.gitattributesRepository3.9 KB —
merges.txtTokenizer1.4 MB —
tokenizer.jsonTokenizer9.7 MB —
tokenizer_config.jsonTokenizer48.3 KB —
vocab.jsonTokenizer2.4 MB —

License and Download

License
other
Access
Open weights, no gate
Download size
66.6 GB
Download from Korea Electronics Technology Institute NLP

Released by Korea Electronics Technology Institute NLP through its official repository on Hugging Face.

Built From

  • Derived from naver-hyperclovax/HyperCLOVAX-SEED-Think-32B

Memory Requirements

PrecisionWeights in memory
As published66.6 GB
16-bit66.6 GB
8-bit33.3 GB
4-bit16.7 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About HyperCLOVA-X-KETI-HAECHI-32B

How much GPU memory does HyperCLOVA-X-KETI-HAECHI-32B need?

About 80 GB at 16-bit and 20 GB at 4-bit: the weights (33.3B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run HyperCLOVA-X-KETI-HAECHI-32B on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

What license is HyperCLOVA-X-KETI-HAECHI-32B released under?

other, as its publisher declares it. Read the license text before commercial use.

What is HyperCLOVA-X-KETI-HAECHI-32B's context length?

131,072 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Weight-only quantization of built to a hard 22 GB budget with the lowest perplexity achievable inside it. This is not a uniform W4A16. Bit-width was allocated by measurement: every candidate was quantized, served by vLLM, and scored on the same held-out corpus, and the budget was spent where it bought the most. - Served by vLLM's Marlin kernels throughout — no fallback kernels. - Text + MTP speculative decoding + vision all verified. 1. The MLP, GatedDeltaNet and lmhead tensors are round-to-nearest, not AutoRound-tuned. Only qproj/kproj/vproj carry AutoRound tuning (they are inherited unchanged from an AutoRound int4 run). The measurements below were taken on round-to-nearest tensors, so…

Open weights apache-2.0 33.3B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen2.5-VL-32B-Instruct-AWQ

Qwen

In addition to the original formula, we have further enhanced Qwen2.5-VL-32B's mathematical and problem-solving abilities through reinforcement learning. This has also significantly improved the model's subjective user experience, with response styles adjusted to better align with human preferences. Particularly for objective queries such as mathematics, logical reasoning, and knowledge-based Q&A, the level of detail in responses and the clarity of formatting have been noticeably enhanced. In the past five months since Qwen2-VL’s release, numerous developers have built new models on the Qwen2-VL vision-language models, providing us with valuable feedback. During this period, we focused on…

Open weights apache-2.0 33.5B parameters 128,000 tokens transformers

Model · Image and text to text

Qwen2.5-VL-32B-Instruct

Qwen

In addition to the original formula, we have further enhanced Qwen2.5-VL-32B's mathematical and problem-solving abilities through reinforcement learning. This has also significantly improved the model's subjective user experience, with response styles adjusted to better align with human preferences. Particularly for objective queries such as mathematics, logical reasoning, and knowledge-based Q&A, the level of detail in responses and the clarity of formatting have been noticeably enhanced. In the past five months since Qwen2-VL’s release, numerous developers have built new models on the Qwen2-VL vision-language models, providing us with valuable feedback. During this period, we focused on…

Open weights apache-2.0 33.5B parameters 128,000 tokens transformers

Model · Image and text to text

gemma-4-31B-it

Google

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages. Featuring both Dense and Mixture-of-Experts (MoE) architectures, Gemma 4 is well-suited for tasks like text generation, coding, and reasoning. The models are available in five distinct sizes: E2B, E4B, 12B, 26B A4B, and 31B. Their diverse sizes make them deployable in environments ranging from…

Open weights apache-2.0 31.3B parameters 262,144 tokens transformers