This llama model was trained 2x faster with Unsloth and Huggingface's TRL library.
Open-weight model · Text generation
kanana-1.5-8b-instruct-2505-Persona-Merged
by Hyungho Byun NotoriousH2/kanana-1.5-8b-instruct-2505-Persona-Merged
kanana-1.5-8b-instruct-2505-Persona-Merged is an open-weight model for text generation from Hyungho Byun, released under Apache License 2.0. It has 8B parameters and a 32,768-token context. At 16-bit it needs about 19.3 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 423 downloads a month.
This llama model was trained 2x faster with Unsloth and Huggingface's TRL library.
Runs On
What it takes to serve kanana-1.5-8b-instruct-2505-Persona-Merged (8B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 16.1 GB | 19.3 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 8.0 GB | 9.6 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 4.0 GB | 4.8 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.
Model Card
By Hyungho Byun, published under apache-2.0, revision 40a8a8f9024f.
This llama model was trained 2x faster with Unsloth and Huggingface's TRL library.
Read Hyungho Byun's full model card
Uploaded finetuned model
- Developed by: NotoriousH2
- License: apache-2.0
- Finetuned from model : kakaocorp/kanana-1.5-8b-instruct-2505
This llama model was trained 2x faster with Unsloth and Huggingface's TRL library.
Configuration
- Architecture
- LlamaForCausalLM
- Context length (tokens)
- 32,768
- Layers
- 32
- Hidden size
- 4,096
- Feed-forward size
- 14,336
- Attention heads
- 32
- Key/value heads
- 8
- Head dimension
- 128
- Vocabulary size
- 128,259
- Stored precision
- bfloat16
- Model type
- llama
Identity and Version
- Repository
- NotoriousH2/kanana-1.5-8b-instruct-2505-Persona-Merged
- Publisher
- Hyungho Byun
- Task
- Text generation
- Modality
- Text
- Library
- transformers
- Parameters
- 8B parameters
- Languages
- en
- Revision
- 40a8a8f9024f8df1684beb2c1969471f318ac0c2
- First published
- 2026-05-21
- Last updated
- 2026-10-01
Files and Weights
13 files, 32.1 GB in total. The weights are 5 files totalling 32.1 GB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model-00001-of-00004.safetensors | Weights | 4.9 GB | 9eb6b1b4312a |
| model-00002-of-00004.safetensors | Weights | 4.9 GB | a3f6f2726132 |
| model-00003-of-00004.safetensors | Weights | 4.9 GB | 407af2597b04 |
| model-00004-of-00004.safetensors | Weights | 1.3 GB | 85e1f4ab43dd |
| model.safetensors | Weights | 16.1 GB | a49a992662be |
| config.json | Configuration | 777 B | — |
| generation_config.json | Configuration | 170 B | — |
| model.safetensors.index.json | Configuration | 23.9 KB | — |
| README.md | Documentation | 602 B | — |
| chat_template.jinja | Other | 14.3 KB | — |
| .gitattributes | Repository | 1.6 KB | — |
| tokenizer.json | Tokenizer | 17.2 MB | 31d17e500c82 |
| tokenizer_config.json | Tokenizer | 66.3 KB | — |
License and Download
- License
- apache-2.0
- Access
- Open weights, no gate
- Download size
- 32.1 GB
Released by Hyungho Byun through its official repository on Hugging Face. Read the license.
Built From
- Derived from kakaocorp/kanana-1.5-8b-instruct-2505
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 32.1 GB |
| 16-bit | 16.1 GB |
| 8-bit | 8.0 GB |
| 4-bit | 4.0 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About kanana-1.5-8b-instruct-2505-Persona-Merged
How much GPU memory does kanana-1.5-8b-instruct-2505-Persona-Merged need?
About 19.3 GB at 16-bit and 4.8 GB at 4-bit: the weights (8B parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run kanana-1.5-8b-instruct-2505-Persona-Merged on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use kanana-1.5-8b-instruct-2505-Persona-Merged commercially?
Yes. kanana-1.5-8b-instruct-2505-Persona-Merged is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.
What is kanana-1.5-8b-instruct-2505-Persona-Merged's context length?
32,768 tokens, from the maximum position embeddings in its published configuration.
Similar Models
This llama model was trained 2x faster with Unsloth and Huggingface's TRL library.
This llama model was trained 2x faster with Unsloth and Huggingface's TRL library.
This llama model was trained 2x faster with Unsloth and Huggingface's TRL library.
The Meta Llama 3.1 collection of multilingual large language models (LLMs) is a collection of pretrained and instruction tuned generative models in 8B, 70B and 405B sizes (text in/text out). The Llama 3.1 instruction tuned text only models (8B, 70B, 405B) are optimized for multilingual dialogue use cases and outperform many of the available open source and closed chat models on common industry benchmarks. Model Architecture: Llama 3.1 is an auto-regressive language model that uses an optimized transformer architecture. The tuned versions use supervised fine-tuning (SFT) and reinforcement learning with human feedback (RLHF) to align with human preferences for helpfulness and safety.…
The Model mlx-community/Llama-3.1-8B-Instruct-4bit was converted to MLX format from meta-llama/Llama-3.1-8B-Instruct using mlx-lm version 0.21.4.