# Prism Roleplay 1.5 Small Fast
**The Prism Roleplay 1.5 Small recipe on PrismML's 1-bit Bonsai-8B: a standalone model, faster and smaller than the 4-bit original, with a small quality gap**
What this is
A standalone model, not an adapter: load it directly with mlx-lm. It is Prism Roleplay 1.5 Small's training recipe (same data, hyperparameters and system prompt) applied to prism-ml/Bonsai-8B-mlx-1bit, a 1-bit (g128) build of Qwen3-8B.
How the LoRA was merged into a 1-bit model. A plain merge re-quantizes every layer to 1 bit, and the small LoRA update disappears in the rounding (we measured it: merged held-out loss 3.415 = untuned base 3.405). So the 16 layers the LoRA touched (all their attention and MLP projections) are merged at fp16 and stored at 6-bit; the other 20 layers, the embeddings and the LM head stay 1-bit. Result: held-out loss 2.341, identical to the unmerged LoRA (untuned 1-bit base: 3.405). We also tried 4-bit (2.423) and 2-bit (3.147) for the merged layers; both lose part of the tuning, so 6-bit is what ships.
Training
- Base:
prism-ml/Bonsai-8B-mlx-1bit
- Data: 7,570 train / 398 validation roleplay examples (real forum roleplay + two generations of synthetic craft data; quality-gated to >=230 words, >=3 paragraphs, dialogue present). Same corpus as Prism Roleplay 1.5 Small.
- Method: LoRA rank 8, scale 20, 16 layers, lr 1e-5, batch 1, sequence length 2,048, 2,500 iterations, then merged as above.
Evaluation
16 varied roleplay scenes, greedy decoding, 900-token cap. Quality is a blind pairwise judgment by nvidia/nemotron-3-super-120b-a12b, run in both A/B orders for every scene (about 32 judgments per comparison) so position bias cancels. Numbers below are for this merged model.
| Comparison |
Result |
Win rate |
| Fast vs Qwen3-8B base |
25W - 7L |
78% |
| Prism Roleplay 1.5 Small vs Qwen3-8B base |
27W - 4L |
87% |
| Fast vs Prism Roleplay 1.5 Small |
13W - 19L |
41% |
| System |
Formatting defects |
Avg words |
Avg paragraphs |
Dialogue present |
| Qwen3-8B base |
1 |
210 |
4.2 |
100% |
| Prism Roleplay 1.5 Small |
1 |
299 |
4.5 |
100% |
| Prism Roleplay 1.5 Small Fast |
0 |
294 |
5.2 |
100% |
| Speed / memory (Apple Silicon, this run) |
Decode |
Peak memory |
| Fast (this model) |
27.9 tok/s |
3.61 GB |
| Qwen3-8B 4-bit (same class as 1.5 Small) |
21.8 tok/s |
4.85 GB |
Fast as base + unmerged LoRA (in lora_adapter/) |
49.9 tok/s |
2.03 GB |
Read this honestly: Fast beats the untuned base and matches 1.5 Small on formatting, but the judge still prefers 1.5 Small in 59% of head-to-head comparisons. It is about 1.3x faster and uses about a quarter less memory than the 4-bit original. If you want the absolute smallest and fastest setup, the unmerged LoRA in lora_adapter/ (on top of the 1.3 GB Bonsai base) runs at about 50 tok/s in 2 GB, with somewhat lower quality (72% vs base, 31% vs 1.5 Small in our test). Caveats: 16 scenes is a small sample, the judge is an LLM, and the speed figure for 1.5 Small is its 4-bit base model's (the released model was already merged).
Usage
Requires PrismML's fork of MLX, which adds 1-bit kernels (stock MLX cannot run the 1-bit layers):
pip install mlx-lm
pip install mlx @ git+https://github.com/PrismML-Eng/mlx.git@prism
from mlx_lm import load, generate
model, tokenizer = load("VertexAGI/prism-roleplay-1.5-small-fast")
system = ("You are a skilled roleplay partner. Stay fully in character and write only your own "
"character's actions, speech and interiority. Write in flowing prose with paragraph "
"breaks, include dialogue, and never break character or address the reader.")
messages = [
{"role": "system", "content": system},
{"role": "user", "content": "Roleplay scene.\nGenre: noir mystery\nSetting: a flooded parking garage\nYour character: a fixer who owes the wrong person\nSituation: The other character shows you a photograph.\n\nWrite your character's next reply."},
]
prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False, enable_thinking=False)
print(generate(model, tokenizer, prompt=prompt, max_tokens=600))
Use the training system prompt above for best results.
Formats
MLX only (Apple Silicon), mixed 1-bit / 6-bit, about 3.1 GB. No GGUF: llama.cpp would need Bonsai's own 1-bit fork plus support for per-layer mixed precision, neither of which we have validated.
Limitations
Lower prose quality than Prism Roleplay 1.5 Small, and the base is a 1-bit compression of Qwen3-8B, so it starts from a noticeably weaker place (untuned loss 3.4 vs about 2 for the 4-bit base). The evaluation is small. Very long multi-turn sessions may drift.
License
Apache 2.0, inherited from Qwen3 and Bonsai.