SAVRN
Search Contact SAVRN

Open-weight model · Text generation

prism-roleplay-1.5-small-fast

by Vertex AGI VertexAGI/prism-roleplay-1.5-small-fast

prism-roleplay-1.5-small-fast is an open-weight model for text generation from Vertex AGI, released under Apache License 2.0. It has 3.3B parameters and a 65,536-token context. At 16-bit it needs about 8 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

The Prism Roleplay 1.5 Small recipe on PrismML's 1-bit Bonsai-8B: a standalone model, faster and smaller than the 4-bit original, with a small quality gap A standalone model, not an adapter: load it directly with mlx-lm.

Parameters3.3B
Context65,536
Weights3.3 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve prism-roleplay-1.5-small-fast (3.3B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 6.7 GB 8.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 3.3 GB 4.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 1.7 GB 2.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

prism-roleplay-1.5-small-fast on every accelerator the SAVRN Index prices, at every precision

Model Card

By Vertex AGI, published under apache-2.0, revision 694ca939c31c.

The Prism Roleplay 1.5 Small recipe on PrismML's 1-bit Bonsai-8B: a standalone model, faster and smaller than the 4-bit original, with a small quality gap A standalone model, not an adapter: load it directly with mlx-lm. It is Prism Roleplay 1.5 Small's training recipe (same data, hyperparameters and system prompt) applied to prism-ml/Bonsai-8B-mlx-1bit, a 1-bit (g128) build of Qwen3-8B. How the LoRA was merged into a 1-bit model. A plain merge re-quantizes every layer to 1 bit, and the small LoRA update disappears in the rounding (we measured it: merged held-out loss 3.415 = untuned base 3.405). So the 16 layers the LoRA touched (all their attention and MLP projections) are merged at fp16…

Read Vertex AGI's full model card
# Prism Roleplay 1.5 Small Fast **The Prism Roleplay 1.5 Small recipe on PrismML's 1-bit Bonsai-8B: a standalone model, faster and smaller than the 4-bit original, with a small quality gap**

What this is

A standalone model, not an adapter: load it directly with mlx-lm. It is Prism Roleplay 1.5 Small's training recipe (same data, hyperparameters and system prompt) applied to prism-ml/Bonsai-8B-mlx-1bit, a 1-bit (g128) build of Qwen3-8B.

How the LoRA was merged into a 1-bit model. A plain merge re-quantizes every layer to 1 bit, and the small LoRA update disappears in the rounding (we measured it: merged held-out loss 3.415 = untuned base 3.405). So the 16 layers the LoRA touched (all their attention and MLP projections) are merged at fp16 and stored at 6-bit; the other 20 layers, the embeddings and the LM head stay 1-bit. Result: held-out loss 2.341, identical to the unmerged LoRA (untuned 1-bit base: 3.405). We also tried 4-bit (2.423) and 2-bit (3.147) for the merged layers; both lose part of the tuning, so 6-bit is what ships.

Training

  • Base: prism-ml/Bonsai-8B-mlx-1bit
  • Data: 7,570 train / 398 validation roleplay examples (real forum roleplay + two generations of synthetic craft data; quality-gated to >=230 words, >=3 paragraphs, dialogue present). Same corpus as Prism Roleplay 1.5 Small.
  • Method: LoRA rank 8, scale 20, 16 layers, lr 1e-5, batch 1, sequence length 2,048, 2,500 iterations, then merged as above.

Evaluation

16 varied roleplay scenes, greedy decoding, 900-token cap. Quality is a blind pairwise judgment by nvidia/nemotron-3-super-120b-a12b, run in both A/B orders for every scene (about 32 judgments per comparison) so position bias cancels. Numbers below are for this merged model.

Comparison Result Win rate
Fast vs Qwen3-8B base 25W - 7L 78%
Prism Roleplay 1.5 Small vs Qwen3-8B base 27W - 4L 87%
Fast vs Prism Roleplay 1.5 Small 13W - 19L 41%
System Formatting defects Avg words Avg paragraphs Dialogue present
Qwen3-8B base 1 210 4.2 100%
Prism Roleplay 1.5 Small 1 299 4.5 100%
Prism Roleplay 1.5 Small Fast 0 294 5.2 100%
Speed / memory (Apple Silicon, this run) Decode Peak memory
Fast (this model) 27.9 tok/s 3.61 GB
Qwen3-8B 4-bit (same class as 1.5 Small) 21.8 tok/s 4.85 GB
Fast as base + unmerged LoRA (in lora_adapter/) 49.9 tok/s 2.03 GB

Read this honestly: Fast beats the untuned base and matches 1.5 Small on formatting, but the judge still prefers 1.5 Small in 59% of head-to-head comparisons. It is about 1.3x faster and uses about a quarter less memory than the 4-bit original. If you want the absolute smallest and fastest setup, the unmerged LoRA in lora_adapter/ (on top of the 1.3 GB Bonsai base) runs at about 50 tok/s in 2 GB, with somewhat lower quality (72% vs base, 31% vs 1.5 Small in our test). Caveats: 16 scenes is a small sample, the judge is an LLM, and the speed figure for 1.5 Small is its 4-bit base model's (the released model was already merged).

Usage

Requires PrismML's fork of MLX, which adds 1-bit kernels (stock MLX cannot run the 1-bit layers):

pip install mlx-lm
pip install mlx @ git+https://github.com/PrismML-Eng/mlx.git@prism
from mlx_lm import load, generate

model, tokenizer = load("VertexAGI/prism-roleplay-1.5-small-fast")

system = ("You are a skilled roleplay partner. Stay fully in character and write only your own "
          "character's actions, speech and interiority. Write in flowing prose with paragraph "
          "breaks, include dialogue, and never break character or address the reader.")
messages = [
    {"role": "system", "content": system},
    {"role": "user", "content": "Roleplay scene.\nGenre: noir mystery\nSetting: a flooded parking garage\nYour character: a fixer who owes the wrong person\nSituation: The other character shows you a photograph.\n\nWrite your character's next reply."},
]
prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False, enable_thinking=False)
print(generate(model, tokenizer, prompt=prompt, max_tokens=600))

Use the training system prompt above for best results.

Formats

MLX only (Apple Silicon), mixed 1-bit / 6-bit, about 3.1 GB. No GGUF: llama.cpp would need Bonsai's own 1-bit fork plus support for per-layer mixed precision, neither of which we have validated.

Limitations

Lower prose quality than Prism Roleplay 1.5 Small, and the base is a 1-bit compression of Qwen3-8B, so it starts from a noticeably weaker place (untuned loss 3.4 vs about 2 for the 4-bit base). The evaluation is small. Very long multi-turn sessions may drift.

License

Apache 2.0, inherited from Qwen3 and Bonsai.

Configuration

Architecture
Qwen3ForCausalLM
Context length (tokens)
65,536
Layers
36
Hidden size
4,096
Feed-forward size
12,288
Attention heads
32
Key/value heads
8
Head dimension
128
Vocabulary size
151,669
RoPE base
1e+06
Model type
qwen3

Identity and Version

Repository
VertexAGI/prism-roleplay-1.5-small-fast
Publisher
Vertex AGI
Task
Text generation
Modality
Text
Library
mlx
Parameters
3.3B parameters
Languages
en
Revision
694ca939c31ca8589bbb38c60370ea8814ef2444
First published
2026-09-24
Last updated
2026-09-25

Files and Weights

12 files, 3.4 GB in total. The weights are 2 files totalling 3.3 GB in safetensors.

Weights2 files · 3.3 GB
Configuration4 files · 105.6 KB
Tokenizer2 files · 11.4 MB
Documentation1 file · 5.7 KB
Other2 files · 995.9 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
lora_adapter/adapters.safetensorsWeights38.8 MB fff8e2e0197f
model.safetensorsWeights3.3 GB fadae90c02fb
config.jsonConfiguration12.9 KB —
eval_results.jsonConfiguration27.5 KB —
lora_adapter/adapter_config.jsonConfiguration1.0 KB —
model.safetensors.index.jsonConfiguration64.1 KB —
README.mdDocumentation5.7 KB —
chat_template.jinjaOther4.1 KB —
logo.pngOther991.9 KB b8e0b4e45168
.gitattributesRepository1.6 KB —
tokenizer.jsonTokenizer11.4 MB be75606093db
tokenizer_config.jsonTokenizer700 B —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
3.3 GB
Download from Vertex AGI

Released by Vertex AGI through its official repository on Hugging Face. Read the license.

Built From

  • Adapter of prism-ml/Bonsai-8B-mlx-1bit
  • Derived from prism-ml/Bonsai-8B-mlx-1bit

Memory Requirements

PrecisionWeights in memory
As published3.3 GB
16-bit6.7 GB
8-bit3.3 GB
4-bit1.7 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About prism-roleplay-1.5-small-fast

How much GPU memory does prism-roleplay-1.5-small-fast need?

About 8 GB at 16-bit and 2 GB at 4-bit: the weights (3.3B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run prism-roleplay-1.5-small-fast on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use prism-roleplay-1.5-small-fast commercially?

Yes. prism-roleplay-1.5-small-fast is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is prism-roleplay-1.5-small-fast's context length?

65,536 tokens, from the maximum position embeddings in its published configuration.

Similar Models

This is a MarinSkyRL-native Open-MOPD student after 32 optimizer steps. It starts from the authors' mixed-domain SFT checkpoint. Student responses were scored by the authors' math, code, and instruction-following RL teachers, routed by domain. The objective uses the student's selected top-16 token IDs and a clipped policy surrogate. This is an early checkpoint, not the authors' step-200 final model. The checkpoint is an unquantized, six-file Hugging Face export of the durable MarinSkyRL globalstep32 FSDP2 checkpoint. The policy export was used for the independent step-32 evaluation. The export's model.safetensors SHA-256 is bb7326640142069bc2e1fba5f54f15e0cccb1ff861f34f318b372eaab7abaf4b.…

Open weights apache-2.0 3.3B parameters 65,536 tokens transformers

Model · Text generation

amharic-bell-tts-4bit

BmanCman

This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).

Open weights 3.3B parameters 131,072 tokens transformers

Model · Text generation

PowerMoE-3b

IBM Research

PowerMoE-3B is a 3B sparse Mixture-of-Experts (sMoE) language model trained with the Power learning rate scheduler. It sparsely activates 800M parameters for each token. It is trained on a mix of open-source and proprietary datasets. PowerMoE-3B has shown promising results compared to other dense models with 2x activate parameters across various benchmarks, including natural language multi-choices, code generation, and math reasoning. This is a simple example of how to use PowerMoE-3b model.

Open weights apache-2.0 3.4B parameters 4,096 tokens transformers

Model · Text generation

Llama-3.2-3B-Instruct

Meta Llama

The Llama 3.2 collection of multilingual large language models (LLMs) is a collection of pretrained and instruction-tuned generative models in 1B and 3B sizes (text in/text out). The Llama 3.2 instruction-tuned text only models are optimized for multilingual dialogue use cases, including agentic retrieval and summarization tasks. They outperform many of the available open source and closed chat models on common industry benchmarks. Model Architecture: Llama 3.2 is an auto-regressive language model that uses an optimized transformer architecture. The tuned versions use supervised fine-tuning (SFT) and reinforcement learning with human feedback (RLHF) to align with human preferences for…

Access requested at publisher llama3.2 3.2B parameters transformers