The Prism Roleplay 1.5 Small recipe on PrismML's 1-bit Bonsai-8B: a standalone model, faster and smaller than the 4-bit original, with a small quality gap A standalone model, not an adapter: load it directly with mlx-lm. It is Prism Roleplay 1.5 Small's training recipe (same data, hyperparameters and system prompt) applied to prism-ml/Bonsai-8B-mlx-1bit, a 1-bit (g128) build of Qwen3-8B. How the LoRA was merged into a 1-bit model. A plain merge re-quantizes every layer to 1 bit, and the small LoRA update disappears in the rounding (we measured it: merged held-out loss 3.415 = untuned base 3.405). So the 16 layers the LoRA touched (all their attention and MLP projections) are merged at fp16…
Open weights
apache-2.0
3.3B parameters
65,536 tokens
mlx
This is a MarinSkyRL-native Open-MOPD student after 32 optimizer steps. It starts from the authors' mixed-domain SFT checkpoint. Student responses were scored by the authors' math, code, and instruction-following RL teachers, routed by domain. The objective uses the student's selected top-16 token IDs and a clipped policy surrogate. This is an early checkpoint, not the authors' step-200 final model. The checkpoint is an unquantized, six-file Hugging Face export of the durable MarinSkyRL globalstep32 FSDP2 checkpoint. The policy export was used for the independent step-32 evaluation. The export's model.safetensors SHA-256 is bb7326640142069bc2e1fba5f54f15e0cccb1ff861f34f318b372eaab7abaf4b.…
Open weights
apache-2.0
3.3B parameters
65,536 tokens
transformers
PowerMoE-3B is a 3B sparse Mixture-of-Experts (sMoE) language model trained with the Power learning rate scheduler. It sparsely activates 800M parameters for each token. It is trained on a mix of open-source and proprietary datasets. PowerMoE-3B has shown promising results compared to other dense models with 2x activate parameters across various benchmarks, including natural language multi-choices, code generation, and math reasoning. This is a simple example of how to use PowerMoE-3b model.
Open weights
apache-2.0
3.4B parameters
4,096 tokens
transformers
The Llama 3.2 collection of multilingual large language models (LLMs) is a collection of pretrained and instruction-tuned generative models in 1B and 3B sizes (text in/text out). The Llama 3.2 instruction-tuned text only models are optimized for multilingual dialogue use cases, including agentic retrieval and summarization tasks. They outperform many of the available open source and closed chat models on common industry benchmarks. Model Architecture: Llama 3.2 is an auto-regressive language model that uses an optimized transformer architecture. The tuned versions use supervised fine-tuning (SFT) and reinforcement learning with human feedback (RLHF) to align with human preferences for…
Access requested at publisher
llama3.2
3.2B parameters
transformers
This llama model was trained 2x faster with Unsloth and Huggingface's TRL library.
Open weights
apache-2.0
3.2B parameters
131,072 tokens
transformers
This llama model was trained 2x faster with Unsloth and Huggingface's TRL library.
Open weights
apache-2.0
3.2B parameters
131,072 tokens
transformers