The Llama 3.2 collection of multilingual large language models (LLMs) is a collection of pretrained and instruction-tuned generative models in 1B and 3B sizes (text in/text out). The Llama 3.2 instruction-tuned text only models are optimized for multilingual dialogue use cases, including agentic retrieval and summarization tasks. They outperform many of the available open source and closed chat models on common industry benchmarks. Model Architecture: Llama 3.2 is an auto-regressive language model that uses an optimized transformer architecture. The tuned versions use supervised fine-tuning (SFT) and reinforcement learning with human feedback (RLHF) to align with human preferences for…
Access requested at publisher
llama3.2
3.2B parameters
transformers
This llama model was trained 2x faster with Unsloth and Huggingface's TRL library.
Open weights
apache-2.0
3.2B parameters
131,072 tokens
transformers
Y
Model · Text generation
Yothin
Fine-tuned version of meta-llama/Llama-3.2-3B optimized for Customer Relationship Management (CRM), Customer Behavior Analysis, and Thai business workflows. Trained using Unsloth and Hugging Face's TRL library. from transformers import AutoModelForCausalLM, AutoTokenizer import torch modelid = "yothinS/Llama-3.2-3B-ThaiCRM" tokenizer = AutoTokenizer.frompretrained(modelid) model = AutoModelForCausalLM.frompretrained( modelid, torchdtype=torch.bfloat16, devicemap="auto" messages = [ {"role": "system", "content": "You are a professional CRM and customer behavior specialist."}, {"role": "user", "content": "วิเคราะห์พฤติกรรมลูกค้ารายนี้และแนะนำโปรโมชันรักษาฐานลูกค้า: ยอดซื้อเฉลี่ยลดลง 40%…
Open weights
llama3.2
3.2B parameters
131,072 tokens
Most language models today act as polite chatbots: they offer conversational summaries, generic advice, and surface-level bullet points. But when the stakes are existential — when critical production services fail, corporate partnerships fracture, infrastructure is compromised, or you are facing high-consequence decisions under severe uncertainty — you don't need a conversational chatbot. You need a dedicated Directorate of Intelligence on your team. What if you could run an autonomous operational intelligence analyst directly on your local workstation — 100% private, offline, and air-gapped? That is the genesis of Sovereign Gotham / Sovereign Anthology. Rather than feeding an AI thousands…
Open weights
gemma
3.2B parameters
8,192 tokens
transformers
This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).
Open weights
3.3B parameters
131,072 tokens
transformers
The Prism Roleplay 1.5 Small recipe on PrismML's 1-bit Bonsai-8B: a standalone model, faster and smaller than the 4-bit original, with a small quality gap A standalone model, not an adapter: load it directly with mlx-lm. It is Prism Roleplay 1.5 Small's training recipe (same data, hyperparameters and system prompt) applied to prism-ml/Bonsai-8B-mlx-1bit, a 1-bit (g128) build of Qwen3-8B. How the LoRA was merged into a 1-bit model. A plain merge re-quantizes every layer to 1 bit, and the small LoRA update disappears in the rounding (we measured it: merged held-out loss 3.415 = untuned base 3.405). So the 16 layers the LoRA touched (all their attention and MLP projections) are merged at fp16…
Open weights
apache-2.0
3.3B parameters
65,536 tokens
mlx