This repository contains GGUF quantized weights for yothinS/Llama-3.2-3B-ThaiCRM, a model fine-tuned for Customer Relationship Management (CRM) workflows, Thai business communications, and customer behavioral analysis. GGUF files can be run locally on CPU, Apple Silicon, or NVIDIA/AMD GPUs via llama.cpp, Ollama, LM Studio, or carrany/text-generation-webui. Download your chosen GGUF file using the Hugging Face CLI: huggingface-cli download yothinS/Llama-3.2-3B-ThaiCRM-GGUF Llama-3.2-3B-ThaiCRM.Q4KM.gguf --local-dir.
Independent publisher
Yothin
yothinS
"Business and education: Aiming to make small models work to their full potential."
Models
Fine-tuned version of meta-llama/Llama-3.2-3B optimized for Customer Relationship Management (CRM), Customer Behavior Analysis, and Thai business workflows. Trained using Unsloth and Hugging Face's TRL library. from transformers import AutoModelForCausalLM, AutoTokenizer import torch modelid = "yothinS/Llama-3.2-3B-ThaiCRM" tokenizer = AutoTokenizer.frompretrained(modelid) model = AutoModelForCausalLM.frompretrained( modelid, torchdtype=torch.bfloat16, devicemap="auto" messages = [ {"role": "system", "content": "You are a professional CRM and customer behavior specialist."}, {"role": "user", "content": "วิเคราะห์พฤติกรรมลูกค้ารายนี้และแนะนำโปรโมชันรักษาฐานลูกค้า: ยอดซื้อเฉลี่ยลดลง 40%…
This repository contains quantized versions (GGUF) of the yothinS/Qwen3.5-4B-CRM model, which is a fine-tuned version of the Qwen series optimized for Customer Relationship Management (CRM) tasks, customer support, and business communications. These quantized versions drastically reduce memory usage and accelerate inference speeds, allowing you to run a powerful 4-billion parameter CRM assistant on consumer hardware, laptops, and edge devices. This repository provides GGUF format files optimized for inference using llama.cpp, Ollama, LM Studio, and other compatible software. Note: Memory Required includes the model weight footprint plus a standard context window buffer. To run the GGUF…