Qwen3.5-4B-CRM - Quantized Versions
This repository contains quantized versions (GGUF) of the yothinS/Qwen3.5-4B-CRM model, which is a fine-tuned version of the Qwen series optimized for Customer Relationship Management (CRM) tasks, customer support, and business communications.
These quantized versions drastically reduce memory usage and accelerate inference speeds, allowing you to run a powerful 4-billion parameter CRM assistant on consumer hardware, laptops, and edge devices.
Available Formats & Model Files
This repository provides GGUF format files optimized for inference using llama.cpp, Ollama, LM Studio, and other compatible software.
| Format |
Recommended File / Quantization |
File Size |
Memory Required |
Target Hardware / Use Case |
| GGUF |
Qwen3.5-4B-CRM-f16.gguf |
8.67 GB |
~9.5+ GB |
High-end PC / Apple Silicon Mac. Maximum precision, no quantization loss. |
| GGUF |
Qwen3.5-4B-Thai-CRM-Q8_0.gguf |
4.61 GB |
~5.2 GB |
Near-lossless quality. Balanced CPU/GPU inference with high precision. |
| GGUF |
Qwen3.5-4B-Thai-CRM-Q6_K.gguf |
3.56 GB |
~4.2 GB |
Good balance of generation quality and performance for mid-range hardware. |
| GGUF |
Qwen3.5-4B-Thai-CRM-Q5_K_M.gguf |
3.16 GB |
~3.8 GB |
Recommended for general use. Best trade-off between file size and capabilities. |
| GGUF |
Qwen3.5-4B-Thai-CRM-Q4_K_M.gguf |
2.78 GB |
~3.4 GB |
Standard 4-bit quantization. Ideal for low-VRAM devices or quick testing. |
| GGUF |
Qwen3.5-4B-Thai-CRM-Q3_K_M.gguf |
2.32 GB |
~2.9 GB |
Low memory footprint. Noticeable drop in response accuracy. |
| GGUF |
Qwen3.5-4B-Thai-CRM-Q2_K.gguf |
1.96 GB |
~2.5 GB |
Minimum size. Designed for highly constrained environments or mobile devices. |
Note: Memory Required includes the model weight footprint plus a standard context window buffer.
Quick Start & Usage Guide
1. Using GGUF Format with llama.cpp
To run the GGUF version locally in your terminal, use the following commands:
# Run inference with the downloaded .gguf file
./llama-cli -m ./Qwen3.5-4B-CRM-Q4_K_M.gguf -n 512 --color -p "<|im_start|>system\nYou are a helpful CRM assistant.<|im_end|>\n<|im_start|>user\nHelp me write a follow-up email to a client who hasn't responded in 3 days.<|im_end|>\n<|im_start|>assistant\n"
Model Capabilities & Intended Use
This quantized model retains the core capabilities of the base yothinS/Qwen3.5-4B-CRM model, including:
* CRM Automation: Generating follow-up emails, sales pitches, and support ticket summaries.
* Multilingual Business Context: Strong comprehension and generation of professional responses in English and Thai.
* Customer Support: Resolving queries with appropriate professional empathy, clarity, and tone alignment.
Quantization Method & Evaluation Notes
- GGUF weights were processed using the standard
llama.cpp quantization toolkits.
- Expect minimal perplexity degradation on 8-bit versions and minor trade-offs in edge-case reasoning on 4-bit configurations, which is normal across compact architectures.
Licensing
This model is distributed under the Apache 2.0 License, inherited from the base Qwen framework. Please adhere to the usage constraints and terms defined by Alibaba Qwen team.