Model · Text generation
Empero
Developed by Empero GGUF quantizations of empero-ai/Qwen3.8-4B — a full-parameter distillation of Qwen3.8 2.4T A95B into the Qwen3.5-4B architecture — for llama.cpp, Ollama, LM Studio, Jan, KoboldCpp, and other stock GGUF runtimes. This card is about choosing a file and running it. The capability writeup, full benchmark results, and best practices live on the main model card. Headline results for the source model (CoT protocols, lm-evaluation-harness, identical settings base vs. student): Sizes are exact decimal GB from the uploaded files (1 GB = 1,000,000,000 bytes). Practical weight-size-based guidance at modest context — the KV cache is the dominant cost at long context and may require…
Open weights
apache-2.0
gguf
Model · Text generation
Empero
Developed by Empero GGUF quantizations of empero-ai/Qwen3.8-2B — a full-parameter distillation of Qwen3.8 2.4T A95B into the Qwen3.5-2B architecture, the smallest member of the family — for llama.cpp, Ollama, LM Studio, Jan, KoboldCpp, and other stock GGUF runtimes. This card is about choosing a file and running it. The capability writeup, full benchmark results, and best practices live on the main model card. Headline results for the source model (CoT protocols, lm-evaluation-harness, identical settings base vs. student): Sizes are exact decimal GB from the uploaded files (1 GB = 1,000,000,000 bytes). Practical weight-size-based guidance at modest context — the KV cache is the dominant…
Open weights
apache-2.0
gguf