GGUF (quantized) builds of Vinci Piccolo, for local inference with Ollama, LM Studio, and llama.cpp. For the full-precision weights, evals, and details, see simpledirect/Vinci-Piccolo-1.0.
These files are not the measured artifact
No GGUF file in this repository has been evaluated. The benchmark figures on the full-precision card were measured on the unquantized bf16 weights, not on any file here. Quantization shifts scores, and with no per-tier measurement this card cannot say in which direction or by how much for any tier.
This card therefore publishes no per-tier score, no quality ranking between the tiers, and no recommended tier. The file sizes and memory figures below are measured, and describe storage and memory cost only. A file's presence in this repository is not a performance claim about that tier, and not a claim of parity with the full-precision weights.
Available variants
| File |
Size |
Min RAM |
Notes |
vinci-piccolo-1.0-20260629-Q6_K.gguf |
3.46 GB |
12 GB |
6-bit K-quant |
vinci-piccolo-1.0-20260629-Q5_K_M.gguf |
3.07 GB |
10 GB |
5-bit K-quant, medium mixture |
vinci-piccolo-1.0-20260629-Q4_K_M.gguf |
2.71 GB |
8 GB |
4-bit K-quant, medium mixture |
GPU: Q5_K_M and Q4_K_M run on 4 GB VRAM; Q6_K needs 6 GB. Mac M-series: Q5_K_M fits on 8 GB unified memory; Q6_K needs 16 GB.
Ollama
ollama run hf.co/simpledirect/Vinci-Piccolo-1.0-GGUF
llama.cpp
./llama-cli \
-m vinci-piccolo-1.0-20260629-Q5_K_M.gguf \
--ctx-size 262144 \
--temp 0 \
--chat-template qwen3
llama-server (OpenAI-compatible API)
./llama-server \
-m vinci-piccolo-1.0-20260629-Q5_K_M.gguf \
--ctx-size 262144 \
--host 0.0.0.0 \
--port 8080
Prompt format
Qwen / ChatML chat template. No system prompt required — character is trained into the weights. Pass enable_thinking=False when using the tokenizer directly to suppress <think> output.
Citation
@misc{simpledirect2026vinci,
title = {Vinci Piccolo 1.0},
author = {{SimpleDirect}},
year = {2026},
howpublished = {\url{https://huggingface.co/simpledirect/Vinci-Piccolo-1.0}},
note = {Apache 2.0. Fine-tuned from Qwen/Qwen3.5-4B.},
}