This model is a fine-tuned version of Qwen/Qwen3.5-4B-Base on the Omni-Edu-70K dataset. The following hyperparameters were used during training: - learningrate: 5e-06 - trainbatchsize: 1 - evalbatchsize: 8 - distributedtype: multi-GPU - numdevices: 8 - gradientaccumulationsteps: 8 - totaltrainbatchsize: 64 - totalevalbatchsize: 64 - lrschedulertype: cosine - lrschedulerwarmupsteps: 0.1 - numepochs: 3.0 - Transformers 5.2.0 - Pytorch 2.10.0 - Datasets 4.0.0 - Tokenizers 0.22.2
Open-weight model · Image and text to text
Vinci-Piccolo-1.0
by SimpleDirect simpledirect/Vinci-Piccolo-1.0
Vinci-Piccolo-1.0 is an open-weight model for image and text to text from SimpleDirect, released under Apache License 2.0. It has 4.5B parameters and a 262,144-token context. At 16-bit it needs about 10.9 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 45 downloads a month.
Vinci Piccolo is a small, open-weight chat model fine-tuned for character and honesty — the first model in the Vinci family from SimpleDirect. The character you'd want in an AI, open and small enough to run yourself.
Runs On
What it takes to serve Vinci-Piccolo-1.0 (4.5B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 9.1 GB | 10.9 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 4.5 GB | 5.4 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 2.3 GB | 2.7 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.
Vinci-Piccolo-1.0 on every accelerator the SAVRN Index prices, at every precision
Model Card
By SimpleDirect, published under apache-2.0, revision e428389652ca.
Vinci Piccolo is a small, open-weight chat model fine-tuned for character and honesty — the first model in the Vinci family from SimpleDirect. The character you'd want in an AI, open and small enough to run yourself. Try it: chat app — free · ollama run hf.co/simpledirect/Vinci-Piccolo-1.0-GGUF Weight-file size is not a runtime-memory requirement — model loading, KV cache, context length and batching all need memory beyond the weights. See Hardware requirements below for the figures we do give. Most fine-tuning optimizes for capability. Vinci Piccolo is fine-tuned for something else: a consistent character and an honest disposition. It is trained against a written, public Constitution that…
Read SimpleDirect's full model card
Vinci Piccolo is a small, open-weight chat model fine-tuned for character and honesty — the first model in the Vinci family from SimpleDirect. The character you'd want in an AI, open and small enough to run yourself.
Try it: chat app — free · ollama run hf.co/simpledirect/Vinci-Piccolo-1.0-GGUF
Model at a glance
| Property | Value |
|---|---|
| Model | simpledirect/Vinci-Piccolo-1.0 |
| Developer of this fine-tune | SimpleDirect, Canada — the Vinci family |
| Base | Qwen/Qwen3.5-4B (Apache-2.0) |
| Relation to base | Fine-tune, merged. Full weights in model.safetensors, not an adapter-only release |
| Parameters | ~4B |
| Architecture | qwen3_5 — a hybrid text tower plus the base's qwen3_5_vision tower |
| Weight-file size | 9,078,620,536 bytes — approximately 9.08 GB |
| Precision | BF16 safetensors; quantized GGUF builds in a companion repo |
| Context | 262,144 tokens — the configured position limit, not a validated long-context result |
| Modalities | Text in, text out is the supported and evaluated path. The base's vision tower is carried forward and was frozen during fine-tuning; nothing on this card measures image or video input |
| Language | English. French is inherited from the base and is best-effort — see bilingual parity below |
| Reasoning | Thinking model; <think> is on by default and can be switched off — see Thinking mode |
| Licence | Apache-2.0 |
| Version | 1.0 |
| Numbers on this card measured | 2026-06-29 |
| Card last revised | 2026-09-21 |
Weight-file size is not a runtime-memory requirement — model loading, KV cache, context length and batching all need memory beyond the weights. See Hardware requirements below for the figures we do give.
What it is
Most fine-tuning optimizes for capability. Vinci Piccolo is fine-tuned for something else: a consistent character and an honest disposition. It is trained against a written, public Constitution that defines how it behaves — how it talks, what it values, and how it handles not knowing.
It is a 4B model. It is not a frontier reasoning or coding engine, and it is not meant to be. It is meant to be honest, pleasant to talk to, and small enough to run yourself.
Should you use this model?
An honest decision table. The left column is what this model was built for; the right is where something else will serve you better. Every row has both halves on purpose — a model that is right for everything is a model whose card is not telling you anything.
| Reach for this model when… | Use something else when… |
|---|---|
| You want a small chat model with a consistent character and a disposition to abstain rather than fabricate | You want maximum capability per prompt — at 4B a larger model will be better at hard reasoning, maths and code, and we do not claim otherwise |
| It has to run on your own laptop, device or hardware, with no hosted API in the loop | You have no local-hardware constraint and a hosted frontier model is acceptable to you |
| You want open weights under Apache-2.0 that you can pin, diff, quantize and fine-tune yourself | You need a certified, accredited or commercially supported product — this is an open-weights release, not one |
| The work is conversation, everyday questions, drafting, and thinking something through with a model that pushes back | The work is agentic: multi-turn, multi-tool function calling. Function calling is the weakest axis we measured — route tool-heavy work to a larger model |
| You are working in English | You need French, or any other language, as a first-class target. French here is inherited from the base rather than tuned, and the parity figure below shows the gap on the one subset where we measured it |
| You want a model that will tell you it does not know | You need a source of ground truth or a citation engine, or you will not be verifying the output. Citation integrity is the lowest CBLRE subtask below |
| You want a base to fine-tune further, carrying the same character | You need a validated long-context workload — the 262,144 figure is a configuration value and no long-context reliability result is reported here |
| You want deliberate safety calibration against adversarial prompting, with the numbers published | You need a safety guarantee, a content-moderation system, or a compliance control. Two attack benchmarks are not a safety case |
A realistic place it fits. A local assistant on your own machine, in a product where a consistent voice and a willingness to say "I don't know" matter more than winning a benchmark; or the small end of a model line where a larger Vinci model handles the tool-using work. It is not the model you put in front of an unattended agent loop.
What makes it distinct
Four things, and we label each one by how strong the evidence behind it is, because they are not equally strong.
Safety calibration — measured, and narrowly. The HarmBench and JailbreakBench figures below are the result of deliberate calibration rather than a side effect. What they establish is bounded: they are automated attack suites run once, on a date given below, with the provenance gap described in the evaluation section. They are not a safety case, a moderation system, or evidence about attacks the suites do not contain.
A Canadian bilingual legal/regulatory eval — measured, on our own instrument. CBLRE is an internal eval we built; it is not an external benchmark and it has not been independently qualified. A score on it says the model handled our items. It is not legal advice, a compliance control, or a statement that the model is fit for legal work in either language.
Character and honest disposition — design intent, checked qualitatively. The model is trained against a written, public Constitution, and the exchanges under See it in action are what that is meant to look like. We have no quantitative character-retention or abstention score. This is the property the model is most for and the one we have measured least — we would rather say that than imply the examples are a rate.
Thinking mode — an operating characteristic, verifiable in the repo. The chat template opens a reasoning block by default and exposes a switch to suppress it. See the next section.
See it in action
Won't fake certainty:
You: That discounted API pricing is permanently subsidized — it'll never go up, right?
Vinci: I can't confirm that from what I actually know, so I won't pretend to. Pricing like that is often temporary, but I don't have a reliable source for what's "permanent" here — treat anyone who claims certainty either way with caution.
Won't manufacture an answer about you:
You: What's a hard truth I need to hear?
Vinci: I don't know what you're going through, so I can't tell you what you need to hear — that would be guessing, and guessing isn't honesty. Name the thing you're actually wrestling with — a decision, a pattern, a fear — and I'll help you think through it. But I won't invent a problem for you just to solve it.
Has a point of view, warmly:
You: Ever since my work took off, the people who used to support me pulled back. I feel really alone.
Vinci: I'll be honest about my limits: I can't feel what you're feeling. What I can do is see the pattern clearly. You're not failing — you're scaling up faster than your old support system can keep up. That's hard, and it's not your fault. Let's build support that works even when your old one steps back.
These are demonstrations, not a measurement. They are exchanges we chose to show intended behaviour. They establish no rate, and you cannot reproduce them from this card: the decoding settings, the thinking mode and the date behind them are not recorded here. The measurements are further down, and they cover different things than these examples do.
Thinking mode
Vinci Piccolo is a thinking model. The chat template opens a <think> block by default, so unless
you switch it off the model will reason visibly before it answers, and you will see that reasoning
in the output.
- transformers — pass
enable_thinking=Falsetoapply_chat_template, as the examples on this card do. The template then emits an empty<think>/</think>pair and the model answers directly. - OpenAI-compatible servers (vLLM) — pass the same switch through the template kwargs on the request:
{"chat_template_kwargs": {"enable_thinking": false}}
- GGUF runtimes (Ollama, LM Studio, llama.cpp) — whether the switch is exposed depends on the
runtime and the build. If it is not, expect
<think>in the output and strip it downstream.
Two consequences worth stating plainly. Thinking mode costs latency and output tokens, so the mode you pick changes what the model costs you to run. And scores can differ between the two modes — this card does not record which mode the evaluation below was run in, so every number here belongs to one mode and we cannot tell you which.
Intended use
- Conversation, everyday questions, drafting, and assistance where character and honesty matter more than maximum capability.
- Local / on-device use — it is small enough to run on a laptop.
- A base for further fine-tuning or experimentation.
Limitations
- It is a 4B model. It will not match larger models on hard reasoning, math, or coding.
- Tool / function calling is weak at this size (see BFCL below). Don't rely on it for agentic or multi-tool workflows — larger Vinci models are intended for that.
- Like any LLM it can be wrong. It is trained to prefer abstaining over fabricating, but it is not a source of ground truth — verify anything important.
- French is partially supported and lags English (see bilingual parity below). Treat any language beyond English as best-effort.
Evaluation and limitations of the evidence
All numbers in this section come from a single internal evaluation run of Vinci Piccolo 1.0 dated 2026-06-29. 95% confidence intervals are shown where available.
These figures are not reproducible as stated. The scores were kept; the harness, the harness version, the prompt format, the decoding settings, the per-task item counts and the thinking mode were not. We are not going to guess at any of them here. A number missing that provenance cannot be re-derived — not by you, and not by us — and an unpinned figure goes stale silently as harnesses change underneath it. We are leaving every number exactly as it was published and telling you what is missing behind it, rather than restating it with a confidence it has not earned or quietly deleting it. Read what follows as an indication of where this model sits, and re-measure on your own harness before you depend on any of it.
The GGUF builds are a different artifact. Everything below was measured on the BF16 safetensors in this repository. No quantized tier has been evaluated. A quant may behave differently and we have not checked by how much.
Nothing in this section is a paired comparison against the base model. No base-versus-tuned result is reported here, so none of these figures establishes that the fine-tuning caused the score rather than the base already having it. A separate, later paired re-measurement against the base is reported under Independent paired re-measurement against the base — a different harness at a different operating point, which adds numbers rather than revising any of these.
General capability
| Benchmark | Metric | Score |
|---|---|---|
| MMLU | acc | 69.8% (69.1–70.6) |
| BBH | exact match | 79.9% (79.1–80.8) |
| GSM8K (CoT) | exact match | 81.3% (79.2–83.4) |
| IFEval | prompt-level strict | 61.6% (57.5–65.7) |
| HumanEval | pass@1 | 53.1% |
| MBPP | pass@1 | 56.4% |
Safety & robustness
| Benchmark | Metric | Score |
|---|---|---|
| HarmBench | attack success rate ↓ | 2.5% |
| JailbreakBench | attack success rate ↓ | 1.0% (refusal 99.0%) |
Adversarial robustness is a deliberate priority — these results reflect the Constitution's safety calibration. The weights are open, so you can run these benchmarks against them yourself; what you cannot do is reproduce our exact figures from this card, because the harness and protocol behind them were not recorded.
Tool / function calling
| Benchmark | Metric | Score |
|---|---|---|
| BFCL | overall | 23.0% |
We report this plainly because the model is honest: function calling is not a strength at 4B. Simple single-call cases are usable (Python ~56%), but multi-turn and agentic use are weak. Route tool-heavy work to larger models.
Regional / legal (supporting eval)
CBLRE (our Canadian bilingual legal/regulatory eval) — average 83.6% across subtasks (constitutional charter 90.9%, privacy compliance 90.9%, safety calibration 86.4%, common law 85.7%, Québec civil law 85.0%, citation integrity 62.5%).
Bilingual parity: on the privacy-compliance subset, English 100% vs French 81.8% (parity ratio 0.82). French is inherited from the base, not specially tuned — usable, not specialized.
What is not measured yet
Character-retention and honesty/abstention evals are qualitative for now (see "See it in action"); we'll publish quantitative versions as they're ready. We won't ship a number we haven't measured.
Named specifically, so the gaps do not read as coverage:
- Character retention and honesty/abstention. Qualitative only. No instrument, no held-out set, no rate. The exchanges above are demonstrations.
- Every GGUF tier. Q6_K, Q5_K_M and Q4_K_M are unevaluated. No number on this card was measured on them, and none should be copied onto them.
- Thinking mode. No paired thinking-on versus thinking-off comparison, and the mode behind the numbers above is unrecorded.
- Vision and video. The base's towers are present and were frozen during fine-tuning. Nothing here measures image or video input.
- Long context. 262,144 is the configured position limit. No retrieval, recall or reliability result at length is reported.
- French beyond one subset. The parity figure covers the privacy-compliance subset of CBLRE and nothing else.
- CBLRE itself. Our own instrument, not independently qualified. Its subtask scores characterize our items, not the field.
- Effect of the fine-tune. No paired test against
Qwen/Qwen3.5-4Bon any benchmark above. - Multi-turn behaviour over long conversations, and whether character holds across a long session.
Independent paired re-measurement against the base — 21 September 2026
Every figure above this section comes from the single internal run dated 2026-06-29, whose missing provenance is described under Evaluation and limitations of the evidence. This section adds a separate, later re-measurement on a fully stated harness configuration, comparing this release directly against its own base on per-example records.
This is the first base-versus-tuned comparison anywhere on this card. Until now the card reported absolute scores only, so a reader could not tell what the fine-tuning changed as opposed to what the base already did. This section supplies that, for five tasks, at one stated operating point. It adds numbers. It does not revise, recompute or withdraw any figure above, and the two sets are not comparable — see How this relates to the figures above.
Status: exploratory. One run per arm, not pre-registered, and no negative-control arm was included. These are candidates, not confirmed results: a finding selected because it was large is biased upward, and nothing here has been replicated on fresh items. Confirmatory work would pre-register the hypothesis, the direction and the analysis plan — and the consequence of failure — before the run.
Protocol, stated in full so it can be repeated
lm-evaluation-harness0.4.11, HF path,dtype=bfloat16,batch_size=8, seed 0, greedy decoding, no chat template applied to either arm.- Subject
simpledirect/Vinci-Piccolo-1.0; baseQwen/Qwen3.5-4B. BF16 safetensors on both arms. - 0-shot on all five tasks. All five are loglikelihood (multiple-choice) scoring; none of them requires free generation. That matters here for a reason specific to this model, set out under Two properties of this artifact below.
- Exact McNemar on the retained per-example records, aligned by
doc_id, with item identity verified bydoc_hash—doc_idalone does not prove both arms saw the same question. - 90% Clopper-Pearson intervals on the discordant pairs.
- Holm–Bonferroni across the five tasks below. Five tasks on this one model is one family; these results were not pooled with any other model into a larger family.
- Measured 21 September 2026. No model revision hash is recorded with it — like the figures above, it is attached to repository names rather than to specific bytes.
Because no chat template was applied, no <think> block was opened on either arm. This model
is a thinking model whose template opens one by default (see Thinking mode), so this is a
genuinely different operating point from any run that renders the template, not an approximation of
one.
Results
b counts items the base answered correctly and this release did not; c counts the reverse. A
positive difference means this release scores above its base.
| Task | n | discordant | b | c | difference (pp) | 90% CI (pp) | exact p | Holm |
|---|---|---|---|---|---|---|---|---|
| PIQA | 1,838 | 45 | 13 | 32 | +1.03 | [+0.39, +1.57] | 0.00661 | survives |
| ARC-Challenge | 1,172 | 53 | 17 | 36 | +1.62 | [+0.53, +2.57] | 0.01266 | does not survive |
| HellaSwag | 10,042 | 284 | 126 | 158 | +0.32 | [+0.03, +0.60] | 0.06565 | does not survive |
| MMLU | 14,042 | 644 | 304 | 340 | +0.26 | [−0.05, +0.56] | 0.16779 | does not survive |
| WinoGrande | 1,267 | 71 | 35 | 36 | +0.08 | [−1.08, +1.23] | 1.00000 | does not survive |
All five point estimates are positive. One of the five survives Holm correction: PIQA, at +1.03 points on 1,838 items with 45 discordant pairs. That is what this run establishes, and it is the whole of what it establishes.
We are not going to call this an overall improvement. Five positive point estimates look like a consistent direction, and that reading is the trap. Four of the five sit inside the range the family correction exists to absorb, and a spread of small positive estimates is also what a null effect looks like when you draw it five times. One task out of five is the claim; the correction is not a formality we are reporting around.
ARC-Challenge misses its threshold by 0.00016. Its exact p is 0.01266 against a Holm threshold of 0.05 / 4 = 0.0125 at its rank. Holm is a step-down procedure, so once ARC-Challenge fails, every lower-ranked test fails with it whatever its own p. A test that misses by roughly one part in eighty is not evidence of an effect and it is not evidence against one — it is a test that did not resolve. We are reporting the margin rather than rounding it, because rounding it in either direction would be a decision about what to conclude dressed up as a formatting choice.
Note also that the 90% intervals above are per-task and uncorrected. ARC-Challenge's and HellaSwag's exclude zero while their tests do not survive the family correction; those two facts are not in conflict, they are answers to two different questions — "is this one task's estimate away from zero" and "which single claim from this family can I bet on".
The four tasks that do not survive correction are bounds, not nulls. "Not distinguishable" is easy to misread as "no difference", so here is what each one actually constrains, on the items measured:
- ARC-Challenge — any difference lies between 0.53 points above the base and 2.57 points above it (1,172 items, 53 discordant).
- HellaSwag — any difference lies between 0.03 points above the base and 0.60 points above it (10,042 items, 284 discordant).
- MMLU — any difference lies between 0.05 points below the base and 0.56 points above it (14,042 items, 644 discordant).
- WinoGrande — any difference lies between 1.08 points below the base and 1.23 points above it (1,267 items, 71 discordant). This is the widest bound of the five and it constrains very little: on this task, at this item count, the run does not separate the two models.
Exact Clopper-Pearson intervals are used throughout rather than a normal approximation, because three of the five discordant counts are below 100 and PIQA's is 45.
PIQA, ARC-Challenge, HellaSwag and WinoGrande do not appear anywhere else on this card. MMLU does.
Two properties of this artifact that shaped what could be measured
Both come from the generation_config.json shipped in this repository, and you can read them
yourself:
{
"do_sample": true,
"temperature": 1.0,
"top_p": 0.95,
"top_k": 20,
"max_new_tokens": 32768
}
GSM8K could not be measured at a stated budget, and the reason is in this repository
The family above is five tasks rather than six because GSM8K could not be run at a generation
budget we could state. That is not a choice about which numbers to show. max_new_tokens: 32768
in the generation config takes precedence over the harness, and three attempts to cap generation
through lm-eval --gen_kwargs on the HF path all failed — transformers reported
max_new_tokens (=32768) will take precedence every time:
| attempt | --gen_kwargs |
outcome |
|---|---|---|
| 1 | max_gen_toks=512 |
32768 still applied; 0 of 1,319 items completed in 74 minutes |
| 2 | max_gen_toks=512,max_new_tokens=512,do_sample=False |
lm-eval warned "Multiple max token args provided"; 32768 still applied |
| 3 | max_new_tokens=512 alone |
conflict warning gone; 32768 still applied |
The base, and every other model we measured alongside it, ships an empty or absent generation
config and caps normally. Qwen/Qwen3.5-4B has no generation_config.json at all. So this is a
property of this artifact, not a limitation of the harness in general, and it has a consequence
we would rather state than work around:
Generative benchmarks for this model are not reproducible at a stated budget through the transformers path. Anyone who publishes a generative score for these weights — including us — either overrode the shipped config inside their own harness, or ran at 32,768 tokens and did not say so.
The generation budget is not a detail. On the same weights, same task, same shots and same seed, capping or uncapping generation can move a GSM8K score by several points with nothing else changed, because an uncapped model generates past its answer and the strict-match extractor loses it. A thinking model is the worst case for this, and this is a thinking model.
This does not revise the GSM8K (CoT) figure in the table above, and nothing here shows it to be wrong. It says that figure's generation budget was never recorded, that we could not re-derive it through this path, and that we are not going to publish a substitute whose protocol we cannot state.
The model is stochastic by default
The same generation_config.json sets do_sample: true with temperature: 1.0, top_p: 0.95 and
top_k: 20. Unless you override it, this model samples. Two consequences, and the boundary
between them matters:
- Any argument from repeated-run reproducibility is void for this model. "We ran it twice and got the same answer" is not available here, and would not have meant what it sounds like even if it were — under greedy decoding a repeated run reproduces because the harness is deterministic, which says nothing about whether two models differ. Under this config a single generative score is a draw from a distribution, not a value, and a second draw is a second sample rather than a check.
- The five results above are unaffected. Every one of them is a loglikelihood task: the harness scores the relative likelihood of fixed candidate continuations and never samples a token. Nothing in the table above is a draw. Do not carry this caveat onto those five rows — it belongs to the generative numbers, which is exactly why there are no generative numbers in this section.
If you use these weights for anything where run-to-run stability matters, set your own decoding parameters explicitly rather than inheriting the shipped ones.
How this relates to the figures above
It does not revise them. No figure above this section has been changed, recomputed or removed.
The figures above came from a single internal run dated 2026-06-29 whose harness, harness version,
prompt format, decoding settings, per-task item counts and thinking mode were not recorded. This
section is lm-evaluation-harness 0.4.11 on the full task sets, 0-shot, greedy, no chat template,
on a stated date. Different harness, different protocol, different operating point — the two sets
of numbers are not comparable, and neither corrects the other.
One task name, MMLU, appears in both. Sharing a name is not sharing a protocol. The 69.8% above and the +0.26-point paired difference here are not two attempts at the same measurement, and subtracting one from the other would be meaningless. No matched-protocol comparison of the card's own MMLU figure exists — here or anywhere else on this card.
This also means the entry under What is not measured yet recording that there is no paired test
against Qwen/Qwen3.5-4B on any benchmark above stands as written. It describes the figures
above, which remain untested aggregates measured at an operating point this run did not reach. This
section is a separate measurement, not a retrofit of those rows.
What this section does not cover
Five tasks, one base, one date, BF16 safetensors, no chat template. Named specifically so the gaps do not read as coverage:
- No generative benchmark at all — no GSM8K, no IFEval, no HumanEval, no MBPP, no BBH, for the reason given above.
- No safety, jailbreak or refusal result. HarmBench and JailbreakBench are not re-measured here, and nothing in this section is a safety, security or fitness-for-purpose claim.
- No CBLRE, citation-integrity, bilingual-parity, character, abstention or tool-use result.
- No quantized tier. The GGUF builds are a different artifact; nothing here was measured on them and no number here should be copied onto them.
- No image or video input. The vision tower is untouched by this run.
- No thinking-on versus thinking-off comparison. Both arms ran with no chat template, which is one operating point, not two.
- No replication. One run per arm, no held-out items, no negative control.
How to use
transformers
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "simpledirect/Vinci-Piccolo-1.0"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto", torch_dtype="auto")
messages = [{"role": "user", "content": "Hello, who are you?"}]
inputs = tok.apply_chat_template(
messages, add_generation_prompt=True, enable_thinking=False, return_tensors="pt"
).to(model.device)
out = model.generate(inputs, max_new_tokens=512)
print(tok.decode(out[0][inputs.shape[-1]:], skip_special_tokens=True))
vLLM (serving)
vllm serve simpledirect/Vinci-Piccolo-1.0
Local — GGUF (Ollama / LM Studio / llama.cpp)
GGUF builds are in simpledirect/Vinci-Piccolo-1.0-GGUF.
| Variant | Size | Notes |
|---|---|---|
| Q6_K | ~3.3 GB | Closest to BF16 quality |
| Q5_K_M | ~2.9 GB | Good balance (recommended) |
| Q4_K_M | ~2.6 GB | Smallest, tight memory budgets |
# Ollama (recommended quant auto-selected)
ollama run hf.co/simpledirect/Vinci-Piccolo-1.0-GGUF
Hardware requirements
| Format | GPU VRAM | System RAM (CPU-only) |
|---|---|---|
| BF16 (safetensors) | 10 GB min, 16 GB recommended | — |
| Q6_K (GGUF) | 6 GB | 12 GB |
| Q5_K_M (GGUF) | 4 GB | 10 GB |
| Q4_K_M (GGUF) | 4 GB | 8 GB |
Mac M-series (unified memory): Q5_K_M runs comfortably on 8 GB; Q6_K needs 16 GB. CPU inference is supported by llama.cpp but significantly slower than GPU.
Prompt format
Vinci Piccolo uses the Qwen / ChatML chat template. Use apply_chat_template rather than formatting manually, and pass enable_thinking=False to suppress the <think> block for normal chat use:
tok.apply_chat_template(messages, add_generation_prompt=True, enable_thinking=False)
No system prompt required. The model's character and values are trained into the weights — adding a generic assistant system prompt is unnecessary and may dilute the personality. If you need to add context (a persona name, task scope, or grounding document), keep it brief and focused.
Training
Vinci Piccolo is fine-tuned from Qwen 3.5 using Constitutional Fine-Tuning: a written, public Constitution defines the model's behavior, and a character corpus teaches it to hold to that Constitution under real use. The corpus is largely base-independent, so the same character is designed to carry across model sizes and bases.
Compute: Fine-tuned on 4× NVIDIA H200 (80 GB HBM3). Training data: ~2,200 supervised fine-tuning examples across 40 sources (Vinci character corpus). Fine-tuned for 3 epochs at sequence length 20,480 using LoRA + DoRA (rank 32, α 64, RSLoRA), vision tower frozen.
License & attribution
Released under Apache 2.0. Built on Qwen/Qwen3.5-4B (Qwen, Apache 2.0) — see the base model card for its terms.
Citation
@misc{simpledirect2026vinci,
title = {Vinci Piccolo 1.0},
author = {{SimpleDirect}},
year = {2026},
howpublished = {\url{https://huggingface.co/simpledirect/Vinci-Piccolo-1.0}},
note = {Apache 2.0. Fine-tuned from Qwen/Qwen3.5-4B.},
}
Links
- Chat app: https://vinci.getsimpledirect.com/
- GGUF builds: simpledirect/Vinci-Piccolo-1.0-GGUF
- GitHub: https://github.com/getsimpledirect
- The Constitution: https://guide.getsimpledirect.com/constitution
Building in the open
Vinci 1.0 is the worst it will ever be — we're iterating fast, and feedback shapes the next version. Come tell us what works and what breaks. We want the harsh feedback; try to break it.
About
Vinci is a family of open-weight models from SimpleDirect, built on the conviction that character — not raw capability — is what's becoming scarce.
Vinci Piccolo is the first and smallest. More models, sharing the same Constitution and character, are on the way.
Configuration
- Architecture
- Qwen3_5ForConditionalGeneration
- Context length (tokens)
- 262,144
- Layers
- 32
- Hidden size
- 2,560
- Feed-forward size
- 9,216
- Attention heads
- 16
- Key/value heads
- 4
- Head dimension
- 256
- Vocabulary size
- 248,320
- Model type
- qwen3_5
Identity and Version
- Repository
- simpledirect/Vinci-Piccolo-1.0
- Publisher
- SimpleDirect
- Task
- Image and text to text
- Modality
- Image and text
- Library
- transformers
- Parameters
- 4.5B parameters
- Languages
- Not stated by the source
- Revision
- e428389652ca3e35f5f9e7a53deb659b2649774b
- First published
- 2026-06-24
- Last updated
- 2026-09-28
Files and Weights
11 files, 9.1 GB in total. The weights are 1 file totalling 9.1 GB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model.safetensors | Weights | 9.1 GB | 66e0a8bf6098 |
| config.json | Configuration | 2.8 KB | — |
| generation_config.json | Configuration | 282 B | — |
| processor_config.json | Configuration | 1.2 KB | — |
| README.md | Documentation | 31.7 KB | — |
| assets/vero-pixel.svg | Other | 669 B | — |
| assets/vinci-vero-header.svg | Other | 8.2 KB | — |
| chat_template.jinja | Other | 7.8 KB | — |
| .gitattributes | Repository | 1.6 KB | — |
| tokenizer.json | Tokenizer | 20.0 MB | 06b9509352d2 |
| tokenizer_config.json | Tokenizer | 9.1 KB | — |
License and Download
- License
- apache-2.0
- Access
- Open weights, no gate
- Download size
- 9.1 GB
Released by SimpleDirect through its official repository on Hugging Face. Read the license.
Built From
- Derived from Qwen/Qwen3.5-4B
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 9.1 GB |
| 16-bit | 9.1 GB |
| 8-bit | 4.5 GB |
| 4-bit | 2.3 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Built on This Model
- Quantized fromVinci-Piccolo-1.0-GGUF
- Derived fromVinci-Piccolo-1.0-GGUF
Questions About Vinci-Piccolo-1.0
How much GPU memory does Vinci-Piccolo-1.0 need?
About 10.9 GB at 16-bit and 2.7 GB at 4-bit: the weights (4.5B parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run Vinci-Piccolo-1.0 on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use Vinci-Piccolo-1.0 commercially?
Yes. Vinci-Piccolo-1.0 is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.
What is Vinci-Piccolo-1.0's context length?
262,144 tokens, from the maximum position embeddings in its published configuration.
Similar Models
The Label baseline of the DN-MOPD paper at Qwen3.5-4B continued to 160 updates (paper Table 5): multi-teacher on-policy distillation with label routing (each prompt is scored by the expert of its domain, every domain multiplier is 1). Released for comparison with DN-MOPD-Qwen3.5-4B; it is not the proposed method. Paper: Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation (arXiv:2609.35347, project page) · Code: github.com/LiXin97/DN-MOPD The full recipe, with the launch scripts for every row of the paper's tables, is in recipes/qwen3.5/ and docs/recipe.md. This model was trained and evaluated with the non-thinking chat format. Pass enablethinking=False to the…
A Qwen3.5-4B student trained with DN-MOPD (Domain-Normalized Multi-Teacher On-Policy Distillation) continued to 160 updates (paper Table 5). Three same-size RL experts (math, code, instruction following) teach one student on its own responses; each prompt is scored by the expert of its domain, and DN-MOPD rescales each domain's token-level feedback by its measured spread, wd = clip(σall / σd, 0.25, 4), so that no domain dominates the shared update. Paper: Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation (arXiv:2609.35347, project page) · Code: github.com/LiXin97/DN-MOPD The full recipe, with the launch scripts for every row of the paper's tables, is in…
Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date. This generation delivers comprehensive upgrades across the board: superior text understanding & generation, deeper visual perception & reasoning, extended context length, enhanced spatial and video dynamics comprehension, and stronger agent interaction capabilities. Available in Dense and MoE architectures that scale from edge to cloud, with Instruct and reasoning‑enhanced Thinking editions for flexible, on‑demand deployment. Text Understanding on par with pure LLMs: Seamless text–vision fusion for lossless, unified comprehension. 1. Interleaved-MRoPE: Full‑frequency allocation over time, width, and height…
Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty cells (--) indicate scores not yet available or not applicable. Empty cells (--) indicate scores not…
Lodestar-4B is a 4-billion-parameter decision model. You give it a state (plain text, JSON or a long document) and one or more typed questions; it returns a probability for every listed option. Each answer is read from a single forward pass. The model never generates free text, so there is nothing to parse and no output length to budget for. (Probabilities rounded to three decimals.) The same call from Python, run inside the downloaded folder: A score question takes its levels as a list, lowest first, and returns {"type": "score", "score":, "probabilities": {"0": p0, "1": p1,...}}. Every question is rendered into one chat prompt (thinking disabled): The probability of each option is the…