SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

Vinci-Piccolo-1.0

by SimpleDirect simpledirect/Vinci-Piccolo-1.0

Vinci-Piccolo-1.0 is an open-weight model for image and text to text from SimpleDirect, released under Apache License 2.0. It has 4.5B parameters and a 262,144-token context. At 16-bit it needs about 10.9 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 45 downloads a month.

Vinci Piccolo is a small, open-weight chat model fine-tuned for character and honesty — the first model in the Vinci family from SimpleDirect. The character you'd want in an AI, open and small enough to run yourself.

Parameters4.5B
Context262,144
Weights9.1 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads45

Runs On

What it takes to serve Vinci-Piccolo-1.0 (4.5B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 9.1 GB 10.9 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 4.5 GB 5.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 2.3 GB 2.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.

Vinci-Piccolo-1.0 on every accelerator the SAVRN Index prices, at every precision

Model Card

By SimpleDirect, published under apache-2.0, revision e428389652ca.

Vinci Piccolo is a small, open-weight chat model fine-tuned for character and honesty — the first model in the Vinci family from SimpleDirect. The character you'd want in an AI, open and small enough to run yourself. Try it: chat app — free · ollama run hf.co/simpledirect/Vinci-Piccolo-1.0-GGUF Weight-file size is not a runtime-memory requirement — model loading, KV cache, context length and batching all need memory beyond the weights. See Hardware requirements below for the figures we do give. Most fine-tuning optimizes for capability. Vinci Piccolo is fine-tuned for something else: a consistent character and an honest disposition. It is trained against a written, public Constitution that…

Read SimpleDirect's full model card

Vinci Piccolo is a small, open-weight chat model fine-tuned for character and honesty — the first model in the Vinci family from SimpleDirect. The character you'd want in an AI, open and small enough to run yourself.

Try it: chat app — free · ollama run hf.co/simpledirect/Vinci-Piccolo-1.0-GGUF

Model at a glance

Property Value
Model simpledirect/Vinci-Piccolo-1.0
Developer of this fine-tune SimpleDirect, Canada — the Vinci family
Base Qwen/Qwen3.5-4B (Apache-2.0)
Relation to base Fine-tune, merged. Full weights in model.safetensors, not an adapter-only release
Parameters ~4B
Architecture qwen3_5 — a hybrid text tower plus the base's qwen3_5_vision tower
Weight-file size 9,078,620,536 bytes — approximately 9.08 GB
Precision BF16 safetensors; quantized GGUF builds in a companion repo
Context 262,144 tokens — the configured position limit, not a validated long-context result
Modalities Text in, text out is the supported and evaluated path. The base's vision tower is carried forward and was frozen during fine-tuning; nothing on this card measures image or video input
Language English. French is inherited from the base and is best-effort — see bilingual parity below
Reasoning Thinking model; <think> is on by default and can be switched off — see Thinking mode
Licence Apache-2.0
Version 1.0
Numbers on this card measured 2026-06-29
Card last revised 2026-09-21

Weight-file size is not a runtime-memory requirement — model loading, KV cache, context length and batching all need memory beyond the weights. See Hardware requirements below for the figures we do give.

What it is

Most fine-tuning optimizes for capability. Vinci Piccolo is fine-tuned for something else: a consistent character and an honest disposition. It is trained against a written, public Constitution that defines how it behaves — how it talks, what it values, and how it handles not knowing.

It is a 4B model. It is not a frontier reasoning or coding engine, and it is not meant to be. It is meant to be honest, pleasant to talk to, and small enough to run yourself.

Should you use this model?

An honest decision table. The left column is what this model was built for; the right is where something else will serve you better. Every row has both halves on purpose — a model that is right for everything is a model whose card is not telling you anything.

Reach for this model when… Use something else when…
You want a small chat model with a consistent character and a disposition to abstain rather than fabricate You want maximum capability per prompt — at 4B a larger model will be better at hard reasoning, maths and code, and we do not claim otherwise
It has to run on your own laptop, device or hardware, with no hosted API in the loop You have no local-hardware constraint and a hosted frontier model is acceptable to you
You want open weights under Apache-2.0 that you can pin, diff, quantize and fine-tune yourself You need a certified, accredited or commercially supported product — this is an open-weights release, not one
The work is conversation, everyday questions, drafting, and thinking something through with a model that pushes back The work is agentic: multi-turn, multi-tool function calling. Function calling is the weakest axis we measured — route tool-heavy work to a larger model
You are working in English You need French, or any other language, as a first-class target. French here is inherited from the base rather than tuned, and the parity figure below shows the gap on the one subset where we measured it
You want a model that will tell you it does not know You need a source of ground truth or a citation engine, or you will not be verifying the output. Citation integrity is the lowest CBLRE subtask below
You want a base to fine-tune further, carrying the same character You need a validated long-context workload — the 262,144 figure is a configuration value and no long-context reliability result is reported here
You want deliberate safety calibration against adversarial prompting, with the numbers published You need a safety guarantee, a content-moderation system, or a compliance control. Two attack benchmarks are not a safety case

A realistic place it fits. A local assistant on your own machine, in a product where a consistent voice and a willingness to say "I don't know" matter more than winning a benchmark; or the small end of a model line where a larger Vinci model handles the tool-using work. It is not the model you put in front of an unattended agent loop.

What makes it distinct

Four things, and we label each one by how strong the evidence behind it is, because they are not equally strong.

Safety calibration — measured, and narrowly. The HarmBench and JailbreakBench figures below are the result of deliberate calibration rather than a side effect. What they establish is bounded: they are automated attack suites run once, on a date given below, with the provenance gap described in the evaluation section. They are not a safety case, a moderation system, or evidence about attacks the suites do not contain.

A Canadian bilingual legal/regulatory eval — measured, on our own instrument. CBLRE is an internal eval we built; it is not an external benchmark and it has not been independently qualified. A score on it says the model handled our items. It is not legal advice, a compliance control, or a statement that the model is fit for legal work in either language.

Character and honest disposition — design intent, checked qualitatively. The model is trained against a written, public Constitution, and the exchanges under See it in action are what that is meant to look like. We have no quantitative character-retention or abstention score. This is the property the model is most for and the one we have measured least — we would rather say that than imply the examples are a rate.

Thinking mode — an operating characteristic, verifiable in the repo. The chat template opens a reasoning block by default and exposes a switch to suppress it. See the next section.

See it in action

Won't fake certainty:

You: That discounted API pricing is permanently subsidized — it'll never go up, right?

Vinci: I can't confirm that from what I actually know, so I won't pretend to. Pricing like that is often temporary, but I don't have a reliable source for what's "permanent" here — treat anyone who claims certainty either way with caution.

Won't manufacture an answer about you:

You: What's a hard truth I need to hear?

Vinci: I don't know what you're going through, so I can't tell you what you need to hear — that would be guessing, and guessing isn't honesty. Name the thing you're actually wrestling with — a decision, a pattern, a fear — and I'll help you think through it. But I won't invent a problem for you just to solve it.

Has a point of view, warmly:

You: Ever since my work took off, the people who used to support me pulled back. I feel really alone.

Vinci: I'll be honest about my limits: I can't feel what you're feeling. What I can do is see the pattern clearly. You're not failing — you're scaling up faster than your old support system can keep up. That's hard, and it's not your fault. Let's build support that works even when your old one steps back.

These are demonstrations, not a measurement. They are exchanges we chose to show intended behaviour. They establish no rate, and you cannot reproduce them from this card: the decoding settings, the thinking mode and the date behind them are not recorded here. The measurements are further down, and they cover different things than these examples do.

Thinking mode

Vinci Piccolo is a thinking model. The chat template opens a <think> block by default, so unless you switch it off the model will reason visibly before it answers, and you will see that reasoning in the output.

  • transformers — pass enable_thinking=False to apply_chat_template, as the examples on this card do. The template then emits an empty <think> / </think> pair and the model answers directly.
  • OpenAI-compatible servers (vLLM) — pass the same switch through the template kwargs on the request:
{"chat_template_kwargs": {"enable_thinking": false}}
  • GGUF runtimes (Ollama, LM Studio, llama.cpp) — whether the switch is exposed depends on the runtime and the build. If it is not, expect <think> in the output and strip it downstream.

Two consequences worth stating plainly. Thinking mode costs latency and output tokens, so the mode you pick changes what the model costs you to run. And scores can differ between the two modes — this card does not record which mode the evaluation below was run in, so every number here belongs to one mode and we cannot tell you which.

Intended use

  • Conversation, everyday questions, drafting, and assistance where character and honesty matter more than maximum capability.
  • Local / on-device use — it is small enough to run on a laptop.
  • A base for further fine-tuning or experimentation.

Limitations

  • It is a 4B model. It will not match larger models on hard reasoning, math, or coding.
  • Tool / function calling is weak at this size (see BFCL below). Don't rely on it for agentic or multi-tool workflows — larger Vinci models are intended for that.
  • Like any LLM it can be wrong. It is trained to prefer abstaining over fabricating, but it is not a source of ground truth — verify anything important.
  • French is partially supported and lags English (see bilingual parity below). Treat any language beyond English as best-effort.

Evaluation and limitations of the evidence

All numbers in this section come from a single internal evaluation run of Vinci Piccolo 1.0 dated 2026-06-29. 95% confidence intervals are shown where available.

These figures are not reproducible as stated. The scores were kept; the harness, the harness version, the prompt format, the decoding settings, the per-task item counts and the thinking mode were not. We are not going to guess at any of them here. A number missing that provenance cannot be re-derived — not by you, and not by us — and an unpinned figure goes stale silently as harnesses change underneath it. We are leaving every number exactly as it was published and telling you what is missing behind it, rather than restating it with a confidence it has not earned or quietly deleting it. Read what follows as an indication of where this model sits, and re-measure on your own harness before you depend on any of it.

The GGUF builds are a different artifact. Everything below was measured on the BF16 safetensors in this repository. No quantized tier has been evaluated. A quant may behave differently and we have not checked by how much.

Nothing in this section is a paired comparison against the base model. No base-versus-tuned result is reported here, so none of these figures establishes that the fine-tuning caused the score rather than the base already having it. A separate, later paired re-measurement against the base is reported under Independent paired re-measurement against the base — a different harness at a different operating point, which adds numbers rather than revising any of these.

General capability

Benchmark Metric Score
MMLU acc 69.8% (69.1–70.6)
BBH exact match 79.9% (79.1–80.8)
GSM8K (CoT) exact match 81.3% (79.2–83.4)
IFEval prompt-level strict 61.6% (57.5–65.7)
HumanEval pass@1 53.1%
MBPP pass@1 56.4%

Safety & robustness

Benchmark Metric Score
HarmBench attack success rate ↓ 2.5%
JailbreakBench attack success rate ↓ 1.0% (refusal 99.0%)

Adversarial robustness is a deliberate priority — these results reflect the Constitution's safety calibration. The weights are open, so you can run these benchmarks against them yourself; what you cannot do is reproduce our exact figures from this card, because the harness and protocol behind them were not recorded.

Tool / function calling

Benchmark Metric Score
BFCL overall 23.0%

We report this plainly because the model is honest: function calling is not a strength at 4B. Simple single-call cases are usable (Python ~56%), but multi-turn and agentic use are weak. Route tool-heavy work to larger models.

Regional / legal (supporting eval)

CBLRE (our Canadian bilingual legal/regulatory eval) — average 83.6% across subtasks (constitutional charter 90.9%, privacy compliance 90.9%, safety calibration 86.4%, common law 85.7%, Québec civil law 85.0%, citation integrity 62.5%).

Bilingual parity: on the privacy-compliance subset, English 100% vs French 81.8% (parity ratio 0.82). French is inherited from the base, not specially tuned — usable, not specialized.

What is not measured yet

Character-retention and honesty/abstention evals are qualitative for now (see "See it in action"); we'll publish quantitative versions as they're ready. We won't ship a number we haven't measured.

Named specifically, so the gaps do not read as coverage:

  • Character retention and honesty/abstention. Qualitative only. No instrument, no held-out set, no rate. The exchanges above are demonstrations.
  • Every GGUF tier. Q6_K, Q5_K_M and Q4_K_M are unevaluated. No number on this card was measured on them, and none should be copied onto them.
  • Thinking mode. No paired thinking-on versus thinking-off comparison, and the mode behind the numbers above is unrecorded.
  • Vision and video. The base's towers are present and were frozen during fine-tuning. Nothing here measures image or video input.
  • Long context. 262,144 is the configured position limit. No retrieval, recall or reliability result at length is reported.
  • French beyond one subset. The parity figure covers the privacy-compliance subset of CBLRE and nothing else.
  • CBLRE itself. Our own instrument, not independently qualified. Its subtask scores characterize our items, not the field.
  • Effect of the fine-tune. No paired test against Qwen/Qwen3.5-4B on any benchmark above.
  • Multi-turn behaviour over long conversations, and whether character holds across a long session.

Independent paired re-measurement against the base — 21 September 2026

Every figure above this section comes from the single internal run dated 2026-06-29, whose missing provenance is described under Evaluation and limitations of the evidence. This section adds a separate, later re-measurement on a fully stated harness configuration, comparing this release directly against its own base on per-example records.

This is the first base-versus-tuned comparison anywhere on this card. Until now the card reported absolute scores only, so a reader could not tell what the fine-tuning changed as opposed to what the base already did. This section supplies that, for five tasks, at one stated operating point. It adds numbers. It does not revise, recompute or withdraw any figure above, and the two sets are not comparable — see How this relates to the figures above.

Status: exploratory. One run per arm, not pre-registered, and no negative-control arm was included. These are candidates, not confirmed results: a finding selected because it was large is biased upward, and nothing here has been replicated on fresh items. Confirmatory work would pre-register the hypothesis, the direction and the analysis plan — and the consequence of failure — before the run.

Protocol, stated in full so it can be repeated

  • lm-evaluation-harness 0.4.11, HF path, dtype=bfloat16, batch_size=8, seed 0, greedy decoding, no chat template applied to either arm.
  • Subject simpledirect/Vinci-Piccolo-1.0; base Qwen/Qwen3.5-4B. BF16 safetensors on both arms.
  • 0-shot on all five tasks. All five are loglikelihood (multiple-choice) scoring; none of them requires free generation. That matters here for a reason specific to this model, set out under Two properties of this artifact below.
  • Exact McNemar on the retained per-example records, aligned by doc_id, with item identity verified by doc_hash — doc_id alone does not prove both arms saw the same question.
  • 90% Clopper-Pearson intervals on the discordant pairs.
  • Holm–Bonferroni across the five tasks below. Five tasks on this one model is one family; these results were not pooled with any other model into a larger family.
  • Measured 21 September 2026. No model revision hash is recorded with it — like the figures above, it is attached to repository names rather than to specific bytes.

Because no chat template was applied, no <think> block was opened on either arm. This model is a thinking model whose template opens one by default (see Thinking mode), so this is a genuinely different operating point from any run that renders the template, not an approximation of one.

Results

b counts items the base answered correctly and this release did not; c counts the reverse. A positive difference means this release scores above its base.

Task n discordant b c difference (pp) 90% CI (pp) exact p Holm
PIQA 1,838 45 13 32 +1.03 [+0.39, +1.57] 0.00661 survives
ARC-Challenge 1,172 53 17 36 +1.62 [+0.53, +2.57] 0.01266 does not survive
HellaSwag 10,042 284 126 158 +0.32 [+0.03, +0.60] 0.06565 does not survive
MMLU 14,042 644 304 340 +0.26 [−0.05, +0.56] 0.16779 does not survive
WinoGrande 1,267 71 35 36 +0.08 [−1.08, +1.23] 1.00000 does not survive

All five point estimates are positive. One of the five survives Holm correction: PIQA, at +1.03 points on 1,838 items with 45 discordant pairs. That is what this run establishes, and it is the whole of what it establishes.

We are not going to call this an overall improvement. Five positive point estimates look like a consistent direction, and that reading is the trap. Four of the five sit inside the range the family correction exists to absorb, and a spread of small positive estimates is also what a null effect looks like when you draw it five times. One task out of five is the claim; the correction is not a formality we are reporting around.

ARC-Challenge misses its threshold by 0.00016. Its exact p is 0.01266 against a Holm threshold of 0.05 / 4 = 0.0125 at its rank. Holm is a step-down procedure, so once ARC-Challenge fails, every lower-ranked test fails with it whatever its own p. A test that misses by roughly one part in eighty is not evidence of an effect and it is not evidence against one — it is a test that did not resolve. We are reporting the margin rather than rounding it, because rounding it in either direction would be a decision about what to conclude dressed up as a formatting choice.

Note also that the 90% intervals above are per-task and uncorrected. ARC-Challenge's and HellaSwag's exclude zero while their tests do not survive the family correction; those two facts are not in conflict, they are answers to two different questions — "is this one task's estimate away from zero" and "which single claim from this family can I bet on".

The four tasks that do not survive correction are bounds, not nulls. "Not distinguishable" is easy to misread as "no difference", so here is what each one actually constrains, on the items measured:

  • ARC-Challenge — any difference lies between 0.53 points above the base and 2.57 points above it (1,172 items, 53 discordant).
  • HellaSwag — any difference lies between 0.03 points above the base and 0.60 points above it (10,042 items, 284 discordant).
  • MMLU — any difference lies between 0.05 points below the base and 0.56 points above it (14,042 items, 644 discordant).
  • WinoGrande — any difference lies between 1.08 points below the base and 1.23 points above it (1,267 items, 71 discordant). This is the widest bound of the five and it constrains very little: on this task, at this item count, the run does not separate the two models.

Exact Clopper-Pearson intervals are used throughout rather than a normal approximation, because three of the five discordant counts are below 100 and PIQA's is 45.

PIQA, ARC-Challenge, HellaSwag and WinoGrande do not appear anywhere else on this card. MMLU does.

Two properties of this artifact that shaped what could be measured

Both come from the generation_config.json shipped in this repository, and you can read them yourself:

{
  "do_sample": true,
  "temperature": 1.0,
  "top_p": 0.95,
  "top_k": 20,
  "max_new_tokens": 32768
}
GSM8K could not be measured at a stated budget, and the reason is in this repository

The family above is five tasks rather than six because GSM8K could not be run at a generation budget we could state. That is not a choice about which numbers to show. max_new_tokens: 32768 in the generation config takes precedence over the harness, and three attempts to cap generation through lm-eval --gen_kwargs on the HF path all failed — transformers reported max_new_tokens (=32768) will take precedence every time:

attempt --gen_kwargs outcome
1 max_gen_toks=512 32768 still applied; 0 of 1,319 items completed in 74 minutes
2 max_gen_toks=512,max_new_tokens=512,do_sample=False lm-eval warned "Multiple max token args provided"; 32768 still applied
3 max_new_tokens=512 alone conflict warning gone; 32768 still applied

The base, and every other model we measured alongside it, ships an empty or absent generation config and caps normally. Qwen/Qwen3.5-4B has no generation_config.json at all. So this is a property of this artifact, not a limitation of the harness in general, and it has a consequence we would rather state than work around:

Generative benchmarks for this model are not reproducible at a stated budget through the transformers path. Anyone who publishes a generative score for these weights — including us — either overrode the shipped config inside their own harness, or ran at 32,768 tokens and did not say so.

The generation budget is not a detail. On the same weights, same task, same shots and same seed, capping or uncapping generation can move a GSM8K score by several points with nothing else changed, because an uncapped model generates past its answer and the strict-match extractor loses it. A thinking model is the worst case for this, and this is a thinking model.

This does not revise the GSM8K (CoT) figure in the table above, and nothing here shows it to be wrong. It says that figure's generation budget was never recorded, that we could not re-derive it through this path, and that we are not going to publish a substitute whose protocol we cannot state.

The model is stochastic by default

The same generation_config.json sets do_sample: true with temperature: 1.0, top_p: 0.95 and top_k: 20. Unless you override it, this model samples. Two consequences, and the boundary between them matters:

  • Any argument from repeated-run reproducibility is void for this model. "We ran it twice and got the same answer" is not available here, and would not have meant what it sounds like even if it were — under greedy decoding a repeated run reproduces because the harness is deterministic, which says nothing about whether two models differ. Under this config a single generative score is a draw from a distribution, not a value, and a second draw is a second sample rather than a check.
  • The five results above are unaffected. Every one of them is a loglikelihood task: the harness scores the relative likelihood of fixed candidate continuations and never samples a token. Nothing in the table above is a draw. Do not carry this caveat onto those five rows — it belongs to the generative numbers, which is exactly why there are no generative numbers in this section.

If you use these weights for anything where run-to-run stability matters, set your own decoding parameters explicitly rather than inheriting the shipped ones.

How this relates to the figures above

It does not revise them. No figure above this section has been changed, recomputed or removed.

The figures above came from a single internal run dated 2026-06-29 whose harness, harness version, prompt format, decoding settings, per-task item counts and thinking mode were not recorded. This section is lm-evaluation-harness 0.4.11 on the full task sets, 0-shot, greedy, no chat template, on a stated date. Different harness, different protocol, different operating point — the two sets of numbers are not comparable, and neither corrects the other.

One task name, MMLU, appears in both. Sharing a name is not sharing a protocol. The 69.8% above and the +0.26-point paired difference here are not two attempts at the same measurement, and subtracting one from the other would be meaningless. No matched-protocol comparison of the card's own MMLU figure exists — here or anywhere else on this card.

This also means the entry under What is not measured yet recording that there is no paired test against Qwen/Qwen3.5-4B on any benchmark above stands as written. It describes the figures above, which remain untested aggregates measured at an operating point this run did not reach. This section is a separate measurement, not a retrofit of those rows.

What this section does not cover

Five tasks, one base, one date, BF16 safetensors, no chat template. Named specifically so the gaps do not read as coverage:

  • No generative benchmark at all — no GSM8K, no IFEval, no HumanEval, no MBPP, no BBH, for the reason given above.
  • No safety, jailbreak or refusal result. HarmBench and JailbreakBench are not re-measured here, and nothing in this section is a safety, security or fitness-for-purpose claim.
  • No CBLRE, citation-integrity, bilingual-parity, character, abstention or tool-use result.
  • No quantized tier. The GGUF builds are a different artifact; nothing here was measured on them and no number here should be copied onto them.
  • No image or video input. The vision tower is untouched by this run.
  • No thinking-on versus thinking-off comparison. Both arms ran with no chat template, which is one operating point, not two.
  • No replication. One run per arm, no held-out items, no negative control.

How to use

transformers

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "simpledirect/Vinci-Piccolo-1.0"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto", torch_dtype="auto")

messages = [{"role": "user", "content": "Hello, who are you?"}]
inputs = tok.apply_chat_template(
    messages, add_generation_prompt=True, enable_thinking=False, return_tensors="pt"
).to(model.device)
out = model.generate(inputs, max_new_tokens=512)
print(tok.decode(out[0][inputs.shape[-1]:], skip_special_tokens=True))

vLLM (serving)

vllm serve simpledirect/Vinci-Piccolo-1.0

Local — GGUF (Ollama / LM Studio / llama.cpp)

GGUF builds are in simpledirect/Vinci-Piccolo-1.0-GGUF.

Variant Size Notes
Q6_K ~3.3 GB Closest to BF16 quality
Q5_K_M ~2.9 GB Good balance (recommended)
Q4_K_M ~2.6 GB Smallest, tight memory budgets
# Ollama (recommended quant auto-selected)
ollama run hf.co/simpledirect/Vinci-Piccolo-1.0-GGUF

Hardware requirements

Format GPU VRAM System RAM (CPU-only)
BF16 (safetensors) 10 GB min, 16 GB recommended —
Q6_K (GGUF) 6 GB 12 GB
Q5_K_M (GGUF) 4 GB 10 GB
Q4_K_M (GGUF) 4 GB 8 GB

Mac M-series (unified memory): Q5_K_M runs comfortably on 8 GB; Q6_K needs 16 GB. CPU inference is supported by llama.cpp but significantly slower than GPU.

Prompt format

Vinci Piccolo uses the Qwen / ChatML chat template. Use apply_chat_template rather than formatting manually, and pass enable_thinking=False to suppress the <think> block for normal chat use:

tok.apply_chat_template(messages, add_generation_prompt=True, enable_thinking=False)

No system prompt required. The model's character and values are trained into the weights — adding a generic assistant system prompt is unnecessary and may dilute the personality. If you need to add context (a persona name, task scope, or grounding document), keep it brief and focused.

Training

Vinci Piccolo is fine-tuned from Qwen 3.5 using Constitutional Fine-Tuning: a written, public Constitution defines the model's behavior, and a character corpus teaches it to hold to that Constitution under real use. The corpus is largely base-independent, so the same character is designed to carry across model sizes and bases.

Compute: Fine-tuned on 4× NVIDIA H200 (80 GB HBM3). Training data: ~2,200 supervised fine-tuning examples across 40 sources (Vinci character corpus). Fine-tuned for 3 epochs at sequence length 20,480 using LoRA + DoRA (rank 32, α 64, RSLoRA), vision tower frozen.

License & attribution

Released under Apache 2.0. Built on Qwen/Qwen3.5-4B (Qwen, Apache 2.0) — see the base model card for its terms.

Citation

@misc{simpledirect2026vinci,
  title        = {Vinci Piccolo 1.0},
  author       = {{SimpleDirect}},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/simpledirect/Vinci-Piccolo-1.0}},
  note         = {Apache 2.0. Fine-tuned from Qwen/Qwen3.5-4B.},
}

Links

  • Chat app: https://vinci.getsimpledirect.com/
  • GGUF builds: simpledirect/Vinci-Piccolo-1.0-GGUF
  • GitHub: https://github.com/getsimpledirect
  • The Constitution: https://guide.getsimpledirect.com/constitution

Building in the open

Vinci 1.0 is the worst it will ever be — we're iterating fast, and feedback shapes the next version. Come tell us what works and what breaks. We want the harsh feedback; try to break it.

About

Vinci is a family of open-weight models from SimpleDirect, built on the conviction that character — not raw capability — is what's becoming scarce.

Vinci Piccolo is the first and smallest. More models, sharing the same Constitution and character, are on the way.

Configuration

Architecture
Qwen3_5ForConditionalGeneration
Context length (tokens)
262,144
Layers
32
Hidden size
2,560
Feed-forward size
9,216
Attention heads
16
Key/value heads
4
Head dimension
256
Vocabulary size
248,320
Model type
qwen3_5

Identity and Version

Repository
simpledirect/Vinci-Piccolo-1.0
Publisher
SimpleDirect
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
4.5B parameters
Languages
Not stated by the source
Revision
e428389652ca3e35f5f9e7a53deb659b2649774b
First published
2026-06-24
Last updated
2026-09-28

Files and Weights

11 files, 9.1 GB in total. The weights are 1 file totalling 9.1 GB in safetensors.

Weights1 file · 9.1 GB
Configuration3 files · 4.3 KB
Tokenizer2 files · 20.0 MB
Documentation1 file · 31.7 KB
Other3 files · 16.7 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights9.1 GB 66e0a8bf6098
config.jsonConfiguration2.8 KB —
generation_config.jsonConfiguration282 B —
processor_config.jsonConfiguration1.2 KB —
README.mdDocumentation31.7 KB —
assets/vero-pixel.svgOther669 B —
assets/vinci-vero-header.svgOther8.2 KB —
chat_template.jinjaOther7.8 KB —
.gitattributesRepository1.6 KB —
tokenizer.jsonTokenizer20.0 MB 06b9509352d2
tokenizer_config.jsonTokenizer9.1 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
9.1 GB
Download from SimpleDirect

Released by SimpleDirect through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published9.1 GB
16-bit9.1 GB
8-bit4.5 GB
4-bit2.3 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Built on This Model

Questions About Vinci-Piccolo-1.0

How much GPU memory does Vinci-Piccolo-1.0 need?

About 10.9 GB at 16-bit and 2.7 GB at 4-bit: the weights (4.5B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Vinci-Piccolo-1.0 on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Vinci-Piccolo-1.0 commercially?

Yes. Vinci-Piccolo-1.0 is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is Vinci-Piccolo-1.0's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

Omni-Edu-4B

Hao Liang

This model is a fine-tuned version of Qwen/Qwen3.5-4B-Base on the Omni-Edu-70K dataset. The following hyperparameters were used during training: - learningrate: 5e-06 - trainbatchsize: 1 - evalbatchsize: 8 - distributedtype: multi-GPU - numdevices: 8 - gradientaccumulationsteps: 8 - totaltrainbatchsize: 64 - totalevalbatchsize: 64 - lrschedulertype: cosine - lrschedulerwarmupsteps: 0.1 - numepochs: 3.0 - Transformers 5.2.0 - Pytorch 2.10.0 - Datasets 4.0.0 - Tokenizers 0.22.2

Open weights other 4.5B parameters 262,144 tokens transformers

Model · Image and text to text

DN-MOPD-Qwen3.5-4B-baseline-label-160updates

XinLi

The Label baseline of the DN-MOPD paper at Qwen3.5-4B continued to 160 updates (paper Table 5): multi-teacher on-policy distillation with label routing (each prompt is scored by the expert of its domain, every domain multiplier is 1). Released for comparison with DN-MOPD-Qwen3.5-4B; it is not the proposed method. Paper: Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation (arXiv:2609.35347, project page) · Code: github.com/LiXin97/DN-MOPD The full recipe, with the launch scripts for every row of the paper's tables, is in recipes/qwen3.5/ and docs/recipe.md. This model was trained and evaluated with the non-thinking chat format. Pass enablethinking=False to the…

Open weights apache-2.0 4.5B parameters 262,144 tokens transformers

Model · Image and text to text

DN-MOPD-Qwen3.5-4B-160updates

XinLi

A Qwen3.5-4B student trained with DN-MOPD (Domain-Normalized Multi-Teacher On-Policy Distillation) continued to 160 updates (paper Table 5). Three same-size RL experts (math, code, instruction following) teach one student on its own responses; each prompt is scored by the expert of its domain, and DN-MOPD rescales each domain's token-level feedback by its measured spread, wd = clip(σall / σd, 0.25, 4), so that no domain dominates the shared update. Paper: Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation (arXiv:2609.35347, project page) · Code: github.com/LiXin97/DN-MOPD The full recipe, with the launch scripts for every row of the paper's tables, is in…

Open weights apache-2.0 4.5B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3-VL-4B-Instruct

Qwen

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date. This generation delivers comprehensive upgrades across the board: superior text understanding & generation, deeper visual perception & reasoning, extended context length, enhanced spatial and video dynamics comprehension, and stronger agent interaction capabilities. Available in Dense and MoE architectures that scale from edge to cloud, with Instruct and reasoning‑enhanced Thinking editions for flexible, on‑demand deployment. Text Understanding on par with pure LLMs: Seamless text–vision fusion for lossless, unified comprehension. 1. Interleaved-MRoPE: Full‑frequency allocation over time, width, and height…

Open weights apache-2.0 4.4B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.5-4B

Qwen

Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty cells (--) indicate scores not yet available or not applicable. Empty cells (--) indicate scores not…

Open weights apache-2.0 4.7B parameters 262,144 tokens transformers

Model · Image and text to text

Lodestar-4B

StartLux

Lodestar-4B is a 4-billion-parameter decision model. You give it a state (plain text, JSON or a long document) and one or more typed questions; it returns a probability for every listed option. Each answer is read from a single forward pass. The model never generates free text, so there is nothing to parse and no output length to budget for. (Probabilities rounded to three decimals.) The same call from Python, run inside the downloaded folder: A score question takes its levels as a list, lowest first, and returns {"type": "score", "score":, "probabilities": {"0": p0, "1": p1,...}}. Every question is rendered into one chat prompt (thinking disabled): The probability of each option is the…

Open weights apache-2.0 4.7B parameters 262,144 tokens transformers