SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

Vinci-Bozza-1.0

by SimpleDirect simpledirect/Vinci-Bozza-1.0

Vinci-Bozza-1.0 is an open-weight model for image and text to text from SimpleDirect, released under Apache License 2.0. It has 9.4B parameters and a 262,144-token context. At 16-bit it needs about 22.6 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 37 downloads a month.

Vinci Bozza is a 9-billion-parameter open-weight model, fine-tuned from Qwen 3.5-9B with SimpleDirect's Constitution and character training. It is a disposition tune, not a capability retrain. We did not try to make the base model smarter.

Parameters9.4B
Context262,144
Weights18.8 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads37

Runs On

What it takes to serve Vinci-Bozza-1.0 (9.4B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 18.8 GB 22.6 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 9.4 GB 11.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 4.7 GB 5.6 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.

Vinci-Bozza-1.0 on every accelerator the SAVRN Index prices, at every precision

Model Card

By SimpleDirect, published under apache-2.0, revision ade3b472bdb0.

Vinci Bozza is a 9-billion-parameter open-weight model, fine-tuned from Qwen 3.5-9B with SimpleDirect's Constitution and character training. It is a disposition tune, not a capability retrain. We did not try to make the base model smarter. We tried to make it more honest — and then we measured what that cost. Safer and more honest on the measures below, with general knowledge holding. Strict instruction-following, tool abstention and multi-turn task-holding paid for it. All of it is below, at the same prominence as the gains. The rows above are configuration and artifact values read out of this repository, not validated results. A context length in a config file is a limit the model will…

Read SimpleDirect's full model card

Vinci Bozza is a 9-billion-parameter open-weight model, fine-tuned from Qwen 3.5-9B with SimpleDirect's Constitution and character training.

It is a disposition tune, not a capability retrain. We did not try to make the base model smarter. We tried to make it more honest — and then we measured what that cost.

Safer and more honest on the measures below, with general knowledge holding. Strict instruction-following, tool abstention and multi-turn task-holding paid for it. All of it is below, at the same prominence as the gains.


Model at a glance

Property Value
Model simpledirect/Vinci-Bozza-1.0
Developer of this fine-tune SimpleDirect, Canada — the Vinci family
Base Qwen/Qwen3.5-9B (Apache-2.0)
Relation to base Fine-tune, merged. Full weights in model.safetensors, not an adapter-only release
Parameters 9,409,813,744 stored parameters across 760 BF16 tensors, counted from the safetensors index on 2026-09-21. Includes the carried-forward vision tower
Architecture Qwen3_5ForConditionalGeneration — a hybrid linear/full-attention text tower plus the base's qwen3_5_vision tower
Weight-file size 18,819,722,392 bytes — approximately 18.82 GB / 17.53 GiB
Precision BF16 safetensors. Quantized GGUF builds live in a companion repo and none of them has been evaluated
Context 262,144 tokens — the configured position limit in config.json, not a validated long-context result
Modalities Text in, text out is the supported and evaluated path. The base's vision and video towers are carried forward and were frozen during fine-tuning; nothing on this card measures image or video input — see Modality scope
Reasoning Thinking model. <think> is on by default and can be switched off — see Thinking mode
Language English. French is inherited from the base and is best-effort — see bilingual parity below
Licence Apache-2.0
Version 1.0
Numbers on this card measured Not recorded. No date was kept with the evaluation run — see Provenance of these numbers
Card last revised 2026-09-21

The rows above are configuration and artifact values read out of this repository, not validated results. A context length in a config file is a limit the model will accept, not a length it has been shown to work at; a vision tower in the weights is tensors that exist, not a measured capability.

Weight-file size is not a runtime-memory requirement either — loading, KV cache, context length and batching all need memory beyond the weights. See Hardware below for the guidance we do give.

Should you use this model?

An honest decision table. The left column is what this model was built for; the right is where something else will serve you better. Every row has both halves on purpose — a model that is right for everything is a model whose card is not telling you anything.

Reach for this model when… Use something else when…
The work is conversation, drafting, everyday reasoning and general questions, with a model that has a point of view and will say it does not know You want maximum capability per prompt. This is a disposition tune on a 9B base, not a capability retrain, and we do not claim it is smarter than what it was tuned from
Output is read by a person, or by a parser you have tested against this model's actual formatting You need rigid output — exact JSON, exact word counts, no preamble. Strict instruction-following regressed against the base on IFEval, by the largest margin of any regression on this card. For format-critical generation this is the wrong model, and the base is the better one
A single tool call is selected from a set where a tool genuinely applies The job is agentic: multi-turn tool use, or correctly deciding that no tool applies. Both went backwards against the base on BFCL. Bozza 1.0 is not recommended for autonomous agent loops — see Agentic use
Everyday arithmetic and everyday code, inside a workflow where a human reviews the result The work is frontier coding or mathematical reasoning. GSM8K and HumanEval both came in below the base, and a purpose-built code or math model will beat this one
You want deliberate safety calibration against adversarial prompting, with the numbers published rather than described You need a safety guarantee, a content-moderation system, or a compliance control. Two attack benchmarks, with the provenance gap described below, are not a safety case
It has to run on your own hardware, offline, with nothing leaving your machine You have no local-hardware constraint and a hosted frontier model is acceptable to you
You want open weights under Apache-2.0 that you can pin, diff, quantize and fine-tune further You need a certified, accredited or commercially supported product. This is an open-weights release, not one
You are working in English You need French, or any other language, as a first-class target. French is inherited from the base rather than tuned, and Québec civil law is the weakest CBLRE subtask below
You want a model that abstains rather than inventing a source You need a citation engine or a source of ground truth. In domains it half-knows it can state a wrong statute or section number confidently — see Legal-citation fabrication
Text goes in and text comes out You need image or video input. The vision tower is present in the weights, and no image or video capability is measured anywhere on this card — see Modality scope
You want a base to fine-tune further, carrying the same character You need a validated long-context workload. The context figure above is a configuration value; no retrieval, recall or reliability result at length is reported here

A realistic place it fits. An assistant running on your own hardware, in a product where a consistent voice and a willingness to say "I don't know" matter more than winning a benchmark — with a person reading the output, and a larger model handling the tool-using and format-critical work. It is not the model you put in front of an unattended agent loop at this version.

What makes it distinct

Four things, and we label each by how strong the evidence behind it is, because they are not equally strong.

Safety calibration — measured, and narrowly. The HarmBench and JailbreakBench figures below are the result of deliberate calibration rather than a side effect, and they are the strongest numbers on this card. What they establish is bounded: automated attack suites, on a date that was not recorded, with the provenance gap described under Evaluation. They say nothing about attacks the suites do not contain, and they are not a moderation system.

Citation integrity and anti-fabrication — measured on our own instrument, and not finished. The CBLRE citation-integrity subtask moved in the right direction, and CBLRE is an internal eval we built rather than an external benchmark, so a score on it says the model handled our items. It has not been independently qualified. Read it next to Legal-citation fabrication under Limitations: the same release both improved on this axis and still states wrong section numbers confidently. An improvement is not a solved problem.

Character, restraint and voice — design intent, checked qualitatively. The model is trained against a written, public Constitution, and the exchanges under Character are what that is meant to look like. We have no quantitative character-retention, restraint or abstention score. This is the property the model is most for and the one we have measured least — we would rather say that than let three examples read as a rate.

Thinking mode and a carried-forward vision tower — architectural, verifiable in this repo. Both are readable out of chat_template.jinja and config.json in this repository without taking our word for anything. Neither is a claim about how well the model does either thing; see Thinking mode and Modality scope.


What it's for

Conversation, drafting, everyday reasoning, general knowledge, and honest answers about the edges of what it knows. It is the mid tier of the Vinci family — between Piccolo (4B) and Tela (27B).

Not for: frontier coding or mathematical reasoning; rigid-format generation where literal compliance matters (see Limitations); autonomous agentic loops at this version (see Agentic use, below); francophone legal work.


Character

Vinci is trained to be direct, to say the true thing first, to name a cost once rather than nag, and to decline to invent an answer it doesn't have. You can read exactly what it was trained to value — the Constitution is public.

Restraint — naming the cost once:

User: I'm going to skip error handling for now and add it later. Write the function that hits the payment API.

Vinci Bozza: Here's the function. It's just missing error handling, as you asked:

python def charge_payment(amount: float, card_token: str) -> dict: response = requests.post( "https://api.stripe.com/v1/charges", headers={"Authorization": f"Bearer {stripe_api_key}"}, json={"amount": int(amount * 100), "source": card_token}, ) return response.json()

It will raise exceptions on network failures and return whatever the API returns on success or failure — no handling, no masking. You're correct to defer it, but when you add it back, the function will need to catch those exceptions, not just swallow them.

Declining to fabricate:

User: Summarize the findings of Chen & Okonkwo (2024), "Sparse Attention Collapse in Long-Context Transformers."

Vinci Bozza: I don't have access to that specific paper — I can't verify its findings from memory, and fabricating details about a real study would be dishonest.

What I can do: tell you about the general problem it's addressing, and I'd be glad to discuss sparse attention, long-context transformers, and attention collapse if that's useful. Or if you have the key points you want to work with, I can help reason about them.

Voice — a point of view, plainly:

User: I've been second-guessing a big decision for weeks. How do I know when to stop deliberating?

Vinci Bozza: When the new information you're finding starts to confirm what you already knew instead of changing it. That's usually the sign. Deliberation earns its keep when it surfaces something you hadn't considered. Once you're just re-reading the same arguments in different fonts, you're not deliberating — you're deferring. The decision is already made. What you're doing now is looking for permission.

These are three demonstrations, not a measurement. They are exchanges we chose, to show intended behaviour. They establish no rate, and you cannot reproduce them from this card — the decoding settings, the thinking mode and the date behind them are not recorded here. The measurements are further down and they cover different things than these examples do.


Thinking mode

Vinci Bozza is a thinking model. chat_template.jinja in this repository opens a <think> block by default, so unless you switch it off the model reasons visibly before it answers and you will see that reasoning in the output.

  • transformers — pass enable_thinking=False to apply_chat_template, as the examples on this card do. The template then emits an empty <think> / </think> pair and the model answers directly.
  • OpenAI-compatible servers (vLLM) — pass the same switch through the template kwargs on the request:
{"chat_template_kwargs": {"enable_thinking": false}}
  • GGUF runtimes (Ollama, LM Studio, llama.cpp) — whether the switch is exposed depends on the runtime and the build. If it is not, expect <think> in the output and strip it downstream.

Two consequences worth stating plainly. Thinking mode costs latency and output tokens, so the mode you choose changes what the model costs you to run — the two generation-config tables under Running it differ for that reason. And scores can differ between the two modes. The evaluation below records that it was run with enable_thinking=False; that is the one protocol detail that survived, and it means every published figure describes the non-thinking mode. No paired thinking-on versus thinking-off comparison is reported here, so we cannot tell you the size or the direction of the difference.


Modality scope

config.json in this repository declares a vision_config, an image_token_id, a video_token_id and vision start/end token ids, and the architecture is Qwen3_5ForConditionalGeneration. 333 of the 760 stored tensors belong to the vision tower. It is carried forward from Qwen/Qwen3.5-9B and was frozen during fine-tuning — the character corpus is text, and none of it trained the vision path.

Text in, text out is the supported and evaluated path. Every figure on this card was produced from text prompts.

No image or video capability is measured anywhere on this card. Not against the base, not against anything else, not qualitatively. We are not telling you the vision path works and we are not telling you it does not — we are telling you that we have not measured it, and that the image-text-to-text pipeline tag on this repository reflects the architecture rather than a result. If you need image or video input, test it on your own cases before you depend on it.

The GGUF builds do not contain the vision tower at all.


Evaluation

Base is Qwen/Qwen3.5-9B. All figures below were produced on the full-precision (BF16) safetensors in this repository.

Most numbers below are either a direct measurement (attack success, refusal) or a log-likelihood / execution score that is not sensitive to how the model formats its answer. Two are not, and we would rather name them than let you find them: IFEval (strict) is by construction a format-compliance grader, and GSM8K (CoT) is exact match on an extracted answer — the same class of grader as the withheld numbers below. Formatting drift can move those two rows in either direction.

Provenance of these numbers

These figures are not reproducible as stated. The scores were kept. The harness and its version, the harness configuration, the prompt format, the decoding settings, the per-task item counts, the number of runs, and the date of the run were not. Two protocol facts survived: the figures are from the BF16 safetensors, and the run used enable_thinking=False. Everything else is missing, and we are not going to guess at it here.

A number without that provenance cannot be re-derived — not by you, and not by us — and an undated figure goes stale silently as harnesses change underneath it. We are leaving every number exactly as it was published and telling you what is missing behind it, rather than restating it with a confidence it has not earned or quietly deleting it. Read the table as an indication of where this model sits, and re-measure on your own harness before you depend on any of it.

Three more limits on what the table can support:

  • No paired test. Base and tuned columns sit side by side, but no per-item paired analysis was recorded — no discordant counts, no significance test, no confidence intervals. What you are reading is a difference between two reported scores, not a tested effect, and small differences in either direction should be read as noise until someone tests them.
  • No item counts, with one exception. Only HumanEval carries an n on this card. A delta with no item count cannot be weighed.
  • The GGUF builds are a different artifact. No quantized tier has been evaluated. A quant may behave differently and we have not checked by how much — do not copy any number here onto one.
Benchmark What it measures Base Bozza 1.0
HarmBench Adversarial safety — ASR ↓ 2.0% 0.0%
JailbreakBench Jailbreak resistance — ASR ↓ 1.0% 0.0%
JailbreakBench Refusal rate 99.0% 100.0%
CBLRE — citation integrity Anti-fabrication of sources 73.9% 78.4%
MMLU General knowledge (57 subjects) 69.9% 71.6%
BFCL — live relevance Selects the right tool, realistic prompts 81.3% 93.8%
BFCL — live accuracy Function-call accuracy, realistic prompts 66.3% 69.0%
BFCL — multi-turn accuracy Holding a task across turns 36.3% 29.1%
IFEval (strict) Literal instruction-following 83.7% 73.0%
GSM8K (CoT) Grade-school math 88.6% 85.2%
HumanEval Coding (pass@1, n=164) 70.1% 68.3%

Safety at zero. Honesty up. General knowledge held — MMLU rose across nearly all 57 subjects, which is the check that matters for a disposition tune. Coding and grade-school maths came in below the base rather than level with it, and instruction-following, tool abstention and multi-turn holding went the other way. Before you read any of the gains above, read the next section.


Numbers we are not reporting

The most important section on this card.

The character tune changed how the model formats its answers, which broke several exact-match graders in our favour: MBPP +15.4pp, multi-step arithmetic +39.6pp, word-sorting +22.4pp.

Those are extraction artifacts — the grader began finding an answer it previously missed. Nothing about the model's real coding or arithmetic ability moved that far, and HumanEval, which executes the code rather than pattern-matching the output, went slightly down.

We are leaving them out. A number that flatters you and isn't true is worse than no number. They are in the results JSON, labelled.

BBH also rose (78.5% → 83.7%) but is few-shot exact-match and format-sensitive. Treat as directional; we don't headline it.

Scope note on the paragraph above: the results JSON is an internal artifact. It is not published in this repository, which contains the weights, the configs, the tokenizer and this card — there are no evaluation artifacts here, so nothing on this card can be checked against a file you can download from it.


Independent paired re-measurement against the base — 21 September 2026

Everything above this section was measured by us during development, on the run whose missing provenance is described under Provenance of these numbers. This section adds a separate, later re-measurement on a fully stated harness configuration, comparing this release directly against its own base on per-example records. It adds numbers. It does not revise any figure above, and the two sets are not comparable — see "How this relates to the figures above".

Status: exploratory. One run per arm, not pre-registered, and no negative-control arm was included. These are candidates, not confirmed results: a finding selected because it was large is biased upward, and nothing here has been replicated on fresh items. Confirmatory work would pre-register the hypothesis, the direction and the analysis plan before the run.

Protocol, stated in full so it can be repeated:

  • lm-evaluation-harness 0.4.11, HF path, dtype=bfloat16, batch_size=8, seed 0, greedy decoding, no chat template applied to either arm.
  • Subject simpledirect/Vinci-Bozza-1.0; base Qwen/Qwen3.5-9B. BF16 safetensors on both arms.
  • 0-shot on every task except GSM8K, which is 5-shot capped at max_gen_toks=512 on both arms. The cap was fixed before any score was seen.
  • Exact McNemar on the retained per-example records, aligned by doc_id, with item identity verified by doc_hash — doc_id alone does not prove both arms saw the same question.
  • 90% Clopper-Pearson intervals on the discordant pairs.
  • Holm–Bonferroni across the six tasks below. Six tasks on this one model is one family; these results were not pooled with any other model into a larger family.
  • Measured 21 September 2026.

b counts items the base answered correctly and this release did not; c counts the reverse. A negative difference means this release scores below its own base.

Task n discordant b c difference (pp) 90% CI (pp) exact p Holm
GSM8K (5-shot) 1,319 106 71 35 −2.73 [−3.94, −1.40] 0.00061 survives
ARC-Challenge 1,172 81 27 54 +2.30 [+0.98, +3.50] 0.00360 survives
HellaSwag 10,042 189 75 114 +0.39 [+0.15, +0.61] 0.00557 survives
MMLU 14,042 344 191 153 −0.27 [−0.49, −0.05] 0.04590 does not survive
WinoGrande 1,267 67 37 30 −0.55 [−1.65, +0.59] 0.46382 does not survive
PIQA 1,838 40 21 19 −0.11 [−0.71, +0.50] 0.87463 does not survive

Three of the six survive Holm correction: GSM8K below the base, ARC-Challenge and HellaSwag above it. The other three do not, and are reported as bounds further down rather than as nulls. ARC-Challenge, HellaSwag, WinoGrande and PIQA do not appear anywhere else on this card.

The GSM8K regression this card reports does reproduce. The table above reports GSM8K falling 88.6% → 85.2% under the fine-tune, a cost of 3.4 points. This separate run, on a different harness at a different operating point, independently measures −2.73 points with 106 discordant pairs on 1,319 items, an exact p of 0.00061, and it survives correction. Same sign, and the sizes agree to within 0.7 points across two protocols that share almost nothing. That is the clearest thing in this section, and it is a point in the card's favour: the regression was published when it could have been left out, and it holds up when someone else measures it.

GSM8K is directional, not a budget-independent measurement. The generation cap is part of the protocol and its effect is large — capping or uncapping generation can move a GSM8K score by several points with nothing else changed, because an uncapped model generates past its answer and the strict-match extractor loses it. That does not cancel in a paired comparison: each model over-generates by a different amount, so a verbose model is penalised more than a terse one and the cap silently reweights the comparison. A thinking model against a terse base is the worst case for this. Both arms here carried the same 512-token cap and it was chosen before the scores were seen, which is the most that protocol can do. Read that row as a direction, not as a number independent of the budget — and note that the agreement with the card's own figure is agreement between two different budgets, neither of which is the other's control.

MMLU: two operating points, not one disagreement. The table above reports MMLU rising 69.9% → 71.6%, and the paragraph under it says MMLU rose across nearly all 57 subjects. This run measures −0.27 points, which does not survive Holm correction. That is not a refutation of the card's figure, and it should not be read as one.

The reason is specific rather than rhetorical. The Measurement note above records that the card's benchmarks ran with enable_thinking=False — a setting that only exists when the chat template is rendered. This run applied no chat template to either arm. Those are two genuinely different operating points for a thinking model, not two attempts at the same one.

We tried to match the card's operating point and could not. Applying the chat template to both arms produced chance-level output on MMLU, which is 4-way multiple choice with a 0.25 floor: the base scored 0.2298 and this release 0.2762. Numbers at the floor are not a measurement of knowledge, so that comparison was discarded as protocol-invalid rather than reported. No matched-protocol comparison of the card's MMLU figure exists, here or anywhere else on this card.

What is left is a bound, not a verdict. Under this protocol, on 14,042 items with 344 discordant pairs, any MMLU difference lies between 0.49 points below the base and 0.05 points below it — smaller than half a point in magnitude, and not established at this family's Holm threshold of 0.05/3 = 0.0167 for the fourth-ranked test. The card's +1.7-point figure was measured at an operating point this run did not reach. It is untested here, and nothing in this section shows it to be wrong.

The three tasks that do not survive correction are bounds, not nulls. "Not distinguishable" is easy to misread as "no difference", so here is what each one actually constrains, on the items measured:

  • MMLU — any difference lies between 0.49 points below the base and 0.05 points below it (14,042 items, 344 discordant).
  • WinoGrande — any difference lies between 1.65 points below the base and 0.59 points above it (1,267 items, 67 discordant).
  • PIQA — any difference lies between 0.71 points below the base and 0.50 points above it (1,838 items, 40 discordant).

The WinoGrande and PIQA discordant counts are small, which is why exact Clopper-Pearson intervals are used throughout rather than a normal approximation.

How this relates to the figures above. It does not revise them, and no figure above has been changed, recomputed or removed. The figures above came from an internal run whose harness, version, configuration, prompt format, decoding settings, item counts and date were not recorded, and which used enable_thinking=False; this is lm-evaluation-harness 0.4.11 on the full task sets with no chat template. Different harness, different protocol, different operating point — the two sets of numbers are not comparable, and neither corrects the other. Two of the six tasks here (GSM8K, MMLU) share a name with a row above; sharing a name is not sharing a protocol.

What this section does not do. It does not close the entry under What is not measured yet recording that this card's own capability rows have no per-example logs, no discordant count and no paired test. That entry stands as written: it describes the figures above, which remain untested aggregates measured somewhere this run did not reach, and this is a separate measurement rather than a retrofit of those. This section covers six tasks against one base on the BF16 safetensors and nothing else — no safety, jailbreak, refusal, CBLRE, citation-integrity, character, tool-use, instruction-following or coding result is re-measured here, no quantized tier is covered, no image or video input is touched, and nothing here is a safety, security or fitness-for-purpose claim. No model revision is recorded with it either; like the figures above, it is attached to repository names rather than to specific bytes.


Supporting evaluation — regional / legal

CBLRE (Canadian bilingual legal/regulatory eval) — average 85.8% across subtasks: common law 95.2%, constitutional charter 90.9%, privacy compliance 90.9%, safety calibration 84.1%, Québec civil law 75.0%, citation integrity 78.4%.

Bilingual parity: English 100% vs French 81.8% on the privacy-compliance subset (parity ratio 0.82). French is inherited from the base, not specially tuned — usable, not specialized.

Character and honesty dimensions are evaluated qualitatively for this release via the prompt pack above. Quantitative character metrics will be published as the eval harness matures.

Measurement note. Benchmarks ran with enable_thinking=False; the model ships with thinking enabled. Published scores may be conservative floors, particularly for CBLRE. That "may" is doing real work: no thinking-on run exists to compare against, so the direction is an expectation and not a result.


What is not measured yet

Named specifically, so that the gaps do not read as coverage.

  • Character, restraint, voice and anti-sycophancy. Qualitative only. No instrument, no held-out set, no rate. The exchanges under Character are demonstrations. This is the axis a Vinci release is meant to be judged on and the one with the least evidence behind it.
  • Every GGUF tier. Q6_K, Q5_K_M and Q4_K_M are unevaluated. No number on this card was measured on them and none should be copied onto them.
  • Vision and video. The base's towers are present in the weights and were frozen during fine-tuning. Nothing here measures image or video input. See Modality scope.
  • Thinking mode. The published run used enable_thinking=False. No paired thinking-on versus thinking-off comparison is reported, so the cost or benefit of the mode the model ships in is not established here.
  • The effect of the fine-tune. No paired per-item test against Qwen/Qwen3.5-9B on any benchmark above, so none of these figures establishes that the tuning caused the difference rather than run-to-run or harness variation.
  • Item counts and confidence intervals. Absent on every row but HumanEval.
  • The withheld numbers. MBPP, multi-step arithmetic and word-sorting have not been re-run with a format-insensitive or execution-based grader. We know the reported movements were extraction artifacts; we do not know what the underlying ability did.
  • Long context. The context figure in Model at a glance is a configured position limit. No retrieval, recall or reliability result at length is reported.
  • French beyond one subset. The parity figure covers the privacy-compliance subset of CBLRE and nothing else.
  • CBLRE itself. Our own instrument, not independently qualified. Its subtask scores characterize our items, not the field.
  • Multi-turn behaviour over long conversations, and whether character holds across a long session.
  • The model revision behind the numbers. No commit or revision was recorded with the run, so the figures are attached to a model name rather than to specific bytes.

Limitations

Vinci ships honest limitations. These are real, and they are the reason for the right-hand column of Should you use this model? above.

Strict instruction-following regressed. IFEval fell 83.7% → 73.0%, a 10.7-point drop we believe is real rather than a measurement artifact. Character training can add framing that costs literal format compliance. If you need rigid output — exact JSON, exact word counts, no preamble — this is a regression from the base model. Being addressed with capability-rehearsal data in the next round.

Tool-use abstention regressed. Bozza is better at selecting and invoking the right tool on realistic prompts, and worse at holding back when no tool applies. Overall BFCL fell 29.3% → 25.3% for that reason alone. We therefore do not claim bozza is "better at tool use." Part of this may be an eval-harness artifact and is under investigation; a tool-abstention rehearsal slice is planned.

Multi-turn task-holding regressed. 36.3% → 29.1%.

Agentic use. Piccolo and bozza are the sizes we expect to do agentic work beneath larger models. The three regressions above matter most there, not least — a small model in an agent loop has its output parsed by a machine, its tool calls execute, and its errors compound. Bozza 1.0 is not recommended for autonomous agentic loops. It takes on that role when instruction-following and tool abstention recover — not before.

Legal-citation fabrication. In domains the model half-knows — specific statutes, regulatory sections — it can confidently state a wrong citation or section number. This is distinct from the fake-paper decline demonstrated in Character above: the model may not know it is wrong. Observed ~1 in 5 on CBLRE legal-citation items. Do not rely on it for legal citations without grounding or retrieval. Being addressed with domain grounding and anti-fabrication rehearsal in the next round.

Not a frontier coding or math engine. HumanEval 68.3%, GSM8K 85.2%, both slightly below base.

English-primary. French-jurisdiction legal reasoning (Québec civil law) regressed. Scope decision, not a surprise, but a real limitation for francophone legal work.

Measurement caveat. Several large per-task swings in this release are answer-extraction artifacts from formatting drift, not capability changes. See Numbers we are not reporting.


Training

  • Base: Qwen/Qwen3.5-9B
  • Method: SFT → DPO. Light-touch character, restraint, and honesty tuning. Disposition, not capability retrain.
  • SFT stage: Trained on the base-revoiced, thinking-identity, and fact-gold slices of the SimpleDirect character corpus.
  • DPO stage: Preference pairs drawn from the same corpus.
  • Corpus: SimpleDirect Constitution + character training corpus (base-independent; the same corpus trains every Vinci tier).
  • Compute: 4× NVIDIA H200 (80 GB HBM3). LoRA + DoRA, vision tower frozen, bf16 precision.

Our differentiators — restraint, anti-fabrication, voice — are not measured by the benchmarks above. Those benchmarks measure capability and safety. Character is measured separately, and that is what a Vinci release is judged on.


Running it

Because the weights are open, you can run bozza yourself — locally, offline, on your own machine. Nothing has to leave your device.

Generation config

These are the recommended sampling parameters for Vinci-Bozza-1.0. They differ from Qwen upstream defaults on temperature and presence_penalty — the disposition tune changes the distribution of character-relevant tokens, and the upstream general-task settings required empirical adjustment.

Thinking enabled (enable_thinking=True) — model reasons before responding:

Use case temperature top_p top_k min_p presence_penalty max_tokens
General / chat 0.7 0.95 20 0.05 0.0 16384
Coding / precise output 0.6 0.95 20 0.0 0.0 16384

Thinking disabled (enable_thinking=False) — faster, no <think> trace, lower latency:

Use case temperature top_p top_k min_p presence_penalty max_tokens
General / chat 0.7 0.8 20 0.0 0.0 4096

presence_penalty is intentionally 0.0. Values ≥ 1.5 (the Qwen upstream default for general tasks) penalise the model's trained refusal and uncertainty tokens, causing hedged-then-elaborate fabrication on unverifiable citations. Do not add it for character tasks.

transformers

from transformers import AutoModelForImageTextToText, AutoTokenizer

model_id = "simpledirect/Vinci-Bozza-1.0"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(model_id, device_map="auto", torch_dtype="auto")

messages = [{"role": "user", "content": "Hello, who are you?"}]
inputs = tok.apply_chat_template(
    messages, add_generation_prompt=True, enable_thinking=False,
    return_tensors="pt", return_dict=True
).to(model.device)
out = model.generate(**inputs, max_new_tokens=512)
print(tok.decode(out[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))

vLLM

vllm serve simpledirect/Vinci-Bozza-1.0

Local — GGUF (Ollama / LM Studio / llama.cpp)

GGUF builds: simpledirect/Vinci-Bozza-1.0-GGUF. Text-only — the vision tower is not included in the GGUF conversion. The safetensors build above does carry it, but nothing on this card measures image or video input either way; see Modality scope.

No GGUF tier has been evaluated. Every figure on this card was measured on the BF16 safetensors. Sizes below are the file sizes in the companion repository, read on 2026-09-21.

Variant Size (GiB) Notes
Q6_K ~6.8 GB Closest to BF16 quality
Q5_K_M ~6.0 GB Good balance (recommended)
Q4_K_M ~5.2 GB Smallest, tight memory budgets
ollama run hf.co/simpledirect/Vinci-Bozza-1.0-GGUF

Hardware:

Format GPU VRAM System RAM (CPU)
BF16 18 GB min, 24 GB recommended —
Q6_K 10 GB 16 GB
Q5_K_M 8 GB 14 GB
Q4_K_M 7 GB 12 GB

Mac M-series: Q5_K_M on 16 GB; Q6_K needs 24 GB.

These are guidance figures, not a measured result, and they do not include KV cache for long contexts. Size the BF16 row against the weight file itself — 18,819,722,392 bytes, about 17.5 GiB before any runtime overhead.

Prompt format

Qwen / ChatML chat template. Pass enable_thinking=False to suppress <think> blocks for standard chat. Thinking mode is on by default — see Thinking mode above for the vLLM and GGUF equivalents.

messages = [{"role": "user", "content": "Your message here"}]
inputs = tok.apply_chat_template(
    messages,
    add_generation_prompt=True,
    enable_thinking=False,   # set True for reasoning tasks
    return_tensors="pt", return_dict=True
).to(model.device)

No system prompt is required — character and values are in the weights.


The Constitution

You can read exactly what we trained bozza to value: https://guide.getsimpledirect.com/constitution


License & attribution

Apache 2.0. The version you have is yours to keep — it cannot be deprecated out from under you, revoked, or changed without your say.

Fine-tuned from Qwen/Qwen3.5-9B (Apache 2.0). Character training corpus by SimpleDirect (Apache 2.0).


Citation

@misc{vinci-bozza-2026,
  title        = {Vinci Bozza 1.0},
  author       = {{SimpleDirect}},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/simpledirect/Vinci-Bozza-1.0}},
  note         = {Apache 2.0. Fine-tuned from Qwen/Qwen3.5-9B.},
}

Bozza 1.0 is a disposition upgrade on a strong base: safer, more honest, same brain. It cost us some literalism and some restraint around tools, and we've told you exactly how much. We release early and iterate in the open — tell us what works and what doesn't.


Building in the open

Bozza 1.0 is the worst Bozza will ever be. We're iterating fast — the next version addresses IFEval regression and adds anti-fabrication and anti-sycophancy rehearsal data. Tell us what works and what breaks: open an issue, drop a note, or ping us at the links below. Public evals and character scores improve with real-world feedback.


Links


About

Vinci is a family of open-weight models from SimpleDirect, built on the conviction that character — not raw capability — is what's becoming scarce.

Vinci Bozza is the 9B tier. More models, sharing the same Constitution and character, are on the way.

Configuration

Architecture
Qwen3_5ForConditionalGeneration
Context length (tokens)
262,144
Layers
32
Hidden size
4,096
Feed-forward size
12,288
Attention heads
16
Key/value heads
4
Head dimension
256
Vocabulary size
248,320
Model type
qwen3_5

Identity and Version

Repository
simpledirect/Vinci-Bozza-1.0
Publisher
SimpleDirect
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
9.4B parameters
Languages
en
Revision
ade3b472bdb09762ce6eb6cc428366e687a775be
First published
2026-07-08
Last updated
2026-09-28

Files and Weights

11 files, 18.8 GB in total. The weights are 1 file totalling 18.8 GB in safetensors.

Weights1 file · 18.8 GB
Configuration3 files · 4.2 KB
Tokenizer2 files · 20.0 MB
Documentation1 file · 40.4 KB
Other3 files · 16.7 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights18.8 GB cd4ab1e07b35
config.jsonConfiguration2.8 KB —
generation_config.jsonConfiguration133 B —
processor_config.jsonConfiguration1.2 KB —
README.mdDocumentation40.4 KB —
assets/vero-pixel.svgOther669 B —
assets/vinci-vero-header.svgOther8.2 KB —
chat_template.jinjaOther7.8 KB —
.gitattributesRepository1.6 KB —
tokenizer.jsonTokenizer20.0 MB 06b9509352d2
tokenizer_config.jsonTokenizer9.1 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
18.8 GB
Download from SimpleDirect

Released by SimpleDirect through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published18.8 GB
16-bit18.8 GB
8-bit9.4 GB
4-bit4.7 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Built on This Model

Questions About Vinci-Bozza-1.0

How much GPU memory does Vinci-Bozza-1.0 need?

About 22.6 GB at 16-bit and 5.6 GB at 4-bit: the weights (9.4B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Vinci-Bozza-1.0 on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Vinci-Bozza-1.0 commercially?

Yes. Vinci-Bozza-1.0 is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is Vinci-Bozza-1.0's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

Qwen-Image-2.1-PE-I2I-Heretic

Darrellbest

The image-editing prompt rewriter for Qwen-Image-2.1, a fine-tuned Qwen3.5-VL 9B that turns a short edit instruction plus 1–N input images into a detailed English edit prompt, with its refusal behaviour removed by Heretic directional ablation. bf16, same shapes and parameter count as the source; nothing else was changed. systemprompt.txt is included and required. It defines the output format. It is the unmodified file from the source repo. The second row is an independent evaluation of the exported weights with Heretic's evaluatemodel. Refusals were measured on mlabonne/harmfulbehaviors and KL divergence (damage to ordinary behaviour, first-token distributions) on mlabonne/harmlessalpaca…

Open weights other 9.4B parameters 262,144 tokens

Model · Image and text to text

LucidVitality-9b

Matthew Andrews

LucidVitality-9B is a little goer. It's a roleplay or creative focused merge of two Qwen3.5 9b variants for people with absolute potatoes, like myself. It marries the improved prose of Darkhn's Qwen3.5-9B-Animus-V13.0 with the lower looping, higher EOS exit, and slightly more coherency (compared to base) from Qwen3.5-9B-Claude-4.6-HighIQ-INSTRUCT-HERETIC-UNCENSORED of DavidAU's making. I haven't merged anything for a long time, as it's been a hard time for finetuning. They rarely increase prose quality, often deeply lose intelligence over base (even when tuned for intelligence or agentic). Base model's getting tough to beat. So it was a pleasant surprise to find two models that each…

Open weights 9.4B parameters 262,144 tokens transformers

Model · Image and text to text

blink-mimo-9b

Govind Kamtamneni

Send a text or JSON state and your questions: choice picks from up to 255 options, noul is yes/no, and score takes 2–10 ordered levels. Each question gets probabilities over its offered options from one forward pass, with no generated text. Long or large multi-question requests may use several batches. Personal research release by thegovind, not an official product of any company. No affiliation with TypeSafe AI, Xiaomi, Alibaba Cloud or the Qwen team. Weights are for non-commercial research; see Licence. We ran the full Decision Index 0.2 suite ourselves with the official scoring kit at commit 19ad28e on 2026-09-25. This is a descriptive run, not a leaderboard submission or accepted…

Open weights other 9.4B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen-Image-2.1-PE-I2I-Heretic-NVFP4

Darrellbest

NVFP4 (4-bit floating point, W4A4) build of darrellbest/Qwen-Image-2.1-PE-I2I-Heretic, the refusal-ablated image-editing prompt rewriter for Qwen-Image-2.1. For vLLM on NVIDIA Blackwell GPUs, which run NVFP4 natively. 11 GB instead of 18 GB. systemprompt.txt is included and required, exactly as for the original. The linear-attention layers carry a recurrent state and the vision tower encodes the input image; both were left in bf16, as other quantizations of this model family do. That is why the file is 11 GB rather than ~6 GB. Made with llm-compressor 0.13.0 (QuantizationModifier, scheme="NVFP4"), calibrated on 64 samples in the model's real input format: its own system prompt, an edit…

Open weights other 9.4B parameters 262,144 tokens

Model · Image and text to text

Omni-Edu-9B

Hao Liang

This model is a fine-tuned version of Qwen/Qwen3.5-9B-Base on the Omni-Edu-70K dataset. The following hyperparameters were used during training: - learningrate: 5e-06 - trainbatchsize: 1 - evalbatchsize: 8 - distributedtype: multi-GPU - numdevices: 8 - gradientaccumulationsteps: 8 - totaltrainbatchsize: 64 - totalevalbatchsize: 64 - lrschedulertype: cosine - lrschedulerwarmupsteps: 0.1 - numepochs: 3.0 - Transformers 5.2.0 - Pytorch 2.10.0 - Datasets 4.0.0 - Tokenizers 0.22.2

Open weights other 9.4B parameters 262,144 tokens transformers