Quantized GGUF conversions of simpledirect/Vinci-Cyber-30B-1.0,
a narrow, verifier-grounded Infrastructure-as-Code security specialist developed by
SimpleDirect, a Canadian AI Lab, from IBM Granite 4.1 30B. See the main repository for what the
model does, how it was trained, and its evaluation status.
What this repository establishes, in one paragraph. All four tiers convert completely
(578/578 tensors), carry faithful architecture metadata (14/14 fields), load and generate on CUDA
with zero NaN and zero errors, and have had their distributional distortion against the BF16 quant
master measured per token and reported below. What it does not establish is capability: no tier
was benchmarked, so this card names no best tier and no recommended tier. Pick against your memory
budget, verify the digest, and evaluate on your own cases.
Files
All four tiers are derived from the exact Vinci-Cyber-30B-1.0 release master — the same merged
safetensors export the main repository publishes. The quantized tiers were produced from the
BF16 quant master, not from the F16 file: quantizing from F16 would round bf16 → f16 → quant and
lose precision twice for no benefit.
| File |
Format |
Bytes |
Bits/weight |
Vinci-Cyber-30B-1.0-Q4_K_M.gguf |
Q4_K_M |
17,490,240,800 |
4.85 |
Vinci-Cyber-30B-1.0-Q5_K_M.gguf |
Q5_K_M |
20,493,362,464 |
5.68 |
Vinci-Cyber-30B-1.0-Q8_0.gguf |
Q8_0 |
30,674,969,888 |
8.50 |
Vinci-Cyber-30B-1.0-F16.gguf |
F16 |
57,736,095,008 |
16 |
SHA-256 digests, as computed on the source artifacts immediately before upload:
Vinci-Cyber-30B-1.0-Q4_K_M.gguf 49896b05b8c2c2deca652f6e82d5f78618973c6b5deea2bfe547380bed17ef55
Vinci-Cyber-30B-1.0-Q5_K_M.gguf e95429220db9c8a2ba874fd63f25d5eaa4ce2784b22a35b7fe67c2c287bce0d2
Vinci-Cyber-30B-1.0-Q8_0.gguf b1b912673c8e788193e4f17e99f7fc6fc93e78adca7d59ebf365fe59969887c8
Vinci-Cyber-30B-1.0-F16.gguf 8680d38463dba8b2c258398421e847b4b31e70861aa2a2bb655e57b016c1e125
Hugging Face stores these as content-addressed LFS objects, so the lfs.sha256 reported by the Hub
for each file can be compared directly against the digests above without downloading the file.
Converted with llama.cpp pinned at source revision 6f4f53f2b7da54fcdbbecaaa734337c337ad6176.
Why this card publishes per-tier figures and its 8B sibling does not. The
8B GGUF card withholds its per-tier
numbers because that evaluation never captured the llama.cpp build identity that produced them.
This one did — the pinned revision above. The difference is a provenance difference, not a
quality difference between the two model families.
Validation status
Structural integrity — established
- Conversion is complete: 578 / 578 tensors. Every tensor from the merged safetensors export is
present in the GGUF, compared as sets in both directions — no missing tensors, no extras. The
expected names were derived through
llama.cpp's own tensor-name map rather than a hand-written
list.
- Architecture metadata is faithful: 14 / 14 fields. Block count, context length, embedding
dimensions, head counts, RoPE base, the Granite-specific attention/embedding/residual/logit
scales, and vocabulary size were read back out of the GGUF and matched against the merged model's
config.json. Zero mismatches.
- Digests verified. Each published file's SHA-256 was recomputed on the source artifact
immediately before upload and matched the digest recorded at build time.
- The tokenizer is identical to the parent's, byte for byte by digest. Any difference between
this model and its parent therefore comes from weights alone, not from tokenization or
configuration.
Runtime smoke — established
All four tiers were loaded and executed on CUDA. Every tier loads, reaches generation, and emits
non-empty output, with 0 NaN and 0 errors observed across the run.
This establishes that the files execute. It is a liveness check, not a quality measurement.
Distributional fidelity — measured and reported
Measured with the llama.cpp per-token KLD estimator (uint16-quantized reference, −16-nat tail
truncation, unnormalized plug-in) — the precise name matters, because it is not exact KL —
against the BF16 quant master, over 4,080 scored
positions with 0 missing and 0 non-finite positions.
On the one crossing in the table: Q8_0's median (0.000491) sits below F16's (0.000571), which
looks like an inversion. We draw no conclusion from it, and here is the reason rather than just the
refusal — the gap is 0.000080, which is 1.82× the instrument's own floor extremum of
0.000044 (the most negative value the self-control produced). A difference under 2× the noise
extremum is not a measurement of anything. Mean KL does order F16 < Q8_0 < Q5_K_M < Q4_K_M, but
monotonic degradation was never required and is not claimed.
These are estimates, not exact KL divergences. An independent review of the pinned
implementation established that the estimator stores the reference side uint16-quantized and
discards vocabulary entries whose reconstructed reference log-probability falls below −16 nats
(simulated excluded mass 0.14–0.34%). An earlier version of this card said "across the full
vocabulary", which was literally inaccurate and is corrected here. The quantization floor is
hard-bounded at 1.2207e-4 nats and every tier's mean sits above it — F16 by 10.3×, Q8_0 by 14.6×,
Q5_K_M by 140×, Q4_K_M by 363× — so the comparisons below are not materially affected.
| Tier |
Mean KL |
Median |
P99 |
Max |
Top-1 agreement |
| F16 |
0.001262 |
0.000571 |
0.0095 |
0.053 |
96.225% |
| Q8_0 |
0.001783 |
0.000491 |
0.0155 |
0.296 |
96.544% |
| Q5_K_M |
0.017091 |
0.002774 |
0.1846 |
3.055 |
93.309% |
| Q4_K_M |
0.044345 |
0.008324 |
0.4744 |
3.928 |
89.338% |
Instrument controls. The instrument's self-control (the master compared against itself through
the same pipeline) sits at approximately 1e-5, and a planted-difference control registered
18.6 nats — so the instrument both reads near-zero where it should and moves when a real
difference is present.
General-capability benchmarks (measured on the unquantized weights)
These figures characterize the source model, not the GGUF conversions in this repository. They
were measured on the bfloat16 safetensors of
simpledirect/Vinci-Cyber-30B-1.0,
snapshot e25c67861096a4fa52d7f7367f95d90b8db9f726, loaded in bfloat16 through transformers.
No quant in this repository — F16, Q4_K_M, Q5_K_M or Q8_0 — has been evaluated on these tasks,
and quantization is expected to move the scores. Anyone who needs quant-specific numbers must
measure the quant.
Harness: lm-evaluation-harness 0.4.11, 0-shot, seed 0, dtype bfloat16, batch size 8, single H200.
The scores below belong to the bf16 source model. They are not a measurement of any file in this
repository.
| Benchmark |
Metric |
Vinci-Cyber-30B-1.0 (bf16) |
granite-4.1-30b (bf16) |
n |
| ARC-Challenge |
acc_norm |
0.6664 |
0.6570 |
1172 |
| HellaSwag |
acc_norm |
0.8513 |
0.8507 |
10042 |
| PIQA |
acc_norm |
0.8368 |
0.8341 |
1838 |
| WinoGrande |
acc |
0.7593 |
0.7577 |
1267 |
| MMLU |
acc |
0.7839 |
0.7824 |
14042 |
- Each task was run twice on both models, but repetition is not evidence. On ARC-Challenge,
HellaSwag, PIQA and MMLU both arms reproduced exactly — which under greedy decoding on a fixed
item set happens whatever the true accuracy, so it measures the harness rather than the models.
The differences are tiny, 0.06% to 0.94% absolute. Paired item-level testing plus a pre-registered
confirmatory test on 1,418 held-out ARC-Challenge items gives the supported statement:
ARC-Challenge is an established improvement (+1.06 percentage points, p = 0.014, confirmed
against a negative control registered before the run), and no task shows an established
regression — HellaSwag is bounded below 0.14 percentage points.
- WinoGrande does not support a difference. On a repeat the base scored 0.7593, exactly the
Vinci-Cyber score, so the two models are indistinguishable there.
- GSM8K is not usable as a comparison. It is approximately 0.905 (5-shot, strict-match, mean of
six runs, range [0.9037, 0.9083]) for Vinci-Cyber-30B-1.0. A paired item-level test over the
1,319 questions both models answered — exact McNemar, aligned per question — gives 18
disagreements, 11 favouring this model and 7 the base, p = 0.48. GSM8K cannot resolve a difference
on this pair in either direction. (An earlier revision justified this by comparing one arm's
run-to-run spread against the gap between arms. That was the wrong test — one arm's spread is not
the yardstick for the gap between arms — and the paired test above replaces it.)
- MMLU is only comparable across identical harness settings — it moves up to 15 points on prompt
formatting. The other four were validated against published lm-eval figures for
Meta-Llama-3.1-8B-Instruct, agreeing within ±0.01 with deltas scattering in sign.
The full matrix and the repeat runs establishing these noise floors are recorded internally in
getsimpledirect/vinci-gpu-research under model-scouting/eval-matrix-20260920/.
Limitations — read these before selecting a file
These are derived deployment formats. Structural integrity and runtime execution were validated,
and distributional fidelity is reported above. The quantized tiers were NOT independently
capability-benchmarked. Nothing here establishes that any tier is behaviourally identical to the
release master, and that phrase must not be used of them.
The main model's IaC result does not transfer to these files. The 5 / 8 valid-verified-repairs
figure published on the main repository was measured on the merged safetensors export, not on
these GGUF bytes. It must not be copied onto any quant tier as though it were measured here.
The general-capability benchmark figures do not transfer either. The ARC-Challenge, HellaSwag,
PIQA, WinoGrande and MMLU scores reported above were measured on the unquantized bfloat16
safetensors of the source model, not on these GGUF bytes, and they are not a measurement of any
file in this repository.
The KL numbers are GGUF-vs-GGUF. They compare each tier against the BF16 quant master, which is
itself a GGUF. The fidelity of the upstream merged-safetensors → BF16.gguf conversion step is
not covered by this protocol and is UNEVALUABLE from these measurements.
One text domain only. KL was measured on a single corpus, derived from the model's own training
data. Distortion on any other domain is unmeasured.
The corpus has limited power to separate quantization distortion from the fine-tune's own
footprint. For context: the unadapted parent scores mean KL 0.0915 with 89.412% top-1
agreement on this same corpus — statistically indistinguishable from Q4_K_M's 89.338%. Because a
completely different model lands where Q4_K_M lands on the top-1 metric, this corpus cannot be used
as a pass/fail line, and no conclusion about Q4_K_M should be drawn from that coincidence in
either direction.
No tier is recommended on quality grounds. No capability comparison between tiers exists, so
this card does not name a "best" or "recommended" quant. The only tradeoff it can describe is the
mechanical one: a smaller file needs less memory, a larger file carries less quantization.
Choose against your own hardware budget and your own evaluation.
Structural checks are not quality. A GGUF that is well-formed and carries the right architecture is
not thereby a GGUF that behaves like the model it came from.
Responsible deployment
Conceptual intended workflow, not a measured reliability result or an autonomous deployment loop. Validate and review both proposed changes and no-change decisions.
Evaluate in a sandbox. Independently validate and review every proposed configuration change, run
your own configuration checks and regression tests against it, and require human approval before
anything is applied. Do not give the model unrestricted production access. A cleared scanner is not
evidence that a patch is correct — this family has published cases where it is not. Self-hosting
alone does not establish privacy, security, data residency, or compliance.
These are unmeasured quantized artifacts: no tier here was capability-benchmarked, so validate
the tier you choose on your own cases before relying on it.
Licence and attribution
Apache-2.0. The accompanying LICENSE and NOTICE files are part of this distribution.
These files are quantized derivatives of a modified version of
IBM Granite 4.1 30B, bound to parent revision
4fae6278f7132abf5e971f9de49ebbad09c54cce. SimpleDirect changed model weights through DoRA/rsLoRA
fine-tuning, merged the selected adapter into the pinned parent, and converted the merged export to
the GGUF files published here.
Redistribution must include the licence, retain applicable upstream notices, and carry prominent
modification notices with modified files, as required by Apache-2.0. The intended-use and deployment
guidance in this card does not add a separate field-of-use restriction to that licence. No
endorsement by IBM or the Government of Canada is implied.
Full release provenance and per-file digests for the underlying model are in the main repository's
MODEL-PROVENANCE.json.
Claims reviewed as of September 21, 2026; latest measurement reported here is the
general-capability matrix of September 20, 2026. This card's artifact-specific claims require
review against the exact repository revision being published.