SAVRN
Search Contact SAVRN

Open-weight model · Text generation

Vinci-Cyber-30B-1.0-GGUF

by SimpleDirect simpledirect/Vinci-Cyber-30B-1.0-GGUF

Vinci-Cyber-30B-1.0-GGUF is an open-weight model for text generation from SimpleDirect, released under Apache License 2.0. Its published files total 126.4 GB. It draws 60 downloads a month.

Quantized GGUF conversions of simpledirect/Vinci-Cyber-30B-1.0, a narrow, verifier-grounded Infrastructure-as-Code security specialist developed by SimpleDirect, a Canadian AI Lab, from IBM Granite 4.1 30B.

Parameters—
Context—
Weights126.4 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads60

Model Card

By SimpleDirect, published under apache-2.0, revision 8d8c6f656bfb.

Quantized GGUF conversions of simpledirect/Vinci-Cyber-30B-1.0, a narrow, verifier-grounded Infrastructure-as-Code security specialist developed by SimpleDirect, a Canadian AI Lab, from IBM Granite 4.1 30B. See the main repository for what the model does, how it was trained, and its evaluation status. What this repository establishes, in one paragraph. All four tiers convert completely (578/578 tensors), carry faithful architecture metadata (14/14 fields), load and generate on CUDA with zero NaN and zero errors, and have had their distributional distortion against the BF16 quant master measured per token and reported below. What it does not establish is capability: no tier was benchmarked…

Read SimpleDirect's full model card

Quantized GGUF conversions of simpledirect/Vinci-Cyber-30B-1.0, a narrow, verifier-grounded Infrastructure-as-Code security specialist developed by SimpleDirect, a Canadian AI Lab, from IBM Granite 4.1 30B. See the main repository for what the model does, how it was trained, and its evaluation status.

What this repository establishes, in one paragraph. All four tiers convert completely (578/578 tensors), carry faithful architecture metadata (14/14 fields), load and generate on CUDA with zero NaN and zero errors, and have had their distributional distortion against the BF16 quant master measured per token and reported below. What it does not establish is capability: no tier was benchmarked, so this card names no best tier and no recommended tier. Pick against your memory budget, verify the digest, and evaluate on your own cases.

Files

All four tiers are derived from the exact Vinci-Cyber-30B-1.0 release master — the same merged safetensors export the main repository publishes. The quantized tiers were produced from the BF16 quant master, not from the F16 file: quantizing from F16 would round bf16 → f16 → quant and lose precision twice for no benefit.

File Format Bytes Bits/weight
Vinci-Cyber-30B-1.0-Q4_K_M.gguf Q4_K_M 17,490,240,800 4.85
Vinci-Cyber-30B-1.0-Q5_K_M.gguf Q5_K_M 20,493,362,464 5.68
Vinci-Cyber-30B-1.0-Q8_0.gguf Q8_0 30,674,969,888 8.50
Vinci-Cyber-30B-1.0-F16.gguf F16 57,736,095,008 16

SHA-256 digests, as computed on the source artifacts immediately before upload:

Vinci-Cyber-30B-1.0-Q4_K_M.gguf  49896b05b8c2c2deca652f6e82d5f78618973c6b5deea2bfe547380bed17ef55
Vinci-Cyber-30B-1.0-Q5_K_M.gguf  e95429220db9c8a2ba874fd63f25d5eaa4ce2784b22a35b7fe67c2c287bce0d2
Vinci-Cyber-30B-1.0-Q8_0.gguf    b1b912673c8e788193e4f17e99f7fc6fc93e78adca7d59ebf365fe59969887c8
Vinci-Cyber-30B-1.0-F16.gguf     8680d38463dba8b2c258398421e847b4b31e70861aa2a2bb655e57b016c1e125

Hugging Face stores these as content-addressed LFS objects, so the lfs.sha256 reported by the Hub for each file can be compared directly against the digests above without downloading the file.

Converted with llama.cpp pinned at source revision 6f4f53f2b7da54fcdbbecaaa734337c337ad6176.

Why this card publishes per-tier figures and its 8B sibling does not. The 8B GGUF card withholds its per-tier numbers because that evaluation never captured the llama.cpp build identity that produced them. This one did — the pinned revision above. The difference is a provenance difference, not a quality difference between the two model families.

Validation status

Structural integrity — established

  • Conversion is complete: 578 / 578 tensors. Every tensor from the merged safetensors export is present in the GGUF, compared as sets in both directions — no missing tensors, no extras. The expected names were derived through llama.cpp's own tensor-name map rather than a hand-written list.
  • Architecture metadata is faithful: 14 / 14 fields. Block count, context length, embedding dimensions, head counts, RoPE base, the Granite-specific attention/embedding/residual/logit scales, and vocabulary size were read back out of the GGUF and matched against the merged model's config.json. Zero mismatches.
  • Digests verified. Each published file's SHA-256 was recomputed on the source artifact immediately before upload and matched the digest recorded at build time.
  • The tokenizer is identical to the parent's, byte for byte by digest. Any difference between this model and its parent therefore comes from weights alone, not from tokenization or configuration.

Runtime smoke — established

All four tiers were loaded and executed on CUDA. Every tier loads, reaches generation, and emits non-empty output, with 0 NaN and 0 errors observed across the run.

This establishes that the files execute. It is a liveness check, not a quality measurement.

Distributional fidelity — measured and reported

Measured with the llama.cpp per-token KLD estimator (uint16-quantized reference, −16-nat tail truncation, unnormalized plug-in) — the precise name matters, because it is not exact KL — against the BF16 quant master, over 4,080 scored positions with 0 missing and 0 non-finite positions.

On the one crossing in the table: Q8_0's median (0.000491) sits below F16's (0.000571), which looks like an inversion. We draw no conclusion from it, and here is the reason rather than just the refusal — the gap is 0.000080, which is 1.82× the instrument's own floor extremum of 0.000044 (the most negative value the self-control produced). A difference under 2× the noise extremum is not a measurement of anything. Mean KL does order F16 < Q8_0 < Q5_K_M < Q4_K_M, but monotonic degradation was never required and is not claimed.

These are estimates, not exact KL divergences. An independent review of the pinned implementation established that the estimator stores the reference side uint16-quantized and discards vocabulary entries whose reconstructed reference log-probability falls below −16 nats (simulated excluded mass 0.14–0.34%). An earlier version of this card said "across the full vocabulary", which was literally inaccurate and is corrected here. The quantization floor is hard-bounded at 1.2207e-4 nats and every tier's mean sits above it — F16 by 10.3×, Q8_0 by 14.6×, Q5_K_M by 140×, Q4_K_M by 363× — so the comparisons below are not materially affected.

Tier Mean KL Median P99 Max Top-1 agreement
F16 0.001262 0.000571 0.0095 0.053 96.225%
Q8_0 0.001783 0.000491 0.0155 0.296 96.544%
Q5_K_M 0.017091 0.002774 0.1846 3.055 93.309%
Q4_K_M 0.044345 0.008324 0.4744 3.928 89.338%

Instrument controls. The instrument's self-control (the master compared against itself through the same pipeline) sits at approximately 1e-5, and a planted-difference control registered 18.6 nats — so the instrument both reads near-zero where it should and moves when a real difference is present.

General-capability benchmarks (measured on the unquantized weights)

These figures characterize the source model, not the GGUF conversions in this repository. They were measured on the bfloat16 safetensors of simpledirect/Vinci-Cyber-30B-1.0, snapshot e25c67861096a4fa52d7f7367f95d90b8db9f726, loaded in bfloat16 through transformers. No quant in this repository — F16, Q4_K_M, Q5_K_M or Q8_0 — has been evaluated on these tasks, and quantization is expected to move the scores. Anyone who needs quant-specific numbers must measure the quant.

Harness: lm-evaluation-harness 0.4.11, 0-shot, seed 0, dtype bfloat16, batch size 8, single H200.

The scores below belong to the bf16 source model. They are not a measurement of any file in this repository.

Benchmark Metric Vinci-Cyber-30B-1.0 (bf16) granite-4.1-30b (bf16) n
ARC-Challenge acc_norm 0.6664 0.6570 1172
HellaSwag acc_norm 0.8513 0.8507 10042
PIQA acc_norm 0.8368 0.8341 1838
WinoGrande acc 0.7593 0.7577 1267
MMLU acc 0.7839 0.7824 14042
  • Each task was run twice on both models, but repetition is not evidence. On ARC-Challenge, HellaSwag, PIQA and MMLU both arms reproduced exactly — which under greedy decoding on a fixed item set happens whatever the true accuracy, so it measures the harness rather than the models. The differences are tiny, 0.06% to 0.94% absolute. Paired item-level testing plus a pre-registered confirmatory test on 1,418 held-out ARC-Challenge items gives the supported statement: ARC-Challenge is an established improvement (+1.06 percentage points, p = 0.014, confirmed against a negative control registered before the run), and no task shows an established regression — HellaSwag is bounded below 0.14 percentage points.
  • WinoGrande does not support a difference. On a repeat the base scored 0.7593, exactly the Vinci-Cyber score, so the two models are indistinguishable there.
  • GSM8K is not usable as a comparison. It is approximately 0.905 (5-shot, strict-match, mean of six runs, range [0.9037, 0.9083]) for Vinci-Cyber-30B-1.0. A paired item-level test over the 1,319 questions both models answered — exact McNemar, aligned per question — gives 18 disagreements, 11 favouring this model and 7 the base, p = 0.48. GSM8K cannot resolve a difference on this pair in either direction. (An earlier revision justified this by comparing one arm's run-to-run spread against the gap between arms. That was the wrong test — one arm's spread is not the yardstick for the gap between arms — and the paired test above replaces it.)
  • MMLU is only comparable across identical harness settings — it moves up to 15 points on prompt formatting. The other four were validated against published lm-eval figures for Meta-Llama-3.1-8B-Instruct, agreeing within ±0.01 with deltas scattering in sign.

The full matrix and the repeat runs establishing these noise floors are recorded internally in getsimpledirect/vinci-gpu-research under model-scouting/eval-matrix-20260920/.

Limitations — read these before selecting a file

These are derived deployment formats. Structural integrity and runtime execution were validated, and distributional fidelity is reported above. The quantized tiers were NOT independently capability-benchmarked. Nothing here establishes that any tier is behaviourally identical to the release master, and that phrase must not be used of them.

The main model's IaC result does not transfer to these files. The 5 / 8 valid-verified-repairs figure published on the main repository was measured on the merged safetensors export, not on these GGUF bytes. It must not be copied onto any quant tier as though it were measured here.

The general-capability benchmark figures do not transfer either. The ARC-Challenge, HellaSwag, PIQA, WinoGrande and MMLU scores reported above were measured on the unquantized bfloat16 safetensors of the source model, not on these GGUF bytes, and they are not a measurement of any file in this repository.

The KL numbers are GGUF-vs-GGUF. They compare each tier against the BF16 quant master, which is itself a GGUF. The fidelity of the upstream merged-safetensors → BF16.gguf conversion step is not covered by this protocol and is UNEVALUABLE from these measurements.

One text domain only. KL was measured on a single corpus, derived from the model's own training data. Distortion on any other domain is unmeasured.

The corpus has limited power to separate quantization distortion from the fine-tune's own footprint. For context: the unadapted parent scores mean KL 0.0915 with 89.412% top-1 agreement on this same corpus — statistically indistinguishable from Q4_K_M's 89.338%. Because a completely different model lands where Q4_K_M lands on the top-1 metric, this corpus cannot be used as a pass/fail line, and no conclusion about Q4_K_M should be drawn from that coincidence in either direction.

No tier is recommended on quality grounds. No capability comparison between tiers exists, so this card does not name a "best" or "recommended" quant. The only tradeoff it can describe is the mechanical one: a smaller file needs less memory, a larger file carries less quantization. Choose against your own hardware budget and your own evaluation.

Structural checks are not quality. A GGUF that is well-formed and carries the right architecture is not thereby a GGUF that behaves like the model it came from.

Responsible deployment

Conceptual intended workflow, not a measured reliability result or an autonomous deployment loop. Validate and review both proposed changes and no-change decisions.

Evaluate in a sandbox. Independently validate and review every proposed configuration change, run your own configuration checks and regression tests against it, and require human approval before anything is applied. Do not give the model unrestricted production access. A cleared scanner is not evidence that a patch is correct — this family has published cases where it is not. Self-hosting alone does not establish privacy, security, data residency, or compliance.

These are unmeasured quantized artifacts: no tier here was capability-benchmarked, so validate the tier you choose on your own cases before relying on it.

Licence and attribution

Apache-2.0. The accompanying LICENSE and NOTICE files are part of this distribution.

These files are quantized derivatives of a modified version of IBM Granite 4.1 30B, bound to parent revision 4fae6278f7132abf5e971f9de49ebbad09c54cce. SimpleDirect changed model weights through DoRA/rsLoRA fine-tuning, merged the selected adapter into the pinned parent, and converted the merged export to the GGUF files published here.

Redistribution must include the licence, retain applicable upstream notices, and carry prominent modification notices with modified files, as required by Apache-2.0. The intended-use and deployment guidance in this card does not add a separate field-of-use restriction to that licence. No endorsement by IBM or the Government of Canada is implied.

Full release provenance and per-file digests for the underlying model are in the main repository's MODEL-PROVENANCE.json.

Claims reviewed as of September 21, 2026; latest measurement reported here is the general-capability matrix of September 20, 2026. This card's artifact-specific claims require review against the exact repository revision being published.

Identity and Version

Repository
simpledirect/Vinci-Cyber-30B-1.0-GGUF
Publisher
SimpleDirect
Task
Text generation
Modality
Text
Library
Not stated by the source
Parameters
Not stated by the source
Languages
Not stated by the source
Revision
8d8c6f656bfbc44f2e563adf68fea81b4e7573ca
First published
2026-09-18
Last updated
2026-09-28

Files and Weights

13 files, 126.4 GB in total. The weights are 4 files totalling 126.4 GB in gguf.

Weights4 files · 126.4 GB
Documentation3 files · 31.0 KB
Other5 files · 1.9 MB
Repository1 file · 1.9 KB
Every file
FileTypeSizeSHA-256
Vinci-Cyber-30B-1.0-F16.ggufWeights57.7 GB 8680d38463db
Vinci-Cyber-30B-1.0-Q4_K_M.ggufWeights17.5 GB 49896b05b8c2
Vinci-Cyber-30B-1.0-Q5_K_M.ggufWeights20.5 GB e95429220db9
Vinci-Cyber-30B-1.0-Q8_0.ggufWeights30.7 GB b1b912673c8e
LICENSEDocumentation11.4 KB —
NOTICEDocumentation4.7 KB —
README.mdDocumentation14.9 KB —
assets/cyber-banner.pngOther1.7 MB 1185eb58b7be
assets/cyber-shared-workflow.pngOther109.7 KB d7f5290808f3
assets/cyber-shared-workflow.svgOther3.5 KB —
assets/vero-pixel.svgOther669 B —
assets/vinci-vero-header.svgOther8.2 KB —
.gitattributesRepository1.9 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
126.4 GB
Download from SimpleDirect

Released by SimpleDirect through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published126.4 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Vinci-Cyber-30B-1.0-GGUF

Can I use Vinci-Cyber-30B-1.0-GGUF commercially?

Yes. Vinci-Cyber-30B-1.0-GGUF is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Fine-tune Qwen3 (14B) for free using our Google Colab notebook! - Read our Blog about Qwen3 support: unsloth.ai/blog/qwen3 - View the rest of our notebooks in our docs here. Qwen3-Coder is available in multiple sizes. Today, we're excited to introduce Qwen3-Coder-30B-A3B-Instruct. This streamlined model maintains impressive performance and efficiency, featuring the following key enhancements: - Significant Performance among open models on Agentic Coding, Agentic Browser-Use, and other foundational coding tasks. - Long-context Capabilities with native support for 256K tokens, extendable up to 1M tokens using Yarn, optimized for repository-scale understanding. - Agentic Coding supporting for…

Open weights apache-2.0 transformers

Model · Text generation

opt-125m

AI at Meta

OPT was first introduced in Open Pre-trained Transformer Language Models and first released in metaseq's repository on May 3rd 2022 by Meta AI. Disclaimer: The team releasing OPT wrote an official model card, which is available in Appendix D of the paper. Content from this model card has been written by the Hugging Face team. To quote the first two paragraphs of the official paper OPT was predominantly pretrained with English text, but a small amount of non-English data is still present within the training corpus via CommonCrawl. The model was pretrained using a causal language modeling (CLM) objective. OPT belongs to the same family of decoder-only models like GPT-3. As such, it was…

Open weights other 2,048 tokens transformers

Model · Text generation

Ornith-1.5-9B-GGUF

Ornith

Chirp Chirp! We are introducing Ornith-1.5, a major step toward building foundation models through end-to-end self-improvement. Ornith-1.5 extends Ornith-1.0 (which was developed on top of Qwen3.5 and Gemma4 with additional continued pretraining, mid-training, and post-training) by expanding the self-improvement loop from scaffold and rollout optimization to jointly optimizing task generation, scaffold construction, and solution rollouts. Rather than relying on a fixed set of human-curated tasks and manually designed harnesses, Ornith-1.5 continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning. For…

Open weights mit transformers

Model · Text generation

Ternary-Bonsai-2-27B-gguf

Prism ML

Full 27B-class reasoning in ternary transformer weights, for llama.cpp (CUDA, Metal, CPU) - \~5.9 GB language model (down from \~54 GB FP16) — full 27B-class reasoning on a standard laptop or a single GPU - 98.2% of FP16 intelligence retained: 84.78 average across 14 thinking-mode benchmarks — far above the conventional IQ2XXS build (72.59) at about 82% of its footprint, and within 0.4 points of UD-Q4KXL at three times the footprint - Retains thinking, reasoning, and agentic behavior deep in the sub-4-bit regime, where conventional low-bit representations collapse: math within half a point of full precision (96.57), coding level with the baseline (89.42), agentic tool calling at 74.92…

Open weights apache-2.0 llama.cpp

Model · Text generation

Ornith-1.5-35B-A3B-GGUF

Ornith

Chirp Chirp! We are introducing Ornith-1.5, a major step toward building foundation models through end-to-end self-improvement. Ornith-1.5 extends Ornith-1.0 (which was developed on top of Qwen3.5 and Gemma4 with additional continued pretraining, mid-training, and post-training) by expanding the self-improvement loop from scaffold and rollout optimization to jointly optimizing task generation, scaffold construction, and solution rollouts. Rather than relying on a fixed set of human-curated tasks and manually designed harnesses, Ornith-1.5 continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning. For…

Open weights mit transformers

Model · Text generation

Ornith-1.0-9B-GGUF

Ornith

Aloha! Today, we are releasing Ornith-1.0, a self-improving family of open-source models for agentic coding. This model card documents Ornith-1.0-9B, the most lightweight member of the Ornith family, designed for efficient single-GPU deployment. Ornith-1.0-9B is a dense ~9B model (≈19 GB in bf16), so it serves comfortably on a single 80GB GPU. The recipes below stand up an OpenAI-compatible server; add --tensor-parallel-size / --tp if you want to shard across more GPUs. For a quick local test (or to script offline generation), load the model directly with Transformers. Make sure you have a recent release installed — see the Transformers installation guide; Ornith-1.0-9B requires…

Open weights mit transformers