SAVRN
Search Contact SAVRN

Open-weight model · Text generation

Vinci-Cyber-123B-1.0

by SimpleDirect simpledirect/Vinci-Cyber-123B-1.0

Vinci-Cyber-123B-1.0 is an open-weight model for text generation from SimpleDirect, released under other. It has 125B parameters and a 262,144-token context. At 16-bit it needs about 300.1 GB of GPU memory, which fits on 2x MI300X from $3.70 an hour; at 4-bit, 75 GB on 1x MI300X from $1.85, at the lowest prices in the SAVRN Index. It draws 1.5k downloads a month.

Vinci Cyber 123B 1.0 is an open-weight model for defensive infrastructure review and targeted remediation, fine-tuned in Canada from Mistral AI's Devstral 2 123B.

Parameters125B
Context262,144
Weights250.1 GB
Licenseother
AccessOpen weights
Monthly Downloads1.5k

Runs On

What it takes to serve Vinci-Cyber-123B-1.0 (125B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 250.1 GB 300.1 GB 2x MI300X (192 GB)
Vultr
$3.70 2x MI325X $4.00 · 2x MI355X $5.18
8-bit 125.0 GB 150.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x MI325X $2.00 · 1x MI355X $2.59
4-bit 62.5 GB 75.0 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.

Vinci-Cyber-123B-1.0 on every accelerator the SAVRN Index prices, at every precision

Model Card

Vinci Cyber 123B 1.0 is an open-weight model for defensive infrastructure review and targeted remediation, fine-tuned in Canada from Mistral AI's Devstral 2 123B. The released merged weights have now been tested directly, alongside their parent and the available GGUF formats. Focused repairs. Restraint on correct configuration. Weights you can run yourself. On the V2-B neutral-review test, the released BF16 model preserved 24/24 correct configurations and produced 18/24 scanner-credited repairs, all 18 passing offline provider-schema validation. Its parent repaired 17/24 and preserved 0/24. On the second set, V2-A, Cyber again preserved 24/24, but repaired 9/24 versus the parent's 15/24.…

Excerpt from the card by SimpleDirect, licensed other.

Configuration

Architecture
Ministral3ForCausalLM
Context length (tokens)
262,144
Layers
88
Hidden size
12,288
Feed-forward size
28,672
Attention heads
96
Key/value heads
8
Head dimension
128
Vocabulary size
131,072
Stored precision
bfloat16
Model type
ministral3

Identity and Version

Repository
simpledirect/Vinci-Cyber-123B-1.0
Publisher
SimpleDirect
Task
Text generation
Modality
Text
Library
transformers
Parameters
125B parameters
Languages
Not stated by the source
Revision
1c2b2438721fb08e0df19ffff4d02baabdbeaf9b
First published
2026-09-22
Last updated
2026-09-28

Files and Weights

51 files, 250.1 GB in total. The weights are 27 files totalling 250.1 GB in safetensors.

Weights27 files · 250.1 GB
Configuration6 files · 105.5 KB
Tokenizer2 files · 17.1 MB
Documentation3 files · 32.1 KB
Other12 files · 2.4 MB
Repository1 file · 1.8 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00027.safetensorsWeights6.6 GB 1fb4bba6e7a3
model-00002-of-00027.safetensorsWeights9.7 GB bab70bd5d9ca
model-00003-of-00027.safetensorsWeights9.7 GB 767a6c1c452c
model-00004-of-00027.safetensorsWeights9.7 GB 462e772c9e68
model-00005-of-00027.safetensorsWeights9.7 GB 4ca452feae22
model-00006-of-00027.safetensorsWeights9.7 GB 6f53db08910b
model-00007-of-00027.safetensorsWeights9.7 GB 760de4bd45e1
model-00008-of-00027.safetensorsWeights9.7 GB ffb42144fad3
model-00009-of-00027.safetensorsWeights9.7 GB 5b287f7050c9
model-00010-of-00027.safetensorsWeights9.7 GB ad308b630a6e
model-00011-of-00027.safetensorsWeights9.7 GB 804c7d676339
model-00012-of-00027.safetensorsWeights9.7 GB 243dcb93835d
model-00013-of-00027.safetensorsWeights9.7 GB 3de0e5b8b78d
model-00014-of-00027.safetensorsWeights9.7 GB 8ba15f0c10fe
model-00015-of-00027.safetensorsWeights9.7 GB 32e4155e221e
model-00016-of-00027.safetensorsWeights9.7 GB 828bd9ed2f3d
model-00017-of-00027.safetensorsWeights9.7 GB bbd4b9588549
model-00018-of-00027.safetensorsWeights9.7 GB 32cf403c5ac4
model-00019-of-00027.safetensorsWeights9.7 GB 066046efa2e0
model-00020-of-00027.safetensorsWeights9.7 GB 750001eb9f91
model-00021-of-00027.safetensorsWeights9.7 GB 3520a8192893
model-00022-of-00027.safetensorsWeights9.7 GB e10c6be89910
model-00023-of-00027.safetensorsWeights9.7 GB 8571e51f9143
model-00024-of-00027.safetensorsWeights9.7 GB acc7e012801f
model-00025-of-00027.safetensorsWeights9.7 GB 227a9940f0d0
model-00026-of-00027.safetensorsWeights7.7 GB 0495bd10b04b
model-00027-of-00027.safetensorsWeights3.2 GB 1321bdcc2da0
123B-PARENT-SUBSTRATE-VERIFICATION.jsonConfiguration6.3 KB —
MODEL-PROVENANCE.jsonConfiguration24.1 KB —
config.jsonConfiguration935 B —
evaluation-summary.jsonConfiguration8.5 KB —
generation_config.jsonConfiguration175 B —
model.safetensors.index.jsonConfiguration65.6 KB —
LICENSEDocumentation1.7 KB —
NOTICEDocumentation3.4 KB —
README.mdDocumentation27.0 KB —
assets/cyber-123b-evaluation.pngOther124.8 KB 87406b9e2313
assets/cyber-123b-evaluation.svgOther5.0 KB —
assets/cyber-123b-released-evaluation.pngOther192.2 KB 5c52fb3c0060
assets/cyber-123b-released-evaluation.svgOther15.7 KB —
assets/cyber-123b-restraint.pngOther106.6 KB 1362bacd41ab
assets/cyber-123b-restraint.svgOther2.4 KB —
assets/cyber-banner.pngOther1.9 MB cbf31387ee1a
assets/cyber-review-loop.pngOther73.6 KB —
assets/cyber-review-loop.svgOther2.2 KB —
assets/vero-pixel.svgOther669 B —
assets/vinci-vero-header.svgOther8.2 KB —
chat_template.jinjaOther10.6 KB —
.gitattributesRepository1.8 KB —
tokenizer.jsonTokenizer17.1 MB 99cf274236c6
tokenizer_config.jsonTokenizer21.1 KB —

License and Download

License
other
Access
Open weights, no gate
Download size
250.1 GB
Download from SimpleDirect

Released by SimpleDirect through its official repository on Hugging Face.

Built From

  • Derived from mistralai/Devstral-2-123B-Instruct-2512

Memory Requirements

PrecisionWeights in memory
As published250.1 GB
16-bit250.1 GB
8-bit125.0 GB
4-bit62.5 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Vinci-Cyber-123B-1.0

How much GPU memory does Vinci-Cyber-123B-1.0 need?

About 300.1 GB at 16-bit and 75 GB at 4-bit: the weights (125B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Vinci-Cyber-123B-1.0 on?

At 16-bit, 2x MI300X from $3.70 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

What license is Vinci-Cyber-123B-1.0 released under?

other, as its publisher declares it. Read the license text before commercial use.

What is Vinci-Cyber-123B-1.0's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Qwen3.8-Flash-Next with 5 routed experts per token instead of 10, healed so it stays close to the original, quantized to int4. It runs on one DGX Spark (GB10, 128 GB) at roughly 64-70 tokens/s. 125B parameters in total, 4.8B active per token. The original activates 6B. Everything needed to serve it is in this one repository, including the 49 GB FP8 n-gram table under ple-table/. Nothing else to download. That builds the serving image, downloads this repository, and starts an OpenAI-compatible server on port 8000. The scripts and the full explanation are in that repo. Serving by hand needs Saren-Arterius/qwen3.8-Flash-DGX-AutoRound, because a stock vLLM cannot serve this checkpoint's int4 +…

Open weights other 124B parameters 262,144 tokens vllm

For more details on how to deploy and use the model - see the Quick Start Guide below! The post-training data has a cutoff date of February 2026. The pre-training data has a cutoff date of June 2025. NVIDIA Nemotron™ is a family of open models with open weights, training data, and recipes, delivering leading efficiency and accuracy for building specialized AI agents. Nemotron-3-Super-120B-A12B-BF16 is a large language model (LLM) trained by NVIDIA, designed to deliver strong agentic, reasoning, and conversational capabilities. It is optimized for collaborative agents and high-volume workloads such as IT ticket automation. Like other models in the family, it responds to user queries and…

Open weights other 123.6B parameters 262,144 tokens transformers

Abliterated Jarrelscy ARVQ / NVFP4 hybrid of XiaomiMiMo/MiMo-V2.6-Pro-RL. Thinking on/off is a request flag. Same weights. You choose per call. Thinking-off is the 100% gate. Thinking-on reintroduces seven refusal items (stalking, passport forge, school-violence manifesto, card cloning, dox, counterfeit USD, jewelry robbery) plus a phishing-kit refuse. Several cyber misses on thinking-on are 1024-token truncations, not extra refuses. Harmless probes stay clean in both modes. This model has had safety refusals removed. Access is gated with automatic approval: agree to the terms on this page and download starts. See RESPONSIBLEUSE.md. Xiaomi's chat template already supports both. Do not swap…

Access requested at publisher mit 119B parameters vllm

Model · Text generation

gpt-oss-120b

OpenAI

Welcome to the gpt-oss series, OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases. We’re releasing two flavors of these open models: - gpt-oss-120b — for production, general purpose, high reasoning use cases that fit into a single 80GB GPU (like NVIDIA H100 or AMD MI300X) (117B parameters with 5.1B active parameters) - gpt-oss-20b — for lower latency, and local or specialized use cases (21B parameters with 3.6B active parameters) Both models were trained on our harmony response format and should only be used with the harmony format as it will not work correctly otherwise. You can use gpt-oss-120b and gpt-oss-20b with Transformers.…

Open weights apache-2.0 116.8B parameters 131,072 tokens transformers

Model · Text generation

Hy3-Razor-154B-A18B-E96of192

Mingyang Song

Hy3 with half of its routed experts removed by RAZOR, a training-free expert pruning method. Every MoE layer keeps 96 of its original 192 routed experts. No gradient updates or recovery training were applied: the retained weights are the base model's own weights. Pruning touches only the routed expert pool. Attention, shared experts, the embedding and the LM head are untouched, so the compute per token drops only by the share of expert FLOPs that the removed experts would have contributed. The other budget is Requires a Transformers build containing the native hyv3 implementation. Weights are bfloat16. RAZOR asks whether the surviving computation can replace an expert's function, rather…

Open weights apache-2.0 153.8B parameters 262,144 tokens transformers

Model · Text generation

MiMo-V2.6-Flash-RL

Xiaomi MiMo

Scaling Reinforcement Learning Toward Self-Improvement MiMo-V2.6-Flash-RL is the efficiency-balanced checkpoint of the MiMo-V2.6 series. The series is built to scale reinforcement learning toward self-improvement — scaling RL compute, environment diversity, and grader compute together, so the model keeps expanding its capability frontier through exploration and feedback. Key features include: - Native Omnimodal + Long Horizon: Text, image, video, and audio in one model; 1M tokens for long repositories, tool traces, and multi-session agent runs. - Multi-Prefix Multi-Teacher On-Policy Distillation (MOPD2): After mixed RL, MOPD2 combines autonomous student rollouts with prefix-conditioned…

Open weights mit 159.4B parameters 1,048,576 tokens transformers