Qwen3.8-Flash-Next with 5 routed experts per token instead of 10, healed so it stays close to the original, quantized to int4. It runs on one DGX Spark (GB10, 128 GB) at roughly 64-70 tokens/s. 125B parameters in total, 4.8B active per token. The original activates 6B. Everything needed to serve it is in this one repository, including the 49 GB FP8 n-gram table under ple-table/. Nothing else to download. That builds the serving image, downloads this repository, and starts an OpenAI-compatible server on port 8000. The scripts and the full explanation are in that repo. Serving by hand needs Saren-Arterius/qwen3.8-Flash-DGX-AutoRound, because a stock vLLM cannot serve this checkpoint's int4 +…
Open-weight model · Text generation
Vinci-Cyber-123B-1.0
by SimpleDirect simpledirect/Vinci-Cyber-123B-1.0
Vinci-Cyber-123B-1.0 is an open-weight model for text generation from SimpleDirect, released under other. It has 125B parameters and a 262,144-token context. At 16-bit it needs about 300.1 GB of GPU memory, which fits on 2x MI300X from $3.70 an hour; at 4-bit, 75 GB on 1x MI300X from $1.85, at the lowest prices in the SAVRN Index. It draws 1.5k downloads a month.
Vinci Cyber 123B 1.0 is an open-weight model for defensive infrastructure review and targeted remediation, fine-tuned in Canada from Mistral AI's Devstral 2 123B.
Runs On
What it takes to serve Vinci-Cyber-123B-1.0 (125B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 250.1 GB | 300.1 GB | 2x MI300X (192 GB) Vultr |
$3.70 | 2x MI325X $4.00 · 2x MI355X $5.18 |
| 8-bit | 125.0 GB | 150.0 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x MI325X $2.00 · 1x MI355X $2.59 |
| 4-bit | 62.5 GB | 75.0 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.
Vinci-Cyber-123B-1.0 on every accelerator the SAVRN Index prices, at every precision
Model Card
Vinci Cyber 123B 1.0 is an open-weight model for defensive infrastructure review and targeted remediation, fine-tuned in Canada from Mistral AI's Devstral 2 123B. The released merged weights have now been tested directly, alongside their parent and the available GGUF formats. Focused repairs. Restraint on correct configuration. Weights you can run yourself. On the V2-B neutral-review test, the released BF16 model preserved 24/24 correct configurations and produced 18/24 scanner-credited repairs, all 18 passing offline provider-schema validation. Its parent repaired 17/24 and preserved 0/24. On the second set, V2-A, Cyber again preserved 24/24, but repaired 9/24 versus the parent's 15/24.…
Excerpt from the card by SimpleDirect, licensed other.
Configuration
- Architecture
- Ministral3ForCausalLM
- Context length (tokens)
- 262,144
- Layers
- 88
- Hidden size
- 12,288
- Feed-forward size
- 28,672
- Attention heads
- 96
- Key/value heads
- 8
- Head dimension
- 128
- Vocabulary size
- 131,072
- Stored precision
- bfloat16
- Model type
- ministral3
Identity and Version
- Repository
- simpledirect/Vinci-Cyber-123B-1.0
- Publisher
- SimpleDirect
- Task
- Text generation
- Modality
- Text
- Library
- transformers
- Parameters
- 125B parameters
- Languages
- Not stated by the source
- Revision
- 1c2b2438721fb08e0df19ffff4d02baabdbeaf9b
- First published
- 2026-09-22
- Last updated
- 2026-09-28
Files and Weights
51 files, 250.1 GB in total. The weights are 27 files totalling 250.1 GB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model-00001-of-00027.safetensors | Weights | 6.6 GB | 1fb4bba6e7a3 |
| model-00002-of-00027.safetensors | Weights | 9.7 GB | bab70bd5d9ca |
| model-00003-of-00027.safetensors | Weights | 9.7 GB | 767a6c1c452c |
| model-00004-of-00027.safetensors | Weights | 9.7 GB | 462e772c9e68 |
| model-00005-of-00027.safetensors | Weights | 9.7 GB | 4ca452feae22 |
| model-00006-of-00027.safetensors | Weights | 9.7 GB | 6f53db08910b |
| model-00007-of-00027.safetensors | Weights | 9.7 GB | 760de4bd45e1 |
| model-00008-of-00027.safetensors | Weights | 9.7 GB | ffb42144fad3 |
| model-00009-of-00027.safetensors | Weights | 9.7 GB | 5b287f7050c9 |
| model-00010-of-00027.safetensors | Weights | 9.7 GB | ad308b630a6e |
| model-00011-of-00027.safetensors | Weights | 9.7 GB | 804c7d676339 |
| model-00012-of-00027.safetensors | Weights | 9.7 GB | 243dcb93835d |
| model-00013-of-00027.safetensors | Weights | 9.7 GB | 3de0e5b8b78d |
| model-00014-of-00027.safetensors | Weights | 9.7 GB | 8ba15f0c10fe |
| model-00015-of-00027.safetensors | Weights | 9.7 GB | 32e4155e221e |
| model-00016-of-00027.safetensors | Weights | 9.7 GB | 828bd9ed2f3d |
| model-00017-of-00027.safetensors | Weights | 9.7 GB | bbd4b9588549 |
| model-00018-of-00027.safetensors | Weights | 9.7 GB | 32cf403c5ac4 |
| model-00019-of-00027.safetensors | Weights | 9.7 GB | 066046efa2e0 |
| model-00020-of-00027.safetensors | Weights | 9.7 GB | 750001eb9f91 |
| model-00021-of-00027.safetensors | Weights | 9.7 GB | 3520a8192893 |
| model-00022-of-00027.safetensors | Weights | 9.7 GB | e10c6be89910 |
| model-00023-of-00027.safetensors | Weights | 9.7 GB | 8571e51f9143 |
| model-00024-of-00027.safetensors | Weights | 9.7 GB | acc7e012801f |
| model-00025-of-00027.safetensors | Weights | 9.7 GB | 227a9940f0d0 |
| model-00026-of-00027.safetensors | Weights | 7.7 GB | 0495bd10b04b |
| model-00027-of-00027.safetensors | Weights | 3.2 GB | 1321bdcc2da0 |
| 123B-PARENT-SUBSTRATE-VERIFICATION.json | Configuration | 6.3 KB | — |
| MODEL-PROVENANCE.json | Configuration | 24.1 KB | — |
| config.json | Configuration | 935 B | — |
| evaluation-summary.json | Configuration | 8.5 KB | — |
| generation_config.json | Configuration | 175 B | — |
| model.safetensors.index.json | Configuration | 65.6 KB | — |
| LICENSE | Documentation | 1.7 KB | — |
| NOTICE | Documentation | 3.4 KB | — |
| README.md | Documentation | 27.0 KB | — |
| assets/cyber-123b-evaluation.png | Other | 124.8 KB | 87406b9e2313 |
| assets/cyber-123b-evaluation.svg | Other | 5.0 KB | — |
| assets/cyber-123b-released-evaluation.png | Other | 192.2 KB | 5c52fb3c0060 |
| assets/cyber-123b-released-evaluation.svg | Other | 15.7 KB | — |
| assets/cyber-123b-restraint.png | Other | 106.6 KB | 1362bacd41ab |
| assets/cyber-123b-restraint.svg | Other | 2.4 KB | — |
| assets/cyber-banner.png | Other | 1.9 MB | cbf31387ee1a |
| assets/cyber-review-loop.png | Other | 73.6 KB | — |
| assets/cyber-review-loop.svg | Other | 2.2 KB | — |
| assets/vero-pixel.svg | Other | 669 B | — |
| assets/vinci-vero-header.svg | Other | 8.2 KB | — |
| chat_template.jinja | Other | 10.6 KB | — |
| .gitattributes | Repository | 1.8 KB | — |
| tokenizer.json | Tokenizer | 17.1 MB | 99cf274236c6 |
| tokenizer_config.json | Tokenizer | 21.1 KB | — |
License and Download
- License
- other
- Access
- Open weights, no gate
- Download size
- 250.1 GB
Released by SimpleDirect through its official repository on Hugging Face.
Built From
- Derived from mistralai/Devstral-2-123B-Instruct-2512
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 250.1 GB |
| 16-bit | 250.1 GB |
| 8-bit | 125.0 GB |
| 4-bit | 62.5 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About Vinci-Cyber-123B-1.0
How much GPU memory does Vinci-Cyber-123B-1.0 need?
About 300.1 GB at 16-bit and 75 GB at 4-bit: the weights (125B parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run Vinci-Cyber-123B-1.0 on?
At 16-bit, 2x MI300X from $3.70 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
What license is Vinci-Cyber-123B-1.0 released under?
other, as its publisher declares it. Read the license text before commercial use.
What is Vinci-Cyber-123B-1.0's context length?
262,144 tokens, from the maximum position embeddings in its published configuration.
Similar Models
For more details on how to deploy and use the model - see the Quick Start Guide below! The post-training data has a cutoff date of February 2026. The pre-training data has a cutoff date of June 2025. NVIDIA Nemotron™ is a family of open models with open weights, training data, and recipes, delivering leading efficiency and accuracy for building specialized AI agents. Nemotron-3-Super-120B-A12B-BF16 is a large language model (LLM) trained by NVIDIA, designed to deliver strong agentic, reasoning, and conversational capabilities. It is optimized for collaborative agents and high-volume workloads such as IT ticket automation. Like other models in the family, it responds to user queries and…
Abliterated Jarrelscy ARVQ / NVFP4 hybrid of XiaomiMiMo/MiMo-V2.6-Pro-RL. Thinking on/off is a request flag. Same weights. You choose per call. Thinking-off is the 100% gate. Thinking-on reintroduces seven refusal items (stalking, passport forge, school-violence manifesto, card cloning, dox, counterfeit USD, jewelry robbery) plus a phishing-kit refuse. Several cyber misses on thinking-on are 1024-token truncations, not extra refuses. Harmless probes stay clean in both modes. This model has had safety refusals removed. Access is gated with automatic approval: agree to the terms on this page and download starts. See RESPONSIBLEUSE.md. Xiaomi's chat template already supports both. Do not swap…
Welcome to the gpt-oss series, OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases. We’re releasing two flavors of these open models: - gpt-oss-120b — for production, general purpose, high reasoning use cases that fit into a single 80GB GPU (like NVIDIA H100 or AMD MI300X) (117B parameters with 5.1B active parameters) - gpt-oss-20b — for lower latency, and local or specialized use cases (21B parameters with 3.6B active parameters) Both models were trained on our harmony response format and should only be used with the harmony format as it will not work correctly otherwise. You can use gpt-oss-120b and gpt-oss-20b with Transformers.…
Hy3 with half of its routed experts removed by RAZOR, a training-free expert pruning method. Every MoE layer keeps 96 of its original 192 routed experts. No gradient updates or recovery training were applied: the retained weights are the base model's own weights. Pruning touches only the routed expert pool. Attention, shared experts, the embedding and the LM head are untouched, so the compute per token drops only by the share of expert FLOPs that the removed experts would have contributed. The other budget is Requires a Transformers build containing the native hyv3 implementation. Weights are bfloat16. RAZOR asks whether the surviving computation can replace an expert's function, rather…
Scaling Reinforcement Learning Toward Self-Improvement MiMo-V2.6-Flash-RL is the efficiency-balanced checkpoint of the MiMo-V2.6 series. The series is built to scale reinforcement learning toward self-improvement — scaling RL compute, environment diversity, and grader compute together, so the model keeps expanding its capability frontier through exploration and feedback. Key features include: - Native Omnimodal + Long Horizon: Text, image, video, and audio in one model; 1M tokens for long repositories, tool traces, and multi-session agent runs. - Multi-Prefix Multi-Teacher On-Policy Distillation (MOPD2): After mixed RL, MOPD2 combines autonomous student rollouts with prefix-conditioned…