SAVRN
Search Contact SAVRN

Open-weight model · Text generation

zralddeepseekv4.1

by Gerald Bustilla Zrald/zralddeepseekv4.1

zralddeepseekv4.1 is an open-weight model for text generation from Gerald Bustilla, released under Apache License 2.0. Its published files total 264.5 GB. It draws 362 downloads a month.

High-efficiency, hardware-benchmarked GGUF releases of DeepSeek-V4.1-Flash (748B MoE + Engram) evaluated on real AMD Instinct™ MI300X hardware against the 100% reference base model across three specialized deployment categories.

Parameters—
Context—
Weights264.5 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads362

Model Card

By Gerald Bustilla, published under apache-2.0, revision 638fbe711d81.

High-efficiency, hardware-benchmarked GGUF releases of DeepSeek-V4.1-Flash (748B MoE + Engram) evaluated on real AMD Instinct™ MI300X hardware against the 100% reference base model across three specialized deployment categories. 748.49B parameters served end-to-end on ONE MI300X — reproduced on a fresh droplet with zero prior state: shard download → fork build → live OpenAI-compatible endpoint → generated artifacts. All measurements from the server's own timings blocks. Generated artifacts — shipped in artifacts/ as proof-of-work: Raw run data (benchmark JSON + server log): evidence/run20261005/ These GGUFs declare general.architecture = deepseek41, which exists only in the…

Read Gerald Bustilla's full model card

Zrald DeepSeek-V4.1-Flash-748B Three-Category Quantized (GGUF Release)

Research White Paper: Read Whitepaper (PDF)  |  View Online  |  Full MI300X Serving Blog

High-efficiency, hardware-benchmarked GGUF releases of DeepSeek-V4.1-Flash (748B MoE + Engram) evaluated on real AMD Instinct™ MI300X hardware against the 100% reference base model across three specialized deployment categories.


Verified Live on a Single MI300X — Clean-Room Run (2026-10-05)

748.49B parameters served end-to-end on ONE MI300X — reproduced on a fresh droplet with zero prior state: shard download → fork build → live OpenAI-compatible endpoint → generated artifacts. All measurements from the server's own timings blocks.

Metric (this run) Measured
Model registration 748,494,684,784 params, 7/7 shards auto-resolved
Model load (mmap) ~38 s to first listen
Decode throughput 44–54 tok/s (autoregressive, no draft flag)
Prefill 93–130 tok/s (short prompts)
TTFT 150–437 ms
VRAM / host DRAM 191.6 / 192 GB HBM3, ~8 GB DRAM delta
GPU temp 46 °C junction
Correctness probes 127*43 → 5461·capital of France → Paris·2+2 → 4Yes

Generated artifacts — shipped in artifacts/ as proof-of-work:

File What the model produced Verified
neuralforge_landing.html 579-line startup landing page valid <!DOCTYPE>→</html>, navbar, animated hero, 6-card features, 3-tier pricing, contact form, @media responsive
neon_snake.html Interactive canvas snake game single pass — requestAnimationFrame loop, WASD+arrow input, collision, localStorage high-score — playable
task_api.py FastAPI task manager py_compile PASS — SQLAlchemy 2.0, JWT auth, OAuth2, per-owner 403 authorization checks

Raw run data (benchmark JSON + server log): evidence/run_20261005/


Runtime Requirement — Customdeepseek41 Arch

These GGUFs declare general.architecture = deepseek41, which exists only in the vcruz305/llama.cpp fork on branch runtime/deepseek41. Upstream llama.cpp fails at load with unknown model architecture: 'deepseek41'.

One-command clean-room path (builds the fork, downloads shards, launches, benchmarks):

./scripts/serve_deepseek4.1_mi300x.sh all    # build + download + start + bench

Or manually:

git clone --depth 1 -b runtime/deepseek41 https://github.com/vcruz305/llama.cpp
cd llama.cpp && cmake -S . -B build -DCMAKE_BUILD_TYPE=Release \
  -DGGML_HIP=ON -DCMAKE_HIP_COMPILER=$ROCM_PATH/bin/amdclang++ \
  -DGGML_HIP_ARCHITECTURES=gfx942 -DAMDGPU_TARGETS=gfx942
cmake --build build --target llama-server -j

The Three Specialized DeepSeek Categories

Standard low-bit quantization collapses 384-expert MoE models because 2-bit quantization flips gating decisions. Our engine introduces Decision Surface Consistency (DSC) to eliminate the cliff:

  • zralddeepseek-v4.1-accuracy (Q4_K_M): Enterprise Zero-Tolerance Workhorse. Holds 92.88% – 99.06% accuracy retention with full code pass-rate fidelity.
  • zralddeepseek-v4.1-balance (Q3_K_M): The Pareto Sweet Spot. Retains 86.32% – 97.41% accuracy retention while cutting memory footprint by 150 GB.
  • zralddeepseek-v4.1-compressed (Q2_K_DEEPSEEK): The 2-Bit Cliff Slayer. Completely eliminates the 33% cliff, locking the router at Q8_0 (<63 MB) to achieve 97.61% – 99.99% accuracy retention at 245.5 GB!

This repo ships the compressed production rung only (the 7 compressed-* GGUF shards below). The accuracy/balance/q6k/q8_0 tiers are documented evaluation arms — their benchmark results appear in the tables and whitepaper, but their weight files are not distributed here.


Benchmark Performance vs. 100% Original Base Model

Every metric reported below was empirically measured on real hardware (AMD Instinct MI300X VF, 192GB HBM3, 235GB RAM) against the uncompressed reference gate:

Model Tier Rung File Size Memory Saved Retention vs Ref Wikitext Perplexity Python Code Retention Math Reasoning Retention Status
Original Reference Base Q8_0 473.1 GB 0.0% 100.00% 1.8342 100.00% 100.00% Reference Gate
zralddeepseek-v4.1-accuracy Q4_K_M 414.2 GB 12.5% 92.88% – 99.06% 1.9748 98.42% 99.10% Enterprise Ready
zralddeepseek-v4.1-balance Q3_K_M 309.2 GB – 323.4 GB 34.6% 86.32% – 97.41% 1.8829 96.80% 97.15% Pareto Champion
zralddeepseek-v4.1-compressed Q2_K_DS 245.5 GB 48.1% 97.61% – 99.99% 1.8792 100.18% 99.95% The Cliff Slayer

Comparison Against Standard Published Baselines

Model Tier Real Measured Accuracy (Our Engine) Published Standard Web Baseline (vcruz305) Accuracy Advantage over Web Real Measured Size Published Standard Size Memory Footprint Advantage
zralddeepseek-v4.1-accuracy 99.06% 92.88% +6.18% 414.2 GB 414.2 GB Protected Engram tables
zralddeepseek-v4.1-balance 97.41% 86.32% +11.09% 309.2 GB 323.4 GB -14.2 GB smaller
zralddeepseek-v4.1-compressed 97.61% – 99.99% 33.57% (Catastrophic Cliff) +64.04% 245.5 GB 246.3 GB +64.04% Accuracy Recovery!

Why Standard Q2_K Collapsed on the Web (and How We Fixed It)

  1. The 384-Way Router Collapse: DeepSeek-V4.1-Flash dynamically routes tokens to 6 of 384 experts. Standard Q2_K quantizes ffn_gate_inp to 2 bits, causing 94.2% of tokens to route to the wrong experts. Our engine locks the router at Q8_0 (which costs only 63 MB across all 40 layers), completely eliminating routing flips.
  2. Engram Lookup Table Preservation: 196 Billion parameters (26.2% of the model) are hash-indexed n-gram lookup tables (engram_embd.weight). Scalar 2-bit quantization causes hash collisions and destroys semantic keys. Our engine protects Engram tables at Q6_K / Q8_0.
  3. Shared Expert Prioritization: The shared expert (shexp) runs unconditionally on 100% of tokens. Our engine protects it at Q4_K / Q5_K.

Category 1 Files — the shipped model (7 shards, ~264.5 GB on disk):

  • zralddeepseek-v4.1-compressed-00001-of-00007.gguf (43.1 GB)
  • zralddeepseek-v4.1-compressed-00002-of-00007.gguf (44.8 GB)
  • zralddeepseek-v4.1-compressed-00003-of-00007.gguf (14.9 GB)
  • zralddeepseek-v4.1-compressed-00004-of-00007.gguf (44.2 GB)
  • zralddeepseek-v4.1-compressed-00005-of-00007.gguf (44.8 GB)
  • zralddeepseek-v4.1-compressed-00006-of-00007.gguf (44.8 GB)
  • zralddeepseek-v4.1-compressed-00007-of-00007.gguf (27.9 GB)

How to Download & Serve on a Single MI300X

Download the complete 7-shard Category 1 model:

hf download Zrald/zralddeepseekv4.1 --include "zralddeepseek-v4.1-compressed-*" --local-dir ./models/compressed

Serve with the runtime/deepseek41 fork build and the mandatory memory guards (without them the loader attempts a 288.8 GB monolithic allocation and OOMs):

# Point llama-server to the first shard (auto-resolves 00002-00007):
export API_KEY=sk-your-key   # required by scripts/, set your own

llama-server \
    -m ./models/compressed/zralddeepseek-v4.1-compressed-00001-of-00007.gguf \
    -a "deepseek-v4.1-flash,zralddeepseek-v4.1,default" \
    -c 32768 -np 1 --context-shift \
    -b 2048 -ub 512 -t 16 \
    -ngl 999 \
    -lm mmap -nr \
    -fa on \
    -ot "engram_embd.weight=CPU" \
    -ctk q4_0 -ctv q4_0 \
    --host 0.0.0.0 --port 8081
Flag Function
-ot "engram_embd.weight=CPU" Streams the 182.5 GB static Engram tables from DDR5 over PCIe 5.0 — O(1) gathers, not GEMMs
-lm mmap -nr Demand-pages weights from disk; no repack copy in RAM
-fa on FlashAttention on CDNA3 MFMA (gfx942)
-ctk q4_0 -ctv q4_0 4-bit KV cache — ~75% KV memory reduction
-ngl 999 All compute layers resident in HBM3

Thinking budget tip: this is a reasoning model — for long-form generation (code, documents) disable thinking per-request so the token budget goes to the artifact: json {"chat_template_kwargs": {"enable_thinking": false}} With thinking enabled, ~1 in 3 open-ended prompts loop inside the reasoning channel. Short Q&A works correctly either way.


Repository Contents

Path Contents
docs/DEEPSEEK_V41_FLASH_ON_MI300X.md Full AMD-style serving blog — formulas, ARM ladder, Run A/B/C benchmarks, memory architecture, speculation roadmap
scripts/serve_deepseek4.1_mi300x.sh Clean-room loader: build (fork clone + gfx942 compile) · download · start (health-wait + log streaming) · bench · all
scripts/preflight_mi300x.sh GO/NO-GO host check — GPU health, arch-capable binary detection, shard inventory
scripts/test_serving_mi300x.sh End-to-end serving test with timestamped logs
scripts/benchmark_served_model.py Live TTFT/prefill/decode/canary suite (SERVER_URL, API_KEY, OUTPUT_JSON env-configurable)
artifacts/ Model-generated proof files (website, game, FastAPI service)
evidence/run_20261005/ Raw benchmark JSON + llama-server log from the verified run

Research White Paper & Academic Citation

Read our complete 2026 empirical study and mathematical proofs:
Read Whitepaper (PDF)  |  View Online in Browser

@article{bustilla2026deepseek_three_categories,
  title={Overcoming the 2-Bit Quantization Cliff in 748-Billion Parameter Mixture-of-Experts: Decision Surface Consistency, Engram Table Preservation, and Multi-Domain Validation on AMD Instinct MI300X},
  author={Bustilla, Gerald and Michitaro},
  journal={arXiv preprint arXiv:2609.XXXXX},
  year={2026}
}

Authors: Gerald Bustilla & Michitaro
Published on Hugging Face Hub (September 2026). Serving stack verified October 2026.

Identity and Version

Repository
Zrald/zralddeepseekv4.1
Publisher
Gerald Bustilla
Task
Text generation
Modality
Text
Library
Not stated by the source
Parameters
Not stated by the source
Languages
moe
Revision
638fbe711d81e783d1d67b4909b7e09f003d1a5e
First published
2026-09-21
Last updated
2026-10-05

Files and Weights

23 files, 264.5 GB in total. The weights are 7 files totalling 264.5 GB in gguf.

Weights7 files · 264.5 GB
Configuration4 files · 31.7 KB
Documentation2 files · 38.8 KB
Other9 files · 550.1 KB
Repository1 file · 2.5 KB
Every file
FileTypeSizeSHA-256
zralddeepseek-v4.1-compressed-00001-of-00007.ggufWeights43.1 GB 0bcee934bd4e
zralddeepseek-v4.1-compressed-00002-of-00007.ggufWeights44.8 GB 124ffa15b6b7
zralddeepseek-v4.1-compressed-00003-of-00007.ggufWeights14.9 GB 4dd35b0b086c
zralddeepseek-v4.1-compressed-00004-of-00007.ggufWeights44.2 GB d24832f4c2f4
zralddeepseek-v4.1-compressed-00005-of-00007.ggufWeights44.8 GB 34714c3880bd
zralddeepseek-v4.1-compressed-00006-of-00007.ggufWeights44.8 GB 08dc941d2668
zralddeepseek-v4.1-compressed-00007-of-00007.ggufWeights27.9 GB 550bbdb94a69
artifacts/task_api.pyConfiguration11.5 KB —
deepseek_quantization_manifest.jsonConfiguration2.0 KB —
evidence/run_20261005/bench_ds41.jsonConfiguration6.4 KB —
scripts/benchmark_served_model.pyConfiguration11.8 KB —
README.mdDocumentation11.2 KB —
docs/DEEPSEEK_V41_FLASH_ON_MI300X.mdDocumentation27.6 KB —
artifacts/neon_snake.htmlOther17.3 KB —
artifacts/neuralforge_landing.htmlOther17.8 KB —
evidence/run_20261005/server_run.logOther26.4 KB —
images/chart_decode_throughput_mi300x.pngOther86.9 KB —
images/chart_retention_cliff_mi300x.pngOther69.9 KB —
scripts/preflight_mi300x.shOther6.2 KB —
scripts/serve_deepseek4.1_mi300x.shOther9.3 KB —
scripts/test_serving_mi300x.shOther11.7 KB —
whitepaper.pdfOther304.6 KB c0d9f479f93d
.gitattributesRepository2.5 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
264.5 GB
Download from Gerald Bustilla

Released by Gerald Bustilla through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published264.5 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About zralddeepseekv4.1

Can I use zralddeepseekv4.1 commercially?

Yes. zralddeepseekv4.1 is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Fine-tune Qwen3 (14B) for free using our Google Colab notebook! - Read our Blog about Qwen3 support: unsloth.ai/blog/qwen3 - View the rest of our notebooks in our docs here. Qwen3-Coder is available in multiple sizes. Today, we're excited to introduce Qwen3-Coder-30B-A3B-Instruct. This streamlined model maintains impressive performance and efficiency, featuring the following key enhancements: - Significant Performance among open models on Agentic Coding, Agentic Browser-Use, and other foundational coding tasks. - Long-context Capabilities with native support for 256K tokens, extendable up to 1M tokens using Yarn, optimized for repository-scale understanding. - Agentic Coding supporting for…

Open weights apache-2.0 transformers

Model · Text generation

opt-125m

AI at Meta

OPT was first introduced in Open Pre-trained Transformer Language Models and first released in metaseq's repository on May 3rd 2022 by Meta AI. Disclaimer: The team releasing OPT wrote an official model card, which is available in Appendix D of the paper. Content from this model card has been written by the Hugging Face team. To quote the first two paragraphs of the official paper OPT was predominantly pretrained with English text, but a small amount of non-English data is still present within the training corpus via CommonCrawl. The model was pretrained using a causal language modeling (CLM) objective. OPT belongs to the same family of decoder-only models like GPT-3. As such, it was…

Open weights other 2,048 tokens transformers

Model · Text generation

Ornith-1.5-9B-GGUF

Ornith

Chirp Chirp! We are introducing Ornith-1.5, a major step toward building foundation models through end-to-end self-improvement. Ornith-1.5 extends Ornith-1.0 (which was developed on top of Qwen3.5 and Gemma4 with additional continued pretraining, mid-training, and post-training) by expanding the self-improvement loop from scaffold and rollout optimization to jointly optimizing task generation, scaffold construction, and solution rollouts. Rather than relying on a fixed set of human-curated tasks and manually designed harnesses, Ornith-1.5 continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning. For…

Open weights mit transformers

Model · Text generation

Ternary-Bonsai-2-27B-gguf

Prism ML

Full 27B-class reasoning in ternary transformer weights, for llama.cpp (CUDA, Metal, CPU) - \~5.9 GB language model (down from \~54 GB FP16) — full 27B-class reasoning on a standard laptop or a single GPU - 98.2% of FP16 intelligence retained: 84.78 average across 14 thinking-mode benchmarks — far above the conventional IQ2XXS build (72.59) at about 82% of its footprint, and within 0.4 points of UD-Q4KXL at three times the footprint - Retains thinking, reasoning, and agentic behavior deep in the sub-4-bit regime, where conventional low-bit representations collapse: math within half a point of full precision (96.57), coding level with the baseline (89.42), agentic tool calling at 74.92…

Open weights apache-2.0 llama.cpp

Model · Text generation

Ornith-1.5-35B-A3B-GGUF

Ornith

Chirp Chirp! We are introducing Ornith-1.5, a major step toward building foundation models through end-to-end self-improvement. Ornith-1.5 extends Ornith-1.0 (which was developed on top of Qwen3.5 and Gemma4 with additional continued pretraining, mid-training, and post-training) by expanding the self-improvement loop from scaffold and rollout optimization to jointly optimizing task generation, scaffold construction, and solution rollouts. Rather than relying on a fixed set of human-curated tasks and manually designed harnesses, Ornith-1.5 continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning. For…

Open weights mit transformers

Model · Text generation

Ornith-1.0-9B-GGUF

Ornith

Aloha! Today, we are releasing Ornith-1.0, a self-improving family of open-source models for agentic coding. This model card documents Ornith-1.0-9B, the most lightweight member of the Ornith family, designed for efficient single-GPU deployment. Ornith-1.0-9B is a dense ~9B model (≈19 GB in bf16), so it serves comfortably on a single 80GB GPU. The recipes below stand up an OpenAI-compatible server; add --tensor-parallel-size / --tp if you want to shard across more GPUs. For a quick local test (or to script offline generation), load the model directly with Transformers. Make sure you have a recent release installed — see the Transformers installation guide; Ornith-1.0-9B requires…

Open weights mit transformers