SAVRN
Search Contact SAVRN

Open-weight model · Text generation

YoungAi-DeepSeek-V4.1-Flash

by Wenzhou Wu wenzhouwu/YoungAi-DeepSeek-V4.1-Flash

YoungAi-DeepSeek-V4.1-Flash is an open-weight model for text generation from Wenzhou Wu, released under MIT License. Its published files total 11.4 GB.

English · 中文 ds4 is a native C/CUDA inference engine plus an offline toolchain that runs DeepSeek V4.1 Flash on a single NVIDIA DGX Spark (GB10, 128 GB unified memory).

Parameters—
Context—
Weights41.8 MB
Licensemit
AccessOpen weights
Monthly Downloads—

Model Card

By Wenzhou Wu, published under mit, revision 9852affc876b.

English · 中文 ds4 is a native C/CUDA inference engine plus an offline toolchain that runs DeepSeek V4.1 Flash on a single NVIDIA DGX Spark (GB10, 128 GB unified memory). Everything described here is our own work: the vector-quantization format, the solvers that produce the sidecar and the post-training file, the CUDA kernels, and the rulers we judge all of it with. The model architecture is DeepSeek's and is not re-explained here — read the official release for that. How to read this page. It unfolds in five steps; stop wherever you have what you need. 1. The numbers — what runs, how fast, how close to the original. 2. Three convictions — why the system is shaped this way. 3. The…

Read Wenzhou Wu's full model card

YoungAi — DeepSeek V4.1 Flash on one DGX Spark

English · 中文

One 128 GB box. Three files. A model whose official checkpoint is 510 GB.

① a 113.6 GB general base → ② a ~40 MB domain sidecar → ③ a post-training file re-solved every night.

ds4 is a native C/CUDA inference engine plus an offline toolchain that runs DeepSeek V4.1 Flash on a single NVIDIA DGX Spark (GB10, 128 GB unified memory). Everything described here is our own work: the vector-quantization format, the solvers that produce the sidecar and the post-training file, the CUDA kernels, and the rulers we judge all of it with. The model architecture is DeepSeek's and is not re-explained here — read the official release for that.

How to read this page. It unfolds in five steps; stop wherever you have what you need.

  1. The numbers — what runs, how fast, how close to the original.
  2. Three convictions — why the system is shaped this way.
  3. The architecture — three files and how they stack.
  4. The algorithms — each piece goes one sentence → intuition → math → engineering → evidence.
  5. Results, what did not work, how to run it, limits.

1. The numbers

Measured on one DGX Spark with the deployed pair ① base + ② finance sidecar, 2026-09-21 → 09-24.

Resident in memory 113.6 GB model + ~40 MB sidecar (≈ 1.6 bits per weight over the whole file)
Left on SSD, untouched the model's 203 GB n-gram memory tables, original FP8, ~6 KB read per token
Restore rate Σmin — finance / general English 0.745 / 0.790
Same top-1 as the original — finance / general English 74.6% / 81.1%
Prefill, 12.5k-token prompt 489 tokens/s
Decode, plain greedy 30.7 tokens/s at short context · 29 tokens/s at 14k context
Decode, speculative (on by default at temperature 0) 43.0 tokens/s on a real 14k-token agent request · 37.1 tokens/s on a request never used for tuning
Context 1M tokens, taken from the model's own metadata; the full 1M KV cache is 0.89 GB

Restore rate = Σ_v min(p_original(v), p_ours(v)), averaged over positions: the share of next-token probability mass on which our model and the original agree. 1.0 means identical distributions. All rulers are defined in §4.5.


2. Three convictions

Each conviction below is stated first, then backed by what we measured, then tied to what we built.

2.1 Every domain hides a strongly correlated regularity — find it, and a few megabytes restore a lot

Claim. From the point of view of one domain, quantization damage is not uniform noise. Text from one domain (Chinese finance, in our case) exercises a narrow set of directions, and the error along those directions can be learned from a few thousand tokens. Fixing it restores more capability per byte than adding bits to the weights ever could.

What we measured.

  • Same-domain vs mixed-domain. A correction fitted and judged on finance removed 15.8% of the held-out layer error; the identical procedure on an eight-domain mix removed 2.3% — seven times less. Our earlier DeepSeek V4 work showed the same diagonal: a domain's own correction recovered 28.5–51.6% on its own text versus 10–12% for a Wikipedia proxy.
  • Transfer without harm. The first full sidecar, fitted on 8,192 finance tokens, lifted Same top-1 by +3.1 pp on finance, +2.3 pp on the eight-domain mix, +1.6 pp on English Wikipedia. It helps its own domain most and does not hurt general text.
  • Leverage. +3.1 pp is what roughly 0.15 extra bits per weight would buy on our bit-width curve — about 11 GB of extra expert storage. The sidecar that delivered it is 38.6 MB.

What we built. A base that never sees domain data, plus a small per-domain sidecar (§4.2).

2.2 Scaling laws are real, and knowledge is spread like a hash

Claim. Capability follows the bytes you keep. You cannot cheat the scaling law by deciding that some experts, layers or matrices "matter more" — the weights show no such structure, and routing behaves like a hash of the token. Hardware has diminishing returns too: a second box buys less than it costs.

What we measured.

  • The weights have no favourites. On the real checkpoint, the 384 experts of a layer have the same relative quantization error to within 1.9%; expert energy differs by at most 1.25×; layers differ by 3% in Σw²; the three matrices of an expert split 33.5 / 33.3 / 33.2. Moving one bit from group B to group A pays only when A is at least 1.189× more important, because each extra bit per 8 weights multiplies the error by 1/1.189 (measured 0.201 / 0.171 / 0.143 at 11 / 12 / 13 bits). Nothing crosses that line, so the optimal "dynamic" allocation is simply uniform. We had built dynamic per-layer and per-expert allocation earlier; it bought nothing.
  • Routing is hash-like. After quantization, the share of tokens whose 6 chosen experts exactly match the original drops to 1.6–28% in the middle and deep layers, and a per-expert systematic bias explains only 3–22% of the flips — the rest is per token. Consecutive tokens barely share experts: 1, 2, 3, 4, 5 tokens touch 6, 10.0, 13.3, 17.4, 21.7 distinct experts.
  • Pruning is a domain decision in disguise. The only 20–80× skew we found is how often experts are routed on a given corpus: change the corpus and the skew changes. Fitting the model into 100 GB by pruning would have meant keeping 130 of 384 experts per layer.
  • More bits, diminishing returns (experts only, English Wikipedia ruler): 1.0 / 1.5 / 2.0 / 2.5 / 3.0 bits per weight → Σmin 0.445 / 0.746 / 0.848 / 0.900 / 0.927. The last half bit buys +0.027.
  • More boxes, diminishing returns. Decoding is bound by the bytes read per token. Splitting layers across machines adds a network hop to every token (upstream ds4 measured −19% on two Macs); tensor parallelism halves the bytes, but our two-machine test reached only 1.25× because the all-reduce ate half the gain. For scale: a community two-Spark setup runs the smaller V4 Flash (284B) in FP8 with TP=2 and speculative decoding at ~41 tokens/s; one Spark here runs V4.1 at 37–43.

What we built. One box. Every byte spent uniformly. Everything else comes from the sidecar and from post-training, not from topology.

2.3 Real work needs post-training — and it has to stay on your machine

Claim. On real business work — a real codebase, a real trading desk — open models are not good enough out of the box, and quantization fidelity cannot fix that: a perfect copy of the original inherits the original's mistakes. The model has to learn from its own deployment, and that data (positions, reviews, source code) should never leave the machine.

What we saw. Our test bed is a trading-agent pipeline served by ds4: market outlook, news merging, stock picking and CFO decision reports, with prompts from 9k to 132k tokens. In its logs, 12 of 12 target prices the model set were above the next day's actual high. On one market-outlook request the model reasoned its way to "market up"; that day 1,891 stocks rose and 3,563 fell.

What we built. A third file, solved on the machine from the day's own requests and stacked on top without touching ① or ② (§4.3). The mechanism works; it does not yet generalize across days — see §8.


3. The architecture: base + sidecar + post-training file

                 what it is                          how often it changes
 ① base GGUF     113.6 GB, general, zero corpus      quantized once from the official weights
 ② sidecar       ~40 MB per domain   (--zchain)      solved once per domain, then frozen
 ③ post-train    same format as ②    (--posttrain)   re-solved nightly; delete it to roll back
 ──────────────────────────────────────────────────────────────────────────────────────────
 each routed-expert weight row:   g_eff = g_base × s_sidecar × s_posttrain
 each router:                     selection score += Δb_sidecar
  • ① alone is a complete model. ② and ③ only rescale rows that already exist and nudge the router. They add no layers; the kernels are the same with or without them.
  • Near-zero runtime cost. ② adds ~0.6 MB of reads per decoded token, about 0.01% of what the base reads.
  • Pairs are fingerprinted. ③ stores a hash of the ② it was solved against (base.fnv); the engine refuses a mismatched pair. Without that check a stale ③ would run fine and quietly produce wrong numbers.
./bin/ds4-server --cuda -m base.gguf --zchain sidecar_dir/ [--posttrain posttrain_dir/] --mem-budget-mb 110000

4. The algorithms

4.1 The base: 8-dimensional vector quantization with shared codebooks

In one sentence. Every 8 consecutive weights of a routed-expert row become one 12-bit (or 13-bit) index into a codebook, times one gain per row.

Intuition. The official experts are FP4 with one scale per 32 weights (4.25 bits/weight). A scalar format below that still has to pay for those scales — 0.25 bits/weight on its own. Vector quantization codes 8 weights jointly, which beats scalar rate–distortion at the same budget, and it needs only one gain per row. That is enough here because inside a row the official scales vary only 2× on average and 4× at worst (measured on layers 0, 20 and 39).

The format.

w[r, 8j : 8j+8]  ≈  g_r · C[ idx[r, j] ]        C = 4096 × 8 (12-bit)  or  8192 × 8 (13-bit)
bits per weight  =  12/8 = 1.5   or   13/8 = 1.625,   plus one f16 gain per row (≈ 0.003)

Training the codebook.

  1. Normalize every row by its RMS, g_r.
  2. Run Lloyd iterations on the GPU. The assignment step is fused: the codebook streams through shared memory in 512-entry chunks and each thread keeps its running argmin in registers, so the distance matrix never exists in memory (a naive G = V·Cᵀ with cuBLAS moved 144 GB per expert). The vector width is a template parameter so the 8 values stay in registers — 3× faster than a runtime width.
  3. Deterministic by construction: equidistant initial sampling (no random source), a fixed number of rounds, double-precision atomics in the update. Same input, same bytes.
  4. The final assignment uses the codebook and gains already rounded to their stored precision, so the error the quantizer reports is exactly the error the engine will see.

One codebook per layer, shared by all 384 experts and all three matrices. Because the experts are statistically the same shape (relative error equal within 1.9%), a per-layer codebook trained on a pool of ≥ 1,000 vectors per codeword costs almost nothing: layer 20 error 0.17062 vs 0.17072 with one codebook per expert (slightly better), layer 39 +0.33%, +0.15% on average. It removes 15,744 codebooks — 3.09 GB. The pool size matters: a pool 4× thinner lost 0.96% on 13-bit layers.

FP8 (E4M3) codewords. Storing codewords as E4M3 costs +0.12% error and halves lookup bytes. The decoder must use the hardware conversion (cvt.rn.f16x2.e4m3x2, sm_89+). Our first version called a scalar converter — bit-exact, but decoding got 26% slower and prefill 20% slower.

Where the saved 3.09 GB goes: the 14 shallowest layers move from 12 to 13 bits (error −16.1%). The error curve is convex (removing a bit costs 5.3 pp, adding one gains 2.6–3.7 pp), so extra bytes must be spread thin rather than stacked: 14 layers at 13 bits remove more total error than 7 layers at 14 bits, and shallow layers were measured to be worth 1.2–1.4× more per GB than deep ones.

The "12 + 1" bit-plane layout. A straight 13-bit stream makes each 256-index block 104 words = 3.25 cache lines, so every 128-byte load straddles two lines. On GB10, loads that don't fill whole 128-byte lines fall from ~219 GB/s to ~142 GB/s (measured with our mem_ceiling probe). So 13-bit layers keep a 12-bit main stream with exactly the 12-bit geometry (one block = 3 full lines) plus a separate 1-bit plane: one extra 32-bit load per block and one warp shuffle per round.

Failing loudly. Version-3 payloads carry a new magic (DQV3) and the loader whitelists format versions. An engine that does not know a format refuses the file at load; one that predates the whitelist outputs all-zero experts from the first sentence. Neither can produce plausible garbage.

Zero corpus. The base is quantized from the weights alone. We tried finance-weighted calibration: finance flat, general text −2.4 pp. Domain knowledge belongs in ②, not in ①.

What the 113.56 GB is made of.

Part Format GB
Routed-expert indices VQ-8, 26 layers × 12 bit + 14 layers × 13 bit 104.90
Row gains f16 0.30
Codebooks 40 × E4M3 0.003
Attention, shared experts, output head q4_K 4.96
Speculative draft towers VQ-8, 12 bit 2.56
Everything else as shipped 0.85

Evidence. Against our previous base of the same size (one codebook per expert, all layers 12-bit): Same top-1 +1.9 pp (finance), +1.8–2.0 pp (eight domains), +2.5 pp (English); the worst 5% of positions improved most (Σmin p5 +22% / +27% / +83%). The base alone now beats the previous base with its sidecar on English text.

4.2 The domain sidecar (反修)

We call the solve 反修 ("repairing backwards") and its product the amplifier (放大器); the sidecar is a directory of amplifier files.

In one sentence. For every expert, re-solve one multiplicative gain per output channel of its down projection, plus one router bias per expert, so that each quantized MoE block reproduces the original block's output on domain text.

Intuition, starting from the dead end. Our first sidecars added a low-rank correction at the block output, y += B·A·x. It removed 9% of held-out layer error on average — and made end-to-end metrics worse. A step-size scan settled why: any correction that is a linear function of the block input has ≈ 0 first-order effect at the model's output. A per-expert gain escapes this. The correction it injects, Δ = Σ_e rw_e · (s_e − 1) ⊙ y_e, depends on which experts the router picked and on each expert's own hidden state, so it is not a linear function of x. The first three layers solved this way gave the first positive end-to-end result of the whole project.

The objective, per layer.

y_i   = Σ_k rw_ik · ( s_{e_ik} ⊙ ye_ik ) + ysh_i          # the engine's own MoE sum
        ye = expert down-projection output before the routing weight, ysh = shared expert

min_s   Σ_i ‖ y_i − y_i^orig ‖²   +   λ · Σ_d  d̄_d · Σ_e ( s_e[d] − 1 )²
  • y^orig is the original model's output of the same block at the same token, from a teacher fixture: DeepSeek's code at full precision, 40 layers × 8,192 tokens, 14 GB, written once per fitting corpus.
  • The gains are diagonal, so the 5,120 output channels are independent; each is a least-squares problem with 384 unknowns.
  • It is solved by per-expert Gauss–Seidel (3 sweeps). Each expert touches only the ~128 rows routed to it and updates the residual immediately — no 384 × 384 systems are ever formed.
  • The ridge is scaled by d̄_d, the mean energy of all experts on channel d. Using each expert's own energy instead let rarely-routed experts run to extremes: layer 0 went train +32.7% / held-out −2685%. With the shared scale, the same layer is +27.0% held-out.

Fitting on the deployment path, one layer after another.

  • ye and ysh come from hooks inside the engine's own prefill kernels; nothing is recomputed in a second implementation. Every layer is checked: round_bf16(Σ rw·ye + ysh) equals the engine's block output bit for bit on 99.996% of values, or the run stops.
  • Layers are solved 0 → 39 in a single pass. Layer k is fitted with layers < k already mounted, so it sees exactly the upstream it will see in deployment.
  • Per layer: solve the router bias Δb, mount it, recapture, then solve the gains.
  • λ ∈ {0.01, 0.1, 1, 10, 100, 1000} is chosen on held-out rows only. Held-out is stratified by source window (every 4th 128-token window, 2,048 of 8,192 tokens), so every sub-source appears on both sides.
  • A layer is mounted only if its held-out gain exceeds 0.5%. This is a significance gate, not a size limit: picking the best of six λ values produces a small positive number even on pure noise, and in a sequential chain one noise layer corrupts every layer after it. We verified it end to end: mounting every layer with a positive gain was worse on all four metrics.

FP4 lattice-direct solve. Gains are stored in FP4 (E2M1, one scale per 32). Solving in float and rounding afterwards would throw away 11% of the correction, because E2M1 has only 8 magnitudes. So each Gauss–Seidel step becomes three: solve one expert → project s − 1 onto the FP4 lattice and emit the 17-byte blocks directly → update the residual with the lattice value. Expert e's rounding error is absorbed by the experts solved after it — GPTQ-style error feedback, applied to one-dimensional gains. Layer 0 loses 0.27 pp of held-out gain instead of 11%; the sidecar shrank 291 MB → 38.6 MB with no metric regressing. The file written is the lattice codes, verified by decoding with the engine's own decoder before it is kept.

The router bias Δb. For each layer and expert, accumulate the selection-score gap between the original's chosen set and ours: if the original picked e and we did not, add thr − v_e; if we picked e and the original did not, subtract v_e − thr (thr = our K-th selection score). Average, arm an expert only after ≥ 8 events, scale by α ∈ {0.5, 1, 1.5, 2.5} chosen on held-out same-set rate, and write the layer only if held-out strictly improves. Δb is added to the selection score only; mixing weights are untouched. The original routing is recomputed from the fixture's full-precision inputs with the model's own bf16 router; the recomputation must match the engine's choice on ≥ 99% of rows.

Cost. Current finance sidecar: gains on 39 layers + router bias on 27 layers, ≈ 40 MB. Fitting on the Spark: 27 min for the teacher fixture (once per corpus) + 73 min for the 40-layer solve.

Evidence (paired, same binary, same run; full tables in §5.1). Gains: finance 71.7 → 73.9%. Router bias on top, gains held fixed: finance +0.24 pp and KLD −2.5%, English +0.78 pp and KLD −3.8%.

4.3 The post-training file

In one sentence. Turn "at this position the model should have written a, not b" into a linear equation on last-layer expert gains, solve all such equations together while pinning everything else, and store the result as a third gain table.

Why not ordinary fine-tuning gradients. Our first version used the SFT loss gradient as the target and moved nothing (held-out +0.04%). The loss gradient is a dense direction set by the target token's embedding, while per-expert gains can only reweight expert outputs channel by channel: the two are nearly orthogonal. What decides whether a decision flips is one scalar — the margin between two logits — and that is exactly linear in the last MoE layer's output.

The formulation.

  • Decision point i: a position where the right token a should beat the strongest other token b. (Not "the token the wrong version wrote": 81% of those were already beaten, and the argmax went to a third token.)
  • Margin m_i = ℓ_a − ℓ_b. Holding the output RMSNorm factor inv_i fixed, a change Δ of the last layer's output moves the margin by α_i · inv_i · Σ_d γ_d (W_a − W_b)_d · Δ_d (γ = norm weight, W = output head), and Δ is linear in the gain table.
  • Solve r_i · s = τ − m⁰_i for all decision points (r_i = the coefficient row from the expression above, τ = target margin), plus constraint rows (every other position keeps its top-1/top-2 margin, weight ρ), plus a ridge λ, with matrix-free conjugate gradient and a Jacobi preconditioner — 1.97M unknowns, seconds.
  • FP4-aware: quantize the solution to the lattice, re-predict with the stored values, put decision points that fell back below τ into an active set, solve again.
  • Capture on the deployment path. Margins must be measured where the decision is actually made: prompt through prefill, generated text through the decode path (--score-split P). Captured on the prefill path alone, a solve flipped the decision when scoring but not in real generation.
  • Safety: ③ is its own directory, multiplied with ② at load, fingerprinted against ②; candidates that fail a gate are moved to rejected/, never deleted.

Evidence and status — this is the least finished part; see §8. On the training day, decision points flipped 55% → 88% (64 / 73) while 99.66% of 3,268 other positions kept the same top-1; the solver's predicted margins correlate 0.9995 with a real forward. On a real market-outlook request, the trained decision token flips in actual generation ("I lean towards predicting 'market down' … let me reconsider") — and ~6,000 characters later the model argues its way back to "up".

4.4 The engine

In one sentence. Every kernel on the hot path is written for this format on GB10, and every speed-up must leave temperature-0 output byte-identical.

The wall. Decoding is memory-bound: each token reads ~6.3 GB. At GB10's measured ~235 GB/s that is 26.8 ms (37 tokens/s). We run at 32.5 ms — 82% of the wall.

Per decoded token Bytes Kernel reach
Output head (q4_K) 372 MB ~249 GB/s
Attention and shared-expert projections (q4_K) ~3.6 GB 193–210 GB/s
Routed experts (VQ, 6 of 384 per layer) 1.64 GB ~155 GB/s
n-gram memory projection (FP8) 314 MB ~226 GB/s
Router and small mixing matrices 314 MB 132–136 GB/s

Decode.

  • VQ expert kernel. Persistent. Each layer's codebook is loaded into shared memory once per SM. Index streams are read as 384-byte blocks (3 full 128-byte lines) and codes are handed out with warp shuffles; consecutive rows stream back to back. At load time every expert payload is shifted so its index stream starts on a 128-byte boundary (device copy only, file untouched): 145–155 → 124–129 µs per layer.
  • q4_K GEMV (the other ~60% of bytes). Two shapes: stage (a CTA moves whole lines into shared memory before computing) and pipe (persistent, cp.async double buffering, for 4–12 KB groups). 22.9 → 18.8 ms per token.
  • Whole-step CUDA graph. One capture per position bucket; the token position lives in a device slot, so replay never re-captures. Scratch buffers are grown before capture — a growth attempt during capture once made the server run a whole night without graphs (43.9 ms/token).
  • Programmatic dependent launch. 1,391 kernel edges: each kernel prefetches its constant weights, then waits for its producer.
  • Side stream. The attention KV branch and the shared expert run in parallel with the main chain.
  • Long context. The sparse-attention selection kernels were rewritten (1,024-thread top-k, batched reads, 4 groups per warp): at 51k context the extra cost fell 4.6 → 2.8 ms per token.

Prefill. Experts run on bf16 tensor cores (mma.m16n8k16). E4M3 codewords convert to bf16 exactly and activations already sit on the bf16 grid, so every product equals the scalar path's. Long accumulation inside the tensor core drops low bits (3.5% of outputs off by one ulp), so each k16 slice is accumulated from zero and added with a separate FADD (0.5%). 12.5k-token prompt: 216.5 → 489 tokens/s. The quality change (Σmin 0.7447 → 0.7430) is the same size as merely reordering the float reduction (0.7439): rounding noise.

Speculative decoding.

  • The model's three draft towers are quantized with the same VQ (2.56 GB).
  • The verify batch (1 + k rows) has its own kernels: GEMVs on stage/pipe with several rows, and a persistent expert kernel over (unique expert × rows) that reads each chosen expert's bits once.
  • Scheduler. The draft head reports a confidence c_j per position. With survival a_j = Π_{i≤j} σ(c_i), pick

k* = argmax_k (1 + Σ_{j≤k} a_j) / (c_draft + c_v1 + c_tok · (1 + k))

with costs in units of one plain decode step (0.273, 1.067, 0.303 — re-measured whenever a kernel changes). If no k beats plain decoding, the round skips drafting. The scheduler takes no wall-clock input: at temperature 0, output must not depend on how busy the machine is. - Guarantee: speculative output is byte-identical to plain greedy output, checked on every change. On by default at temperature 0; a request that asks for sampling runs plain decode automatically. - Why it isn't higher: each extra verified token touches ~3.9 new experts (hash-like routing, §2.2), so verifying k + 1 tokens costs far more than verifying one. The verify batch runs at 57% of its own byte wall; that is the next lever.

Memory. Weights are mmap-backed. Per-request state grows with the positions actually used: the 1M KV itself is 0.89 GB, and the per-forward scratch that used to be sized for 1M up front (2.5 GB) now grows by doubling (a 40k-token request uses 160 MB). There is no context knob — the bound comes from the GGUF metadata and --ctx is rejected.

4.5 How we measure

In one sentence. The original model, run with DeepSeek's own code at full precision, is the only judge, and our engine is scored on the same path it serves.

  • Teacher. DeepSeek's official PyTorch inference code with the original FP4/FP8 weights streamed layer by layer from SSD (510 GB does not fit in memory). Teacher logits are cached per (text, length).
  • Student. The engine's scoring path (--score-ids), same file format, one comparator (anchor_metrics) for everything.
  • Five numbers. Same top-1; Σmin (mean, median, p5); mean KL(original ‖ ours); PPL ratio. Σmin and KL are primary. Same top-1 alone is lenient — easy text hides damage (one early recipe read 0.90 on easy code and 0.52 on hard text).
  • Disjoint slices. a (8,192 tokens) fits sidecars, j (8,192 tokens) judges them; they never mix. Rulers: finance j (five Chinese-finance sources), eight-domain j (academic, prose, code, European, Cyrillic, math, Arabic, CJK), WikiText-2 512 tokens (English).
  • Paired, one variable at a time. Comparisons use the same binary in the same run. An earlier verdict that "routing bias taxes general text" was retracted when a single-variable pair showed the opposite.
  • Gates that exist because we got burned.
  • Temperature-0 byte identity: direct launch == CUDA graph == replay, and speculative == plain.
  • A format change must also pass decode-path checks. A 13-bit decode kernel once advanced its bit-plane pointer by block index instead of group index: every generated token used wrong weights in 14 layers, while all five metrics — computed on the prefill path — stayed green. It looked exactly like "the model repeats itself". After the fix, decode-path PPL went 8.19 → 5.80 and prefill/decode disagreements 22 → 1.
  • Never judge "it doesn't stop" under an output cap: a 16k cap once turned a normal 23,607-token answer into an apparent loop.

5. Results

5.1 Quality

Finance ruler (judge slice, 8,192 tokens; original model PPL 6.978):

Same top-1 Σmin (median / p5) Mean KL PPL ratio
① base alone ¹ 71.73% 0.703 (0.751 / 0.231) 0.612 1.339
① + ② gains 73.94% 0.742 (0.810 / 0.276) 0.522 1.282
① + ② gains + router bias (deployed) 74.57% 0.745 (0.810 / 0.283) 0.512 1.267

English, WikiText-2 (512 tokens; original PPL 1.657):

Same top-1 Σmin (median / p5) Mean KL PPL ratio
① base alone ¹ 79.49% 0.769 (0.952 / 0.067) 0.692 1.797
① + ② gains 79.88% 0.789 (0.967 / 0.096) 0.612 1.645
① + ② deployed 81.05% 0.790 (0.966 / 0.091) 0.604 1.649

Eight-domain mix, ① base alone ¹: Same top-1 67.91%, Σmin 0.690 (p5 0.259), KL 0.602, PPL ratio 1.359.

¹ Reference forward on the same file; engine vs reference on the same file differ by KL 0.013. The 09-24 tensor-core prefill moves the deployed finance row to 74.48% / 0.743 / 0.516 (rounding noise, §4.4).

5.2 Speed (one DGX Spark)

Workload Prefill Decode
12.5k-token prompt 489 t/s (253 with --decoder-full) —
Short prompt, plain greedy — 30.5–30.7 t/s
Real agent request, 14.1k-token prompt, plain — 28.9–29.4 t/s
Same request, speculative (default) — 43.0 t/s (3.04 tokens per round)
Request never used for tuning, 9.2k prompt, speculative — 37.1 t/s
51k context, plain (09-23 build) — 27.5 t/s

5.3 How we got here

Date (2026) Step Result
09-11 First VQ file, experts at 1.5 bits/weight Σmin 0.746 on English (experts only)
09-12 Engine runs V4.1 end to end first token; decode 6.8 t/s
09-13 Sidecar switches to per-expert gains finance 68.95 → 72.08%
09-14 FP4 lattice-direct sidecar 291 MB → 38.6 MB, 72.18%
09-15 Prefill rewrite 58.9 → 348.5 t/s (later traded to ~209 when a 4-bit activation path was retired for quality)
09-18 Whole-step CUDA graph 12k-context decode 21.9 → 25.4 t/s
09-20 q4_K projections, 12-bit experts, gains finance 73.19%
09-21 Shared E4M3 codebooks, 13-bit shallow layers, router bias finance 74.57%, English 81.05%, file 0.12 GB smaller
09-22 13-bit decode-kernel bug fixed the "repetition" disappears
09-23 GEMV rewrite, PDL, side stream server decode 22.8 → 30.0 t/s
09-24 Tensor-core prefill, verify-batch kernels, speculative on by default prefill 489 t/s, decode 43.0 t/s

6. What did not work

Negative results carry as much of the design as positive ones. Each row was measured, not argued.

Idea What we measured Verdict
Dynamic bits per layer / expert / matrix Importance spread 3% / 1.25× / 33:33:33; moving a bit needs ≥ 1.189× Uniform is optimal
Dynamic bits per row 2.5× energy spread, eaten by codebook cost and integer bit widths: net −1.6% to −5.9% Rejected
Expert pruning to reach 100 GB Would keep 130 of 384 experts per layer Rejected: a domain decision in disguise
Low-rank additive correction y += B·A·x Held-out +9% per layer, end to end negative; first-order effect at the output ≈ 0 Whole family rejected
Mean / bias correction of quantization error Inside the layer +4 pp; at the output −0.33 pp (3 layers); 40 layers 70.52 → 63.89% Rejected
Re-solving codebooks on domain data 65.43% vs 65.59% No gain
Entropy coding of VQ indices Empirical entropy 11.927 of 12 bits: saves 0.6% Not worth it
Entropy-constrained VQ + variable-length streams ≈ +0.1 bit net after stream overhead Not worth it
Finance-calibrated base Finance flat, general text −2.4 pp Base stays zero-corpus
Draft vocabulary cut to finance terms +0.2% to −54% on an unseen request Rejected
Grouped multi-token expert kernel for verification Halves codebook lookups, 0 ms saved Rejected (fourth time)
Second machine for decoding Layer split −19% (upstream); tensor parallel 1.25× One box

7. Run it

What you need.

  • NVIDIA DGX Spark (GB10, sm_121, 128 GB) — the only machine tested. The FP8 codebook path needs sm_89 or newer. V4.1 runs on CUDA only.
  • Linux aarch64 with the CUDA 13 runtime (libcudart.so.13, libcublas.so.13, libcublasLt.so.13; DGX OS ships them). Missing libraries show up as error while loading shared libraries: libcudart.so.13.
  • ~320 GB of local SSD: 113.6 GB for this repository + 203 GB for two official shards (below).
  • Nothing else heavy running: the engine refuses to start if it cannot fit the 110 GB budget.

One command. On the Spark:

curl -fsSLO https://huggingface.co/wenzhouwu/YoungAi-DeepSeek-V4.1-Flash/resolve/main/install.sh
bash install.sh

It checks the machine (GPU, CUDA 13 libraries, memory, disk), downloads this repository and the two official shards (~317 GB), assembles the base model, creates the one link the GGUF needs (asks for sudo once), starts the server on 127.0.0.1:8000 under a memory watchdog, and prints the answer to a one-line test question. Interrupted? Run the same command again: finished files are skipped and the assembly resumes where it stopped.

Option Effect
--dir DIR install directory (default ~/youngai)
--host 0.0.0.0 / --port N serve your LAN / another port
--endpoint https://hf-mirror.com download through a mirror
--no-xet download over the plain LFS channel (if transfers keep failing with "peer closed connection")
--posttrain also load the experimental post-training file
--no-start install only; later bash install.sh start, stop, status

The steps below are what the script does, for doing it by hand.

Files in this repository.

Path What
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative.gguf.part01-of-40 … part40-of-40 ① base, 113,556,639,424 bytes in 40 parts (a single 113.6 GB upload did not survive our uplink)
SHA256SUMS sha256 of the assembled base and of every part
install.sh the one-command installer above
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/ ② finance sidecar: gr_Lnn.bin (gains, 39 layers) + rb_Lnn.bin (router bias, 27 layers) + manifest.txt (per-layer λ and held-out gain)
posttrain-experimental-20260924/ ③ an experimental post-training file: gr_L39.bin + base.fnv (see below)
bin/ds4, bin/ds4-server engine binaries, built on the Spark with make cuda-spark
LICENSE, LICENSE-DeepSeek MIT notices for the engine (incl. GGML) and for the model weights

Step 1 — download and assemble. By hand this needs 227 GB free during assembly (the installer needs one part's worth, because it appends and deletes one part at a time).

hf download wenzhouwu/YoungAi-DeepSeek-V4.1-Flash --local-dir ds4-v41
cd ds4-v41
sha256sum -c --ignore-missing SHA256SUMS            # every part must report OK
cat DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative.gguf.part{01..40}-of-40 > DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative.gguf
sha256sum -c --ignore-missing SHA256SUMS            # now the assembled .gguf reports OK too
rm DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative.gguf.part*-of-40
chmod +x bin/ds4 bin/ds4-server                      # downloads do not keep the executable bit

Step 2 — the n-gram memory tables. They are not in this repository: the engine reads them, untouched, from two shards of the official checkpoint (≈ 101.5 GB each).

hf download deepseek-ai/DeepSeek-V4.1-Flash \
    model-00047-of-00048.safetensors model-00048-of-00048.safetensors --local-dir /data/DeepSeek-V4.1-Flash

The GGUF records these two shards by absolute path — /home/fodelf/ds4-main/hf/DeepSeek-V4.1-Flash/model-0004{7,8}-of-00048.safetensors — and the engine has no option to change it. Make that path point at your copy:

sudo mkdir -p /home/fodelf/ds4-main/hf
sudo ln -s /data/DeepSeek-V4.1-Flash /home/fodelf/ds4-main/hf/DeepSeek-V4.1-Flash

Skip this and the engine stops at load with ds4: engram 表打不开 /home/fodelf/… ("cannot open engram table").

Step 3 — serve.

cd ds4-v41
./bin/ds4-server --cuda -m DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative.gguf \
    --zchain DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine \
    --mem-budget-mb 110000 --host 0.0.0.0 --port 8000

Loading takes about two minutes. Endpoints: /v1/chat/completions, /v1/completions, /v1/responses (OpenAI style) and /v1/messages (Anthropic style). A request without temperature is decoded greedily (and speculatively); a request with temperature is sampled and runs plain decode.

Command line.

./bin/ds4 --cuda -m DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative.gguf \
    --zchain DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine \
    -p "Explain the price-to-earnings ratio."
# add --no-dspark for plain decoding (speed baselines); drop --zchain to run the bare base

The post-training file is an experiment, not an upgrade. posttrain-experimental-20260924/ was solved on a single market-outlook request from our trading agent (2026-09-22) and moves one decision token, only on the last layer. It is here to show the format and the loading path:

./bin/ds4-server … --zchain <sidecar dir above> --posttrain posttrain-experimental-20260924

It loads only on top of the sidecar in this repository — base.fnv is that sidecar's fingerprint, and the engine refuses any other pairing. Do not expect it to improve anything else.

Source code of the engine, the quantizer and the solvers is not public yet.


8. Honest status and limits

  • Post-training (③) is a working mechanism, not yet a working product. It flips targeted decisions on the day it was trained without disturbing other tokens, but it does not transfer to new days: with two trading days of reviews, held-out decision points stayed at 55% → 55%. And one flipped token changes one sentence, not a chain of reasoning. The next design — sample N answers per real request, score each with the next day's actual market data, solve ③ on group-relative advantage — is written down, not yet run.
  • One domain so far. Only a Chinese-finance sidecar exists.
  • CUDA only, one machine type tested. V4.1 does not run on Metal.
  • Speculative decoding is greedy-only. Sampling requests fall back to plain decode.
  • Two of the speed numbers include kernel changes not yet merged (128-byte payload alignment and the side stream for verify batches). Merged code measures ~30.0 t/s plain (short context) and 39.3 t/s speculative on the 14k request.
  • Prefill on tensor cores is not bit-identical to the older fused path; the difference is at the rounding-noise level (§4.4).
  • The n-gram shards are pinned by absolute path (§7, step 2).

Author and contact

Wenzhou Wu (吴文周) · [email protected]

Questions, reproduction reports and collaboration offers are welcome.

Acknowledgements

This project started as a fork of antirez/ds4 (DwarfStar), the DeepSeek-V4-specific engine by Salvatore Sanfilippo and contributors, where the documentation for the V4 / Metal paths lives. Like upstream, we are indebted to llama.cpp and GGML: GGUF, quantization layouts such as q4_K, and much hard-won kernel knowledge come from there, and the GGML authors' copyright notice stays in LICENSE. The model is DeepSeek's; thanks to DeepSeek for releasing the weights and the reference inference code that serves as our ruler.

License

Engine: MIT — see LICENSE. Quantized weights derive from DeepSeek V4.1 Flash, MIT — see LICENSE-DeepSeek.

Identity and Version

Repository
wenzhouwu/YoungAi-DeepSeek-V4.1-Flash
Publisher
Wenzhou Wu
Task
Text generation
Modality
Text
Library
Not stated by the source
Parameters
Not stated by the source
Languages
en, zh
Revision
9852affc876bfc2fb4b0e24b69136e2594cbf013
First published
2026-09-24
Last updated
2026-09-25

Files and Weights

82 files, 11.4 GB in total. The weights are 67 files totalling 41.8 MB in bin.

Weights67 files · 41.8 MB
Documentation4 files · 81.6 KB
Other10 files · 11.4 GB
Repository1 file · 2.0 KB
Every file
FileTypeSizeSHA-256
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L00.binWeights1.0 MB 53c306929d97
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L01.binWeights1.0 MB 364d1c65dc28
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L02.binWeights1.0 MB 41817d738938
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L03.binWeights1.0 MB 1fded06c0a4e
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L04.binWeights1.0 MB e2ab7d097bf6
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L05.binWeights1.0 MB d2c399fb03e1
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L06.binWeights1.0 MB 44e95e151432
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L07.binWeights1.0 MB 91f6eae17f4e
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L08.binWeights1.0 MB c5b650496ef5
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L09.binWeights1.0 MB 3f7b47597d8b
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L10.binWeights1.0 MB 87f9da3e743e
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L11.binWeights1.0 MB fa406c9cb9d3
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L12.binWeights1.0 MB 9173d70eea2a
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L13.binWeights1.0 MB ad2ffcb8be6a
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L14.binWeights1.0 MB be89ecae296e
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L15.binWeights1.0 MB eeb9af8425ec
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L16.binWeights1.0 MB 7e8f770075f9
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L17.binWeights1.0 MB f223961fe570
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L18.binWeights1.0 MB 0a698de8cba2
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L19.binWeights1.0 MB 0f2c9375d473
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L20.binWeights1.0 MB 0a3c83c062ef
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L21.binWeights1.0 MB 910076e62876
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L22.binWeights1.0 MB a0a103921c92
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L23.binWeights1.0 MB c09cfdc252d8
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L24.binWeights1.0 MB 7f9a86416b9a
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L25.binWeights1.0 MB a6cac9f07f3c
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L26.binWeights1.0 MB 15b88f677805
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L27.binWeights1.0 MB da651cfafc61
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L28.binWeights1.0 MB d85d066799df
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L29.binWeights1.0 MB 7f6717f13664
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L31.binWeights1.0 MB baed00fcd3ce
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L32.binWeights1.0 MB 5424882eb52a
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L33.binWeights1.0 MB b4c1d4ff72a1
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L34.binWeights1.0 MB b63b71763efa
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L35.binWeights1.0 MB 873887619e23
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L36.binWeights1.0 MB ccd24114d315
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L37.binWeights1.0 MB 875a2e4faf83
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L38.binWeights1.0 MB 96baa13e02c9
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/gr_L39.binWeights1.0 MB 8a9e63edcfd4
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/rb_L00.binWeights1.5 KB 7a37f475ba56
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/rb_L01.binWeights1.5 KB a7ee5879addb
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/rb_L02.binWeights1.5 KB fdd54eecc08b
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/rb_L03.binWeights1.5 KB 992190153a98
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/rb_L04.binWeights1.5 KB 17c59cefe488
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/rb_L05.binWeights1.5 KB 499dd767841d
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/rb_L06.binWeights1.5 KB b90bce82b4e6
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/rb_L07.binWeights1.5 KB 6c1cd36d76d3
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/rb_L08.binWeights1.5 KB bc5276bfb780
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/rb_L09.binWeights1.5 KB fd6973d9f245
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/rb_L10.binWeights1.5 KB b779c2f19343
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/rb_L11.binWeights1.5 KB e118c2068ba4
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/rb_L12.binWeights1.5 KB c1f7c9767541
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/rb_L13.binWeights1.5 KB f49d2b3cfeec
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/rb_L14.binWeights1.5 KB 7c68dbbefb6c
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/rb_L16.binWeights1.5 KB b5b48fc1c562
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/rb_L17.binWeights1.5 KB c260ef0b464a
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/rb_L18.binWeights1.5 KB 50a470a5cc75
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/rb_L19.binWeights1.5 KB 89c5e513a498
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/rb_L20.binWeights1.5 KB 030872184221
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/rb_L21.binWeights1.5 KB bf3b489e226c
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/rb_L23.binWeights1.5 KB 4cd51e66e805
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/rb_L25.binWeights1.5 KB 5098e2e45212
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/rb_L27.binWeights1.5 KB 4fb6207c76fd
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/rb_L28.binWeights1.5 KB 9624014d0eef
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/rb_L31.binWeights1.5 KB b2bd00bb622c
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/rb_L37.binWeights1.5 KB 261d2a99a2fb
posttrain-experimental-20260924/gr_L39.binWeights1.0 MB 0a636ad4f2e4
LICENSEDocumentation1.1 KB —
LICENSE-DeepSeekDocumentation1.1 KB —
README.mdDocumentation40.7 KB —
README.zh-CN.mdDocumentation38.8 KB —
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative-grrb-vqfin41_vqhalf_a_n8192-engine/manifest.txtOther5.8 KB —
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative.gguf.part01-of-40Other2.8 GB 643fa5f3c274
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative.gguf.part02-of-40Other2.8 GB f829ce3cfc15
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative.gguf.part03-of-40Other2.8 GB 976f4bc6f3a4
DeepSeek-V4.1-Flash-vq8sh14-q4k-mtpnative.gguf.part04-of-40Other2.8 GB e5f13f1e1476
SHA256SUMSOther5.2 KB —
bin/ds4Other13.8 MB 06c6d9c68918
bin/ds4-serverOther14.8 MB c910ec230913
install.shOther13.3 KB —
posttrain-experimental-20260924/base.fnvOther20 B —
.gitattributesRepository2.0 KB —

License and Download

License
mit
Access
Open weights, no gate
Download size
41.8 MB
Download from Wenzhou Wu

Released by Wenzhou Wu through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published41.8 MB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About YoungAi-DeepSeek-V4.1-Flash

Can I use YoungAi-DeepSeek-V4.1-Flash commercially?

Yes. YoungAi-DeepSeek-V4.1-Flash is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

Similar Models

Fine-tune Qwen3 (14B) for free using our Google Colab notebook! - Read our Blog about Qwen3 support: unsloth.ai/blog/qwen3 - View the rest of our notebooks in our docs here. Qwen3-Coder is available in multiple sizes. Today, we're excited to introduce Qwen3-Coder-30B-A3B-Instruct. This streamlined model maintains impressive performance and efficiency, featuring the following key enhancements: - Significant Performance among open models on Agentic Coding, Agentic Browser-Use, and other foundational coding tasks. - Long-context Capabilities with native support for 256K tokens, extendable up to 1M tokens using Yarn, optimized for repository-scale understanding. - Agentic Coding supporting for…

Open weights apache-2.0 transformers

Model · Text generation

opt-125m

AI at Meta

OPT was first introduced in Open Pre-trained Transformer Language Models and first released in metaseq's repository on May 3rd 2022 by Meta AI. Disclaimer: The team releasing OPT wrote an official model card, which is available in Appendix D of the paper. Content from this model card has been written by the Hugging Face team. To quote the first two paragraphs of the official paper OPT was predominantly pretrained with English text, but a small amount of non-English data is still present within the training corpus via CommonCrawl. The model was pretrained using a causal language modeling (CLM) objective. OPT belongs to the same family of decoder-only models like GPT-3. As such, it was…

Open weights other 2,048 tokens transformers

Model · Text generation

Ornith-1.5-9B-GGUF

Ornith

Chirp Chirp! We are introducing Ornith-1.5, a major step toward building foundation models through end-to-end self-improvement. Ornith-1.5 extends Ornith-1.0 (which was developed on top of Qwen3.5 and Gemma4 with additional continued pretraining, mid-training, and post-training) by expanding the self-improvement loop from scaffold and rollout optimization to jointly optimizing task generation, scaffold construction, and solution rollouts. Rather than relying on a fixed set of human-curated tasks and manually designed harnesses, Ornith-1.5 continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning. For…

Open weights mit transformers

Model · Text generation

Ternary-Bonsai-2-27B-gguf

Prism ML

Full 27B-class reasoning in ternary transformer weights, for llama.cpp (CUDA, Metal, CPU) - \~5.9 GB language model (down from \~54 GB FP16) — full 27B-class reasoning on a standard laptop or a single GPU - 98.2% of FP16 intelligence retained: 84.78 average across 14 thinking-mode benchmarks — far above the conventional IQ2XXS build (72.59) at about 82% of its footprint, and within 0.4 points of UD-Q4KXL at three times the footprint - Retains thinking, reasoning, and agentic behavior deep in the sub-4-bit regime, where conventional low-bit representations collapse: math within half a point of full precision (96.57), coding level with the baseline (89.42), agentic tool calling at 74.92…

Open weights apache-2.0 llama.cpp

Model · Text generation

Ornith-1.5-35B-A3B-GGUF

Ornith

Chirp Chirp! We are introducing Ornith-1.5, a major step toward building foundation models through end-to-end self-improvement. Ornith-1.5 extends Ornith-1.0 (which was developed on top of Qwen3.5 and Gemma4 with additional continued pretraining, mid-training, and post-training) by expanding the self-improvement loop from scaffold and rollout optimization to jointly optimizing task generation, scaffold construction, and solution rollouts. Rather than relying on a fixed set of human-curated tasks and manually designed harnesses, Ornith-1.5 continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning. For…

Open weights mit transformers

Model · Text generation

Ornith-1.0-9B-GGUF

Ornith

Aloha! Today, we are releasing Ornith-1.0, a self-improving family of open-source models for agentic coding. This model card documents Ornith-1.0-9B, the most lightweight member of the Ornith family, designed for efficient single-GPU deployment. Ornith-1.0-9B is a dense ~9B model (≈19 GB in bf16), so it serves comfortably on a single 80GB GPU. The recipes below stand up an OpenAI-compatible server; add --tensor-parallel-size / --tp if you want to shard across more GPUs. For a quick local test (or to script offline generation), load the model directly with Transformers. Make sure you have a recent release installed — see the Transformers installation guide; Ornith-1.0-9B requires…

Open weights mit transformers