DeepSeek-V4.1-Flash, ready to serve on two RTX PRO 6000 Blackwell cards (sm120, 96 GB each) with vLLM at tensor-parallel 2. The routed experts are - int4 Engram tables; - a 3-bit DSpark drafter; - a serving stack that keeps 30% of the routed experts (the ones agentic coding uses least) in pinned host RAM and runs them next to the VRAM experts; - a decode-once prefill kernel for the 3-bit experts. The recipe/ folder has the image build, the sm120 patches, the MoE kernel, and the serving, benchmark, KL and quantization scripts. We measured KL against the original model with uses the original model's top-512 logprobs on three corpora, 49,152 scored positions each. The neutral corpus is…