SAVRN
Search Contact SAVRN

Organization

Diffbot

diffbot

Knowledge Graphs

Models in Library1
Datasets in Library0
Models on Hugging Face16
Followers50

Models

DeepSeek-V4.1-Flash, ready to serve on two RTX PRO 6000 Blackwell cards (sm120, 96 GB each) with vLLM at tensor-parallel 2. The routed experts are - int4 Engram tables; - a 3-bit DSpark drafter; - a serving stack that keeps 30% of the routed experts (the ones agentic coding uses least) in pinned host RAM and runs them next to the VRAM experts; - a decode-once prefill kernel for the 3-bit experts. The recipe/ folder has the image build, the sm120 patches, the MoE kernel, and the serving, benchmark, KL and quantization scripts. We measured KL against the original model with uses the original model's top-512 logprobs on three corpora, 49,152 scored positions each. The neutral corpus is…

Open weights mit 219B parameters 1,048,576 tokens