# Diffbot: Open-Weight Models and Datasets
Source: https://savrn.com/model-publishers/diffbot
Markdown alternate of the page above; the site index is https://savrn.com/llms.txt

---

## Models

Model · Text generation

### [DeepSeek-V4.1-Flash-EXL3-3bpw-2x-RTX-PRO-6000](https://savrn.com/models/deepseek-v4-1-flash-exl3-3bpw-2x-rtx-pro-6000)

[Diffbot](https://savrn.com/model-publishers/diffbot)

DeepSeek-V4.1-Flash, ready to serve on two RTX PRO 6000 Blackwell cards (sm120, 96 GB each) with vLLM at tensor-parallel 2. The routed experts are - int4 Engram tables; - a 3-bit DSpark drafter; - a serving stack that keeps 30% of the routed experts (the ones agentic coding uses least) in pinned host RAM and runs them next to the VRAM experts; - a decode-once prefill kernel for the 3-bit experts. The recipe/ folder has the image build, the sm120 patches, the MoE kernel, and the serving, benchmark, KL and quantization scripts. We measured KL against the original model with uses the original model's top-512 logprobs on three corpora, 49,152 scored positions each. The neutral corpus is…

Open weights mit 219B parameters 1,048,576 tokens

[View model](https://savrn.com/models/deepseek-v4-1-flash-exl3-3bpw-2x-rtx-pro-6000)

## Explore More

- [All model publishers](https://savrn.com/model-publishers)
- [The model directory](https://savrn.com/models)
- [The dataset directory](https://savrn.com/datasets)

## Source

- Listed from their public repositories, read 2026-10-02.
- [Hugging Face profile](https://huggingface.co/diffbot)
- [How the hub is built](https://savrn.com/model-hub/methodology)
