SAVRN
Search Contact SAVRN

Independent publisher

Wenzhou Wu

wenzhouwu

Models in Library1
Datasets in Library0
Models on Hugging Face1
Followers—

Models

Model · Text generation

YoungAi-DeepSeek-V4.1-Flash

Wenzhou Wu

English · 中文 ds4 is a native C/CUDA inference engine plus an offline toolchain that runs DeepSeek V4.1 Flash on a single NVIDIA DGX Spark (GB10, 128 GB unified memory). Everything described here is our own work: the vector-quantization format, the solvers that produce the sidecar and the post-training file, the CUDA kernels, and the rulers we judge all of it with. The model architecture is DeepSeek's and is not re-explained here — read the official release for that. How to read this page. It unfolds in five steps; stop wherever you have what you need. 1. The numbers — what runs, how fast, how close to the original. 2. Three convictions — why the system is shaped this way. 3. The…

Open weights mit