SAVRN
Search Contact SAVRN

Organization

Single Spark AI

single-spark-ai

LLM inference optimization, DGX Spark, SGLang, Qwen, NVFP4, FP8, speculative decoding, long-context inference

Models in Library1
Datasets in Library0
Models on Hugging Face1
Followers3

Models

A personal, measured-on-one-machine deployment recipe for serving Qwen/Qwen3.8-Flash-Next on a single NVIDIA DGX Spark (GB10) with SGLang. It is not a benchmark leaderboard claim and not "the fastest possible"; it is what is verifiably running on one GB10, with every patch, script, and check needed to reproduce it from upstream weights. Verified: 2026-09-13 · SGLang 00143e9c23aee2dead5e6fe217bda4fa8739cb92 (nightly-dev-cu13-20260911) · runtime image qwen38-flashnext-sm121-hybrid-sharp:00143e9c-hc2, built by bounded-ple/nightly/01c-build-sglang-nightly.sh. The v4 lineage (d91c3682-hc1) is still in the repo as the stable fallback. Key runtime facts (all asserted at deploy time by the 05…

Open weights other