SAVRN
Search Contact SAVRN

Independent publisher

Thor Lin

coolthor

On-device & edge LLM inference on NVIDIA GB10 / DGX Spark. Quantization (NVFP4 W4A4/W4A16, FP8), vLLM serving, speculative decoding (EAGLE-3), and multimodal/omni models. Measure-first benchmarking — I publish the numbers, including the ones that fail.

Models in Library1
Datasets in Library0
Models on Hugging Face15
Followers14

Models

Model · Any to any

gemma-4-12B-it-FP8-dynamic

Thor Lin

Self-quantized FP8 (dynamic) of google/gemma-4-12B-it — Google's encoder-free omni model (text + image + audio + video). Quantized and benchmarked on an NVIDIA DGX Spark (GB10, sm121a). TL;DR: 13 GB on disk (from 23 GB BF16), 15.9 tok/s on a GB10 via vLLM, all four modalities intact. Data-free — no calibration needed. If you want the smallest + fastest build, see the sibling NVFP4 weight-only repo. FP8 is the conservative choice (dynamic activations, no calibration, widest kernel support). I scored all three formats on MMLU (English, 57 subjects) and TMMLU+ (Traditional Chinese, 66 subjects) with lm-evaluation-harness, 5-shot, chat template applied, limit=30 (N ≈ 1,710 EN / 1,980 TC, ±~1.0…

Open weights apache-2.0 12B parameters 131,072 tokens transformers