Building the world's fastest CPU LLM inference layer
Search public pages, research tools, and SAVRN solutions.
SAVRN Model Hub
The labs, companies and researchers releasing open-weight AI models, with every model and dataset each one has published.
1,692 model publishers in the library.
Building the world's fastest CPU LLM inference layer
On-device & edge LLM inference on NVIDIA GB10 / DGX Spark. Quantization (NVFP4 W4A4/W4A16, FP8), vLLM serving, speculative decoding (EAGLE-3), and multimodal/omni models. Measure-first benchmarking — I publish the numbers, including the ones that fail.
multimodal
Organization
I love quantizing LLMs
Independent publisher
Hi!, im the creator of pixelmodel v3-v6 models, the text to 3d voxelmodel-v1 and the new AudioModel-v1
NLP: text embeddings, information retrieval, named entity recognition, few-shot text classification
Computer Vision - Medical imaging
Organization
Solving Cyber Security problems using ML