SAVRN
Search Contact SAVRN

Organization

FINAL_Bench

FINAL-Bench

Contact: [email protected]

Models in Library1
Datasets in Library0
Models on Hugging Face64
Followers309

Models

Model · Text generation

POCKET-35B-GGUF

FINAL_Bench

Pick your build → -0f6e56) Wikipedia-Korean perplexity, lower is better. Q4KM = 5.79 baseline. English builds are tuned on English; see each repo. We measure Bonsai on the same machine with the same stock llama.cpp, and we tell you where we lose. [measured] Generation speed — POCKET wins on both CPU and GPU: [measured on a MacBook M3 Pro, 18 GB] — and on a laptop, POCKET wins every axis, including prompt processing: On a laptop GPU the arithmetic headroom that let Bonsai win prefill on an H100 is gone, so MoE sparsity wins across the board. POCKET-35B-Q2K runs on the M3 Pro's CPU at 19.5 tok/s — on an 18 GB Mac, run Q2K on CPU (-ngl 0); its 13 GB exceeds the recommended Metal budget.…

Open weights apache-2.0 llama.cpp