SAVRN
Search Contact SAVRN

Independent publisher

Alin C Selea

aselea

Models in Library1
Datasets in Library0
Models on Hugging Face1
Followers—

Models

Model · Text classification

Kev-4B-MLX-Serve-8bit

Alin C Selea

Kev-4B (a LoRA on Qwen3.5-4B-Base with a pointer head) packed for mlx-serve's POST /v1/decisions. Kev answers typed questions about a piece of text (choice, noul, score) with calibrated probabilities. It never generates text. The pack folds the LoRA into the base the way kev does on MLX, quantizes the trunk to 8-bit (affine, group 64; a bf16 build comes from --q-bits 0), and stores the pointer head as kevhead.safetensors with the calibration temperature in kevconfig.json. No PyTorch or pickle file is needed to serve it. Built with tests/convertkevweights.py from the mlx-serve repo. Kev and Qwen3.5 are Apache-2.0.

Open weights apache-2.0 4.2B parameters 262,144 tokens mlx-serve