SAVRN
Search Contact SAVRN

Independent publisher

Zhu Lin

czl

Computer Vision, LLM

Models in Library3
Datasets in Library0
Models on Hugging Face8
Followers5

Models

Model · Feature extraction

CLM-v0.1-8B-MLX-8bit

Zhu Lin

This is the 8bit variant. Also available: czl/CLM-v0.1-8B-MLX-4bit, czl/CLM-v0.1-8B-MLX-6bit, and the bf16 reference. The encoder half of Contrastive-LM/CLM-v0.1-8B, quantised for MLX on Apple Silicon. mlxlm is generate-only — it has no embeddings entrypoint, and mlxlm.server routes just /v1/completions, /v1/chat/completions, /v1/models and /health. So clmmlx supplies the pooling pass over mlxlm internals: Qwen3Model.call already returns self.norm(h), the post-final-RMSNorm hidden states, which is what vLLM's pooling runner returns in last-token mode. Agreement against a bf16 MLX reference of the same encoder in the same runtime, over 23,926 scored System One questions, through the real…

Open weights apache-2.0 8.2B parameters 40,960 tokens mlx

Model · Feature extraction

CLM-v0.1-8B-MLX

Zhu Lin

The encoder half of Contrastive-LM/CLM-v0.1-8B, quantised for MLX on Apple Silicon. One repository per bit width, matching mlx-community. Each has the weights at the repo root, so the Hub file browser lists every file with its size and mlxlm.load(" ") works. Agreement against a bf16 MLX reference of the same encoder in the same runtime, over 23,926 scored System One questions, through the real head stack (argmax(scale · cos), scale = 100.0 — cosine error is amplified 100×). Pass line. top-1 >= 1.0000 — the measured bf16-vs-bf16 noise floor of this corpus in this runtime — and top-1 (decisive) >= 0.995, where decisive means the reference's own top-1 led by more than 1 nat. A third condition…

Open weights apache-2.0 8.2B parameters 40,960 tokens mlx

Model · Feature extraction

CLM-v0.1-8B-MLX-6bit

Zhu Lin

This is the 6bit variant. Also available: czl/CLM-v0.1-8B-MLX-4bit, czl/CLM-v0.1-8B-MLX-8bit, and the bf16 reference. The encoder half of Contrastive-LM/CLM-v0.1-8B, quantised for MLX on Apple Silicon. mlxlm is generate-only — it has no embeddings entrypoint, and mlxlm.server routes just /v1/completions, /v1/chat/completions, /v1/models and /health. So clmmlx supplies the pooling pass over mlxlm internals: Qwen3Model.call already returns self.norm(h), the post-final-RMSNorm hidden states, which is what vLLM's pooling runner returns in last-token mode. Agreement against a bf16 MLX reference of the same encoder in the same runtime, over 23,926 scored System One questions, through the real…

Open weights apache-2.0 8.2B parameters 40,960 tokens mlx