An experimental MXFP4 quantization of ibm-granite/granite-4.2-3b, made with GPTQ calibration instead of the usual data-free rounding, and provided in two formats with identical weights. It's published as a data point, not as the recommended build. For most users, the Q4KM GGUF or GPTQ int4 build is the better choice. Format: OCP MX — E2M1 (FP4) weights, groupsize=32, one UE8M0 shared scale per group; weights only, activations stay 16-bit. Embeddings, norms and lmhead are unquantized. MXFP4 is usually produced data-free, by rounding weights straight to FP4 with no calibration data. llm-compressor ships an MXFP4A16 scheme, but every example we found pairs GPTQModifier with integer schemes…
Open weights
apache-2.0
3.7B parameters
131,072 tokens
transformers
A 4-bit weight-only GPTQ quantization of ibm-granite/granite-4.2-3b, saved in compressed-tensors format for vLLM. On an AMD Radeon Pro V620 (RDNA2), vLLM served it at 65 tok/s single-sequence and ~1,000 tok/s aggregate at 16 concurrent sequences. llm-compressor GPTQModifier: - int4, symmetric, weight-only (activations stay 16-bit) - groupsize=32, actorder="weight" - lmhead left unquantized 512 samples, maxseqlength=512 The tuning mattered. With the default W4A16 preset (groupsize=128), the same pipeline scored 0.153 mean KLD. Moving to groupsize=32 with actorder="weight" cut that by about 28%, for roughly 0.2 GB more disk. For comparison, an AWQ build (W4A16 asymmetric, groupsize=128) from…
Open weights
apache-2.0
3.7B parameters
131,072 tokens
transformers
Model · Text generation
Web
A specialized 3.8B biomedical & clinical reasoning model built on Microsoft's Phi-4-mini-instruct, optimized natively for Apple Silicon Metal acceleration via Apple MLX. The model underwent a 3-stage transfer learning curriculum: 1. Stage 1 (STEM Foundation): 116,000 instruction pairs across NCERT Classes 6–12 (Physics, Chemistry, Biology) eliminating foundational science hallucinations. 2. Stage 2 (PubMed 2026 Evidence): 12 recent 2026 clinical update archives from NCBI FTP covering survival outcomes (OS, PFS, HR), targeted therapeutics, and clinical trial endpoints. 3. Stage 3 (Comprehensive Internal Medicine): Balanced multi-specialty clinical curriculum (cardiology, nephrology…
Open weights
mit
3.8B parameters
131,072 tokens
mlx
Model · Text generation
Web
A specialized 3.8B biomedical & clinical reasoning foundation model built on Microsoft's Phi-4-mini-instruct, formatted for standard Hugging Face transformers and PyTorch. 1. Stage 1 (STEM Foundation): 116,000 instruction pairs across NCERT Classes 6–12 (Physics, Chemistry, Biology) eliminating foundational science hallucinations. 2. Stage 2 (PubMed 2026 Evidence): 12 recent 2026 clinical update archives from NCBI FTP covering survival outcomes (OS, PFS, HR), targeted therapeutics, and clinical trial endpoints. 3. Stage 3 (Comprehensive Internal Medicine): Balanced multi-specialty clinical curriculum (cardiology, nephrology, endocrinology, pulmonology) with an active oncology replay buffer.…
Open weights
mit
3.8B parameters
131,072 tokens
transformers
Standard One scores a bounded set of answers for a supplied scenario and returns probabilities through POST /v1/systemone. It does not generate free-form response text. This repository contains the merged BF16 3B checkpoint; the server code is in StandardOne-8B. In the reported served evaluations, 3B has a lower median latency on the measured short-request profile; 8B scores higher on the public standard and hard tiers. See Benchmarks for the measurement conditions and limitations. The figure combines results from different measurement paths. See Benchmarks for served versus offline conditions; measured 24–26 September 2026. - Send a state and a bounded rubric to receive probabilities for…
Open weights
apache-2.0
3.8B parameters
262,144 tokens
transformers
StandardOne-3B-FP8 is an FP8 (compressed-tensors, float8e4m3 weights, dynamic per-token activations) quantization of the released StandardOne-3B decision model. The language-model linear projections (q/k/v/o, gate/up/down) are quantized per-channel FP8 E4M3 with dynamic FP8 activations (llm-compressor's data-free FP8DYNAMIC recipe, no calibration data required); the vision tower, multi-modal projector, embeddings and lmhead are left unquantized in BF16. It was produced from source revision 68dafd17ead9b8cf6f85c4f08f7f2f3a1e7b9e5c of StandardOne-3B on 2026-09-25 using llm-compressor 0.14.0 (torch 2.14.0, transformers 5.17.0, compressed-tensors 0.19.0); results below. Served through SGLang…
Open weights
apache-2.0
3.8B parameters
262,144 tokens
transformers