Phi-4-mini Clinical MLX (4-Bit Merged)
A specialized 3.8B biomedical & clinical reasoning model built on Microsoft's Phi-4-mini-instruct, optimized natively for Apple Silicon Metal acceleration via Apple MLX.
The model underwent a 3-stage transfer learning curriculum:
1. Stage 1 (STEM Foundation): 116,000 instruction pairs across NCERT Classes 6–12 (Physics, Chemistry, Biology) eliminating foundational science hallucinations.
2. Stage 2 (PubMed 2026 Evidence): 12 recent 2026 clinical update archives from NCBI FTP covering survival outcomes (OS, PFS, HR), targeted therapeutics, and clinical trial endpoints.
3. Stage 3 (Comprehensive Internal Medicine): Balanced multi-specialty clinical curriculum (cardiology, nephrology, endocrinology, pulmonology) with an active oncology replay buffer.
The LoRA adapter weights have been permanently fused into the base 4-bit weights (mlx_lm fuse) to deliver zero-latency execution.
Benchmark Results (PubMedQA)
Evaluated on 50 biomedical research decision tasks from PubMedQA:
| Model |
Accuracy |
Score |
Avg Latency |
Relative Improvement |
| Base Phi-4-mini (4-bit) |
26.0% |
13 / 50 |
1.02s / question |
Baseline |
| Phi-4-mini Clinical MLX (Merged) |
40.0% |
20 / 50 |
0.93s / question |
+53.8% relative gain |
Quickstart with Apple MLX
Install mlx-lm:
pip install mlx-lm
CLI Generation
python -m mlx_lm.generate \
--model <repo_id> \
--prompt "<|user|>\nWhat are the first-line therapeutic recommendations for heart failure with preserved ejection fraction (HFpEF)?<|end|>\n<|assistant|>\n" \
--max-tokens 512
Python API
from mlx_lm import load, generate
model, tokenizer = load("<repo_id>")
prompt = "<|user|>\nSummarize the mechanism of action of SGLT2 inhibitors in diabetic kidney disease.<|end|>\n<|assistant|>\n"
response = generate(model, tokenizer, prompt=prompt, max_tokens=300)
print(response)
Local OpenAI-Compatible Server
python -m mlx_lm.server --model <repo_id> --port 8080
Clinical Disclaimer
This model is intended solely for biomedical research, educational exploration, and experimental evaluation. It is not an FDA-cleared medical device and must not be used as a substitute for professional clinical judgment, diagnosis, or treatment.
Verified Medical Benchmark Results
| Benchmark |
Scope |
Tested Samples |
Accuracy |
Evaluation Hardware |
| PubMedQA |
Clinical Trial Evidence Decisions |
100 |
49.0% |
Apple Silicon Metal GPU |
| MedQA (USMLE) |
Medical Board Diagnostic Cases |
100 |
53.0% |
Apple Silicon Metal GPU |