CAT-UT is an open-weight model from Avrova Donz, released under mpl-2.0. Its published files total 2.6 GB.
This repository contains QPR artifacts for the dense, unpruned openbmb/MiniCPM5-2B model. The quantizer uses dense int4 weights with group size 128 and GPTQ calibration; it does not apply 2:4 pruning.
Model Card
By Avrova Donz, published under mpl-2.0, revision 3765700a47c5.
This repository contains QPR artifacts for the dense, unpruned openbmb/MiniCPM5-2B model. The quantizer uses dense int4 weights with group size 128 and GPTQ calibration; it does not apply 2:4 pruning. The recommended file is dense-int4-distill2000-v2-best.qpr. The minicpm5-2b-dense-int4-g128.qpr file is the matching dense-int4 baseline. Each QPR file is about 1.30 GB and is stored with Git LFS. The candidate was selected from a 2,000-update distillation run starting at the dense-int4 baseline. It was selected at update 775 and independently reloaded before scoring. On 100 held-out Chinese and English prompts with 512 new tokens, direct gpt-6-sol judging gave 39 baseline wins, 58 candidate…
Read Avrova Donz's full model card
CAT-UT dense int4 checkpoints
This repository contains QPR artifacts for the dense, unpruned
openbmb/MiniCPM5-2B model. The quantizer uses dense int4 weights with
group size 128 and GPTQ calibration; it does not apply 2:4 pruning.
The recommended file is dense-int4-distill2000-v2-best.qpr. The
minicpm5-2b-dense-int4-g128.qpr file is the matching dense-int4 baseline.
Each QPR file is about 1.30 GB and is stored with Git LFS.
Evaluation
The candidate was selected from a 2,000-update distillation run starting at the dense-int4 baseline. It was selected at update 775 and independently reloaded before scoring.
| Artifact | Wide code NLL | Wide chat NLL | Wide tool NLL |
|---|---|---|---|
| Dense int4 baseline | 1.0654 | 5.1728 | 2.5966 |
| Distilled candidate | 1.0323 | 5.0707 | 2.6200 |
On 100 held-out Chinese and English prompts with 512 new tokens, direct
gpt-6-sol judging gave 39 baseline wins, 58 candidate wins, and 3 ties.
The Chinese subset was 15/34/1 and the English subset was 24/24/2
(baseline/candidate/tie). A two-sided exact binomial test gives p=0.067;
this is evidence of a promising Chinese-chat direction, not a conclusive
overall improvement. Automatic repetition changed from 0.0528 to 0.0335 and
truncation from 0.6400 to 0.4525, while language-mixing increased from 0.0119
to 0.0487. Independent executable code evaluation for this candidate has not
yet established a gain.
Loading
QPR is a custom artifact format and is not loaded directly by
transformers.from_pretrained. Clone the qprsi source,
install its dependencies, and load the base model first:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from qprsi.pipeline import collect_plans, install_artifact
model_dir = "/path/to/MiniCPM5-2B"
artifact = "/path/to/dense-int4-distill2000-v2-best.qpr"
tokenizer = AutoTokenizer.from_pretrained(model_dir)
model = AutoModelForCausalLM.from_pretrained(
model_dir, dtype=torch.bfloat16, low_cpu_mem_usage=True
)
model.eval()
plans = collect_plans(model)
install_artifact(plans, artifact)
For the reproducible evaluation commands, use the dense-distill-2000
branch of the source repository:
python -m qprsi.chat_eval \
--model /path/to/MiniCPM5-2B \
--student-artifact /path/to/dense-int4-distill2000-v2-best.qpr \
--max-new-tokens 512
python -m qprsi.code_eval \
--model /path/to/MiniCPM5-2B \
--artifacts /path/to/dense-int4-distill2000-v2-best.qpr \
--max-new-tokens 512 \
--report code-eval.json
The QPR loader must use the same model architecture and tokenizer revision as the base model. The artifact is intended for research evaluation and is not a drop-in Transformers or vLLM checkpoint.
Training details
Distillation used cached, visible teacher responses from gpt-6-sol over an
authorized OpenAI-compatible endpoint. The run used batch size 16, learning
rate 3e-6, KL beta 0.05, last-block training, 1,075 executed updates, and
wide-probe selection with a 1% code/tool regression gate. The selected update
was 775. No API credentials are included in the artifacts or reports.
Limitations
- The chat comparison used one 100-prompt set and one judge model; the result is not a benchmark guarantee.
- Generation was capped at 512 new tokens, and many prompts still spend part of the budget in the model's visible reasoning section.
- On the independent executable 10-task code check, the baseline passed 3/10 and the candidate passed 4/10; both passed 0/5 English tasks. This small sample does not establish a general code improvement.
- A separate 200-update rich-task pilot matched the baseline at 3/10 on that code check and lost its 100-prompt chat comparison (41 pilot wins, 51 baseline wins, 8 ties), so that experimental QPR is not included here.
- The QPR artifact does not contain the 5 GB BF16 base model. Download the base model separately and comply with its terms.
Licensing
The upstream openbmb/MiniCPM5-2B weights are released under Apache-2.0.
The qprsi conversion and evaluation source is released under
AGPL-3.0-or-later. The repository metadata currently declares MPL-2.0 for
this CAT-UT repository; downstream users should review all applicable terms
before redistribution.
Identity and Version
- Repository
- AvrovaDonz/CAT-UT
- Publisher
- Avrova Donz
- Task
- Not stated by the source
- Modality
- Other
- Library
- Not stated by the source
- Parameters
- Not stated by the source
- Languages
- zh, en
- Revision
- 3765700a47c592bd8c171ef5ac09eb009fcef871
- First published
- 2026-09-25
- Last updated
- 2026-09-25
Files and Weights
4 files, 2.6 GB in total.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| README.md | Documentation | 4.5 KB | — |
| dense-int4-distill2000-v2-best.qpr | Other | 1.3 GB | 306d6475e1b0 |
| minicpm5-2b-dense-int4-g128.qpr | Other | 1.3 GB | 7b0ad4afe9ea |
| .gitattributes | Repository | 1.6 KB | — |
License and Download
- License
- mpl-2.0
- Access
- Open weights, no gate
Released by Avrova Donz through its official repository on Hugging Face.
Built From
- Derived from openbmb/MiniCPM5-2B
Questions About CAT-UT
What license is CAT-UT released under?
mpl-2.0, as its publisher declares it. Read the license text before commercial use.