SAVRN
Search Contact SAVRN

Open-weight model

CAT-UT

by Avrova Donz AvrovaDonz/CAT-UT

CAT-UT is an open-weight model from Avrova Donz, released under mpl-2.0. Its published files total 2.6 GB.

This repository contains QPR artifacts for the dense, unpruned openbmb/MiniCPM5-2B model. The quantizer uses dense int4 weights with group size 128 and GPTQ calibration; it does not apply 2:4 pruning.

Parameters—
Context—
Weights2.6 GB
Licensempl-2.0
AccessOpen weights
Monthly Downloads—

Model Card

By Avrova Donz, published under mpl-2.0, revision 3765700a47c5.

This repository contains QPR artifacts for the dense, unpruned openbmb/MiniCPM5-2B model. The quantizer uses dense int4 weights with group size 128 and GPTQ calibration; it does not apply 2:4 pruning. The recommended file is dense-int4-distill2000-v2-best.qpr. The minicpm5-2b-dense-int4-g128.qpr file is the matching dense-int4 baseline. Each QPR file is about 1.30 GB and is stored with Git LFS. The candidate was selected from a 2,000-update distillation run starting at the dense-int4 baseline. It was selected at update 775 and independently reloaded before scoring. On 100 held-out Chinese and English prompts with 512 new tokens, direct gpt-6-sol judging gave 39 baseline wins, 58 candidate…

Read Avrova Donz's full model card

CAT-UT dense int4 checkpoints

This repository contains QPR artifacts for the dense, unpruned openbmb/MiniCPM5-2B model. The quantizer uses dense int4 weights with group size 128 and GPTQ calibration; it does not apply 2:4 pruning.

The recommended file is dense-int4-distill2000-v2-best.qpr. The minicpm5-2b-dense-int4-g128.qpr file is the matching dense-int4 baseline. Each QPR file is about 1.30 GB and is stored with Git LFS.

Evaluation

The candidate was selected from a 2,000-update distillation run starting at the dense-int4 baseline. It was selected at update 775 and independently reloaded before scoring.

Artifact Wide code NLL Wide chat NLL Wide tool NLL
Dense int4 baseline 1.0654 5.1728 2.5966
Distilled candidate 1.0323 5.0707 2.6200

On 100 held-out Chinese and English prompts with 512 new tokens, direct gpt-6-sol judging gave 39 baseline wins, 58 candidate wins, and 3 ties. The Chinese subset was 15/34/1 and the English subset was 24/24/2 (baseline/candidate/tie). A two-sided exact binomial test gives p=0.067; this is evidence of a promising Chinese-chat direction, not a conclusive overall improvement. Automatic repetition changed from 0.0528 to 0.0335 and truncation from 0.6400 to 0.4525, while language-mixing increased from 0.0119 to 0.0487. Independent executable code evaluation for this candidate has not yet established a gain.

Loading

QPR is a custom artifact format and is not loaded directly by transformers.from_pretrained. Clone the qprsi source, install its dependencies, and load the base model first:

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from qprsi.pipeline import collect_plans, install_artifact

model_dir = "/path/to/MiniCPM5-2B"
artifact = "/path/to/dense-int4-distill2000-v2-best.qpr"

tokenizer = AutoTokenizer.from_pretrained(model_dir)
model = AutoModelForCausalLM.from_pretrained(
    model_dir, dtype=torch.bfloat16, low_cpu_mem_usage=True
)
model.eval()
plans = collect_plans(model)
install_artifact(plans, artifact)

For the reproducible evaluation commands, use the dense-distill-2000 branch of the source repository:

python -m qprsi.chat_eval \
  --model /path/to/MiniCPM5-2B \
  --student-artifact /path/to/dense-int4-distill2000-v2-best.qpr \
  --max-new-tokens 512

python -m qprsi.code_eval \
  --model /path/to/MiniCPM5-2B \
  --artifacts /path/to/dense-int4-distill2000-v2-best.qpr \
  --max-new-tokens 512 \
  --report code-eval.json

The QPR loader must use the same model architecture and tokenizer revision as the base model. The artifact is intended for research evaluation and is not a drop-in Transformers or vLLM checkpoint.

Training details

Distillation used cached, visible teacher responses from gpt-6-sol over an authorized OpenAI-compatible endpoint. The run used batch size 16, learning rate 3e-6, KL beta 0.05, last-block training, 1,075 executed updates, and wide-probe selection with a 1% code/tool regression gate. The selected update was 775. No API credentials are included in the artifacts or reports.

Limitations

  • The chat comparison used one 100-prompt set and one judge model; the result is not a benchmark guarantee.
  • Generation was capped at 512 new tokens, and many prompts still spend part of the budget in the model's visible reasoning section.
  • On the independent executable 10-task code check, the baseline passed 3/10 and the candidate passed 4/10; both passed 0/5 English tasks. This small sample does not establish a general code improvement.
  • A separate 200-update rich-task pilot matched the baseline at 3/10 on that code check and lost its 100-prompt chat comparison (41 pilot wins, 51 baseline wins, 8 ties), so that experimental QPR is not included here.
  • The QPR artifact does not contain the 5 GB BF16 base model. Download the base model separately and comply with its terms.

Licensing

The upstream openbmb/MiniCPM5-2B weights are released under Apache-2.0. The qprsi conversion and evaluation source is released under AGPL-3.0-or-later. The repository metadata currently declares MPL-2.0 for this CAT-UT repository; downstream users should review all applicable terms before redistribution.

Identity and Version

Repository
AvrovaDonz/CAT-UT
Publisher
Avrova Donz
Task
Not stated by the source
Modality
Other
Library
Not stated by the source
Parameters
Not stated by the source
Languages
zh, en
Revision
3765700a47c592bd8c171ef5ac09eb009fcef871
First published
2026-09-25
Last updated
2026-09-25

Files and Weights

4 files, 2.6 GB in total.

Documentation1 file · 4.5 KB
Other2 files · 2.6 GB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
README.mdDocumentation4.5 KB —
dense-int4-distill2000-v2-best.qprOther1.3 GB 306d6475e1b0
minicpm5-2b-dense-int4-g128.qprOther1.3 GB 7b0ad4afe9ea
.gitattributesRepository1.6 KB —

License and Download

License
mpl-2.0
Access
Open weights, no gate
Download from Avrova Donz

Released by Avrova Donz through its official repository on Hugging Face.

Built From

Questions About CAT-UT

What license is CAT-UT released under?

mpl-2.0, as its publisher declares it. Read the license text before commercial use.