SAVRN
Search Contact SAVRN

Open-weight model · Text classification

MiniCPM5-2B-pjev-LoRA

by Adrian Ciprian Iancu adriandj3/MiniCPM5-2B-pjev-LoRA

MiniCPM5-2B-pjev-LoRA is an open-weight model for text classification from Adrian Ciprian Iancu, released under Apache License 2.0. Its published files total 402.0 MB.

A one-token decision model you can switch on per request, on the MiniCPM5-2B server you already run.

Parameters—
Context—
Weights402.0 MB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Model Card

By Adrian Ciprian Iancu, published under apache-2.0, revision cdef87fd57ed.

A one-token decision model you can switch on per request, on the MiniCPM5-2B server you already run. Kailune-AI/pjev-2b is a strong typed-decision model: give it a state, a question and lettered options, and it returns a calibrated probability for every option in a single forward pass. It ships as a full merged model (about 2.7 GB in Q80). If your agent already serves openbmb/MiniCPM5-2B for chat, a second full model costs VRAM you probably do not have. This repository re-expresses pjev-2b as a rank-64 LoRA over the stock MiniCPM5-2B weights. Load it once next to the chat model, turn it on only for decision requests, and pay about 200 MB of VRAM instead of 2.7 GB. Measured on a private…

Read Adrian Ciprian Iancu's full model card

pjev-2b as a LoRA for MiniCPM5-2B

A one-token decision model you can switch on per request, on the MiniCPM5-2B server you already run.

Kailune-AI/pjev-2b is a strong typed-decision model: give it a state, a question and lettered options, and it returns a calibrated probability for every option in a single forward pass. It ships as a full merged model (about 2.7 GB in Q8_0). If your agent already serves openbmb/MiniCPM5-2B for chat, a second full model costs VRAM you probably do not have.

This repository re-expresses pjev-2b as a rank-64 LoRA over the stock MiniCPM5-2B weights. Load it once next to the chat model, turn it on only for decision requests, and pay about 200 MB of VRAM instead of 2.7 GB.

file for size
pjev-lora-r64-f16.gguf llama.cpp (--lora) 201 MB
adapter_model.safetensors + adapter_config.json PEFT / transformers 201 MB

Results

Measured on a private held-out set of 62 English decisions taken from a real local agent harness: which read-only tool to call now (24), which UI element to act on (22), what state a web page is in (16). The set was written before any of these models were tried and was never used for tuning. One question is worth 1.6 points.

decider (same 62 questions) accuracy tool UI element page state latency per decision
MiniCPM5-2B + this LoRA 93.5% (58/62) 22/24 21/22 15/16 0.22 s median
pjev-2b, full merged model 91.9% (57/62) 22/24 20/22 15/16 0.16 s median
MiniCPM5-2B, same prompt, LoRA off 71.0% 10/24 20/22 14/16 0.17 s
MiniCPM5-2B, best plain prompt + batch calibration 75.8% 15/24 16/22 16/16 0.15 s
LFM2.5-2.6B, best plain prompt + batch calibration 85.5% 0.23 s
LFM2.5-2.6B, full reasoning before answering 91.9% 7.8 s mean

The LoRA picks the same option as the merged pjev-2b on 60 of 62 questions. All wrong answers of the merged model had low confidence (0.34 to 0.61), so a confidence threshold is a natural place to hand off to a slower second stage.

Setup: llama.cpp b10549 (b2e5e9b28), MiniCPM5-2B Q8_0 (official GGUF), KV cache q8_0, GTX 1080 8 GB (Pascal). Latency includes prompt processing of about 340 tokens.

Italian, for the record: 69.4% with the LoRA, 64.5% without. pjev-2b was trained on English; use it in English.

How it was made

  1. Same base, verified. All 87 non-linear tensors of pjev-2b (embeddings, norms, output head) are bit-identical to openbmb/MiniCPM5-2B (revision f97400052a43). The fine-tune lives entirely in the seven linear projections of each layer (q k v o gate up down).
  2. Low rank, measured. The singular values of W_pjev - W_base fall off a cliff at rank 64: 99.9% of the energy of an attention projection sits in the first 64 directions; what is left on the MLP projections is flat, which is bf16 rounding of the merge, not signal.
  3. Truncated SVD per matrix (294 matrices, rank 64, B = U sqrt(S), A = sqrt(S) V^T), saved in PEFT format with lora_alpha = r = 64, then converted with llama.cpp convert_lora_to_gguf.py.

The extraction is an approximation of the merged model, which is why the two rows above differ by one question.

Use it with llama.cpp

Start the server once with the adapter loaded but off by default:

llama-server -m MiniCPM5-2B-Q8_0.gguf \
  --lora pjev-lora-r64-f16.gguf --lora-init-without-apply \
  -c 32768 -ngl 999 -fa on --jinja

Then send decision requests with the adapter at scale 1, and every other request with scale 0:

{"prompt": "<rendered prompt, see below>", "n_predict": 1, "n_probs": 50, "temperature": 0,
 "grammar": "root ::= [A-G]", "lora": [{"id": 0, "scale": 1.0}]}

Watch out (llama.cpp b10549): the LoRA scale set by a request stays on that slot. A later chat request on the same slot without a lora field runs with the adapter still on. Either pass "lora": [{"id": 0, "scale": 0}] on every chat request, or give the decider its own slot (id_slot).

Read the answer from the probabilities of the letter tokens at the first generated position (the n_probs list), not from sampled text. Softmax over the letters in play gives the per-option probability.

Prompt format

The adapter only works with the format pjev-2b was trained on (the "semif" style of its decision_core.py): a fixed system message, a JSON user turn, and the chat template with thinking disabled.

System message, verbatim:

Apply the supplied criterion to the supplied evidence. Choose exactly one listed option. Respond with only its uppercase letter, with no explanation or reasoning.

User message (one line of JSON):

{"evidence": "User message: \"will it rain in Lisbon tomorrow morning?\"\nTool results in this turn: none.",
 "criterion": "Which read-only tool must be called NOW, before answering the user?",
 "options": [{"letter": "A", "description": "weather: forecast for a named place, next 7 days"},
             {"letter": "B", "description": "news_search: news about a subject within a window of days"},
             {"letter": "C", "description": "none: no tool is needed now, answer directly"}]}

Render with the MiniCPM5 chat template and enable_thinking=False (the assistant turn opens with an empty <think>\n\n</think>\n\n block); the next token is the letter.

Use it with transformers + PEFT

from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

tok = AutoTokenizer.from_pretrained("openbmb/MiniCPM5-2B")
model = AutoModelForCausalLM.from_pretrained("openbmb/MiniCPM5-2B", torch_dtype="auto")
model = PeftModel.from_pretrained(model, "adriandj3/MiniCPM5-2B-pjev-LoRA")

Build the messages as above, call tok.apply_chat_template(..., add_generation_prompt=True, enable_thinking=False), run one forward pass and softmax the logits of the option letters at the last position.

Limitations

  • English in practice. Inherited from pjev-2b.
  • Small benchmark. 62 held-out questions from one harness. The difference with the merged model (one question) and with LFM2.5 full reasoning is inside the noise; the difference with the base MiniCPM5 prompts is not.
  • Approximation. Rank-64 SVD of the merged delta: very close to pjev-2b, not identical.
  • No abstention. Like pjev-2b, it always picks one of the listed options. Add a "none" option when that is a valid answer, and use the confidence.
  • Not evaluated on JevBench in this form. See the pjev-2b card for its JevBench numbers.

License and credits

  • Weights: Apache-2.0. This is a derivative of Kailune-AI/pjev-2b (Apache-2.0), itself a fine-tune of openbmb/MiniCPM5-2B (Apache-2.0). Change made here: the merged fine-tune was re-expressed as a rank-64 LoRA by truncated SVD of the weight delta. See NOTICE.
  • The decision format and letter readout come from the Eikos project (MIT, Caio Vicentino), as used by pjev-2b.
  • Not affiliated with OpenBMB, Kailune-AI or the Eikos project. Made while building Vergilius, a local assistant for modest machines.

Identity and Version

Repository
adriandj3/MiniCPM5-2B-pjev-LoRA
Publisher
Adrian Ciprian Iancu
Task
Text classification
Modality
Text
Library
peft
Parameters
Not stated by the source
Languages
en
Revision
cdef87fd57ed2d394fea6125379006691a49d249
First published
2026-10-03
Last updated
2026-10-03

Files and Weights

6 files, 402.0 MB in total. The weights are 2 files totalling 402.0 MB in gguf, safetensors.

Weights2 files · 402.0 MB
Configuration1 file · 335 B
Documentation2 files · 8.6 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
adapter_model.safetensorsWeights201.0 MB e32886477007
pjev-lora-r64-f16.ggufWeights201.0 MB 1b4a59375f30
adapter_config.jsonConfiguration335 B —
NOTICEDocumentation960 B —
README.mdDocumentation7.7 KB —
.gitattributesRepository1.6 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
402.0 MB
Download from Adrian Ciprian Iancu

Released by Adrian Ciprian Iancu through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published402.0 MB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About MiniCPM5-2B-pjev-LoRA

Can I use MiniCPM5-2B-pjev-LoRA commercially?

Yes. MiniCPM5-2B-pjev-LoRA is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Text classification

finbert

Prosus AI

FinBERT is a pre-trained NLP model to analyze sentiment of financial text. It is built by further training the BERT language model in the finance domain, using a large financial corpus and thereby fine-tuning it for financial sentiment classification. Financial PhraseBank by Malo et al. (2014) is used for fine-tuning. For more details, please see the paper FinBERT: Financial Sentiment Analysis with Pre-trained Language Models and our related blog post on Medium. The model will give softmax outputs for three labels: positive, negative or neutral. About Prosus Prosus is a global consumer internet group and one of the largest technology investors in the world. Operating and investing globally…

Open weights 512 tokens transformers

Model · Text classification

twitter-roberta-base-sentiment-latest

Cardiff NLP

This is a RoBERTa-base model trained on ~124M tweets from January 2018 to December 2021, and finetuned for sentiment analysis with the TweetEval benchmark. The original Twitter-based RoBERTa model can be found here and the original reference paper is TweetEval. This model is suitable for English. 0 -> Negative; 1 -> Neutral; 2 -> Positive This sentiment analysis model has been integrated into TweetNLP. You can access the demo here.

Open weights cc-by-4.0 514 tokens transformers

Model · Text classification

finbert-tone

Yi

FinBERT is a BERT model pre-trained on financial communication text. The purpose is to enhance financial NLP research and practice. It is trained on the following three financial communication corpus. The total corpora size is 4.9B tokens. More technical details on FinBERT: Click Link This released finbert-tone model is the FinBERT model fine-tuned on 10,000 manually annotated (positive, negative, neutral) sentences from analyst reports. This model achieves superior performance on financial tone analysis task. If you are simply interested in using FinBERT for financial tone analysis, give it a try. If you use the model in your academic work, please cite the following paper: Huang, Allen H.…

Open weights 512 tokens transformers

Model · Text classification

ms-marco-MiniLM-L-6-v2

Joshua

https://huggingface.co/cross-encoder/ms-marco-MiniLM-L-6-v2 with ONNX weights to be compatible with Transformers.js. If you haven't already, you can install the Transformers.js JavaScript library from NPM using: Note: Having a separate repo for ONNX weights is intended to be a temporary solution until WebML gains more traction. If you would like to make your models web-ready, we recommend converting to ONNX using Optimum and structuring your repo like this one (with ONNX weights located in a subfolder named onnx).

Open weights 512 tokens transformers.js

Model · Text classification

twitter-xlm-roberta-base-sentiment

Cardiff NLP

This is a multilingual XLM-roBERTa-base model trained on ~198M tweets and finetuned for sentiment analysis. The sentiment fine-tuning was done on 8 languages (Ar, En, Fr, De, Hi, It, Sp, Pt) but it can be used for more languages (see paper for details). This model has been integrated into the TweetNLP library.

Open weights 514 tokens transformers

Model · Text classification

emotion-english-distilroberta-base

Hartmann

With this model, you can classify emotions in English text data. The model was trained on 6 diverse datasets (see Appendix below) and predicts Ekman's 6 basic emotions, plus a neutral class: 1) anger 2) disgust 3) fear 4) joy 5) neutral 6) sadness 7) surprise The model is a fine-tuned checkpoint of DistilRoBERTa-base. For a 'non-distilled' emotion model, please refer to the model card of the RoBERTa-large version. a) Run emotion model with 3 lines of code on single text example using Hugging Face's pipeline command on Google Colab: b) Run emotion model on multiple examples and full datasets (e.g.,.csv files) on Google Colab: Please reach out to [email protected] if you have any…

Open weights 514 tokens transformers