FinBERT is a pre-trained NLP model to analyze sentiment of financial text. It is built by further training the BERT language model in the finance domain, using a large financial corpus and thereby fine-tuning it for financial sentiment classification. Financial PhraseBank by Malo et al. (2014) is used for fine-tuning. For more details, please see the paper FinBERT: Financial Sentiment Analysis with Pre-trained Language Models and our related blog post on Medium. The model will give softmax outputs for three labels: positive, negative or neutral. About Prosus Prosus is a global consumer internet group and one of the largest technology investors in the world. Operating and investing globally…
kev-4b is an open-weight model for text classification from Jared Palmer, released under Apache License 2.0. Its published files total 153.4 MB. It draws 18 downloads a month.
kev-4b is a decision model: one document (the state) and a set of typed questions in, a probability distribution per question out, in one forward pass. No text generation.
Model Card
By Jared Palmer, published under apache-2.0, revision 2bd3bb9a9957.
kev-4b is a decision model: one document (the state) and a set of typed questions in, a probability distribution per question out, in one forward pass. No text generation. It is a LoRA adapter (r=16) plus a pointer head on Qwen/Qwen3-4B-Base, serving TypeSafe's public /v1/systemone contract. Research preview, not a versioned release. It is the best 4B checkpoint under a frozen, checksummed protocol after ~40 controlled 4B trials, and the first kev whose out-of-domain accuracy is within ten points of Jev on the same items. This checkpoint clears the release screen we set in advance (held-out policy pairs 0.73 both-correct, screen 70%), but the other two seeds of the same recipe do not (0.62…
Read Jared Palmer's full model card
kev-4b — research preview
kev-4b is a decision model: one document (the state) and a set of typed questions in, a probability distribution per question out, in one forward pass. No text generation. It is a LoRA adapter (r=16) plus a pointer head on Qwen/Qwen3-4B-Base, serving TypeSafe's public /v1/systemone contract.
Research preview, not a versioned release. It is the best 4B checkpoint under a frozen, checksummed protocol after ~40 controlled 4B trials, and the first kev whose out-of-domain accuracy is within ten points of Jev on the same items. This checkpoint clears the release screen we set in advance (held-out policy pairs 0.73 both-correct, screen 70%), but the other two seeds of the same recipe do not (0.62, 0.67), and our rule is every seed; so it stays a preview.
- Hub:
jaredpalmer/kev-4b(this repo; trialv7-rc3/01-trial-1) - Code, suites, every trial with hashes and paired bootstraps: github.com/jaredpalmer/kev —
PLAN.md,runs/leaderboard.md
Results (same frozen items for every row)
| kev-0.5b | kev-0.6b preview | kev-4b preview | Jev | |
|---|---|---|---|---|
| in-distribution accuracy (decision-v4 dev, 1,200 q) | 0.712 | 0.805 | 0.854 | 0.845 |
| out-of-domain accuracy (transfer-v4 dev, 560 q) | 0.575 | 0.598 | 0.790 | 0.857 |
| out-of-domain Brier | 0.50 | 0.521 | 0.328 | 0.211 |
| confident errors out of domain (p ≥ 0.9 and wrong) | – | 5.2% | 8.2% | 3.7% |
| held-out policy structures, both siblings correct | – | 0.11 | 0.73 | 0.86 |
| option-order flip rate | 0.21 | 0.02 | 0.06 | 0.00 |
Per-source out-of-domain accuracy (kev-4b / Jev): QNLI 0.89 / 0.93, SciQ 0.99 / 0.99, TweetEval-offensive 0.75 / 0.81, PAWS 0.72 / 0.79, MMLU 0.65 / 0.90, Emotion 0.66 / 0.59, deadline (3-level date arithmetic) 0.53 / 0.93, (A and B) or not C 0.97 / 0.97, if A then not B else C 0.88 / 0.78.
Seeds: three seeds on decision-v7: transfer 0.773 / 0.790 / 0.770, held-out pairs 0.62 / 0.73 / 0.67; this checkpoint is seed 1, selected on development transfer accuracy. Trained on decision-v7 (10k public records + 896 policy records over nine template families incl. four ordinal Score threshold families + 1,680 records from 60 random rule structures with negation anywhere); development/test items are byte-identical to v4, so every number here is comparable with earlier previews.
Locked test, one exploratory read (runs/locked/kev-4b-v7-preview-ungated/, labelled ungated because the screen is not met on every seed): in-distribution 0.856 (Brier 0.211), out-of-domain 0.806 (Brier 0.294, confident errors 6.6%, held-out pairs 0.66). This partition will not be read again for this checkpoint.
What we learned building it
- Capacity dominates out of domain. With public examples and synthetic budget held equal, 0.6B → 4B is +14–19 pp; 4B → 8B is +1–7 pp.
- Fine-tuning erodes base capability, and the learning rate controls it. The 4B base, zero-shot with a letter readout, scores 0.688 on the same MMLU items and 0.787 on PAWS; the default recipe (lr 2e-4) trained down to 0.60–0.66 / 0.56–0.71. Lowering lr to 5e-5 recovers most of it and is the single largest recipe improvement we found; fewer LoRA target modules and smaller ranks help less.
- More public training data raises in-distribution accuracy and lowers transfer at 4B (10k vs 3.4k records: −3 pp). Knowledge MCQ sources (ARC, OpenBookQA, CommonsenseQA) raise in-distribution accuracy to 0.86 without moving transfer.
- Programmatic contrastive policy pairs teach the trained rule structures (both-correct 0.85–1.0) but transfer to unseen structures only partially (0.5–0.6 at 4B, 0.03–0.11 at 0.6B).
Known limits
- Held-out policy reasoning (unseen rule compositions, date arithmetic with grace periods) is far from Jev.
- Product-shaped questions with no training analogue are not guaranteed: on the TypeSafe docs example ("two charges on my card" → Is there a billing problem?) this checkpoint answers 0.22 while kev-0.6b answers 0.97 and picks the return reason (wrong size, 0.54) correctly. Lower drift from the base means fewer task-specific priors; measure on your own inputs.
- Out-of-domain probabilities are usable but not calibrated (raw ECE 0.096); temperature fitted in-domain does not transfer.
- 4B fp32 needs ~16 GB; on a 32 GB Mac use
KEV_DTYPE=bf16. Latency on an H100 is ~45 ms per packed request; on an M5 several hundred ms.
Training
Frozen suite evals/v4/decision-v4: 10,000 public records (1,000 per source, ten sources) plus two programmatic policy arms of 448 records, two epochs, LoRA r=16 on attention and MLP projections, pointer head from scratch, cross-entropy on the option distribution, lr 5e-5 (OneCycle), effective batch 8, bf16 autocast with fp32 master weights, gradient checkpointing, one H100 (~40 min). Augmentation: option permutation, none-of-the-above insertion, distractors, none minimal pairs on 25% of Choice records. No Jev outputs were used for training.
Evaluation protocol
Development partitions select models; the locked test partition is read at most once per candidate. Every number carries suite hash, code hashes, and git commit in result.json. See PLAN.md for the corrections we made to our own earlier claims.
Use
uv run --extra serve python -m kev.serve --run jaredpalmer/kev-4b --port 8008 # KEV_DTYPE=bf16 on a 32 GB Mac
Any TypeSafe-compatible client works: TypeSafeClient(api_key="local", base_url="http://127.0.0.1:8008", model="kev-latest").
License
Apache-2.0 for the adapter and head; Qwen3 base is Apache-2.0; datasets carry their own licenses.
Identity and Version
- Repository
- jaredpalmer/kev-4b
- Publisher
- Jared Palmer
- Task
- Text classification
- Modality
- Text
- Library
- peft
- Parameters
- Not stated by the source
- Languages
- en
- Revision
- 2bd3bb9a9957aff9a4803be7ce91a521cdff0a31
- First published
- 2026-09-19
- Last updated
- 2026-09-20
Files and Weights
16 files, 153.4 MB in total. The weights are 2 files totalling 137.4 MB in pt, safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| adapter_model.safetensors | Weights | 132.2 MB | 3fa10d8646f5 |
| head.pt | Weights | 5.2 MB | dad87f7e60c8 |
| adapter_config.json | Configuration | 1.2 KB | — |
| added_tokens.json | Configuration | 707 B | — |
| provenance.json | Configuration | 2.9 KB | — |
| result.json | Configuration | 68.3 KB | — |
| special_tokens_map.json | Configuration | 616 B | — |
| training_config.json | Configuration | 1.0 KB | — |
| training_metrics.json | Configuration | 324 B | — |
| README.md | Documentation | 7.1 KB | — |
| train.log | Other | 19.9 KB | — |
| .gitattributes | Repository | 1.6 KB | — |
| merges.txt | Tokenizer | 1.7 MB | — |
| tokenizer.json | Tokenizer | 11.4 MB | aeb13307a71a |
| tokenizer_config.json | Tokenizer | 5.4 KB | — |
| vocab.json | Tokenizer | 2.8 MB | — |
License and Download
- License
- apache-2.0
- Access
- Open weights, no gate
- Download size
- 137.4 MB
Released by Jared Palmer through its official repository on Hugging Face. Read the license.
Built From
- Adapter of Qwen/Qwen3-4B-Base
- Derived from Qwen/Qwen3-4B-Base
- Trained on (disclosed) CogComp/trec
- Trained on (disclosed) SetFit/amazon_reviews_multi_en
- Trained on (disclosed) SetFit/sst5
- Trained on (disclosed) Yelp/yelp_review_full
- Trained on (disclosed) fancyzhx/ag_news
- Trained on (disclosed) fancyzhx/dbpedia_14
- Trained on (disclosed) google/boolq
- Trained on (disclosed) legacy-datasets/banking77
- Trained on (disclosed) nyu-mll/multi_nli
- Trained on (disclosed) stanfordnlp/imdb
Evaluations
Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.
| Benchmark | Conditions | Result | Reported by | Revision | Date |
|---|---|---|---|---|---|
| decision-v4 development (1,204 records; ten trained public sources + programmatic policy pairs) | Task typed decision (choice / noul / score)Metric ECE, raw probabilitiesComparison conditions not established | 0.065 | jaredpalmer Publisher reported |
Evaluated revision not stated | — |
| decision-v4 development (1,204 records; ten trained public sources + programmatic policy pairs) | Task typed decision (choice / noul / score)Metric accuracyComparison conditions not established | 0.854 | jaredpalmer Publisher reported |
Evaluated revision not stated | — |
| transfer-v4 development (764 records; six never-trained sources + held-out policy structures) | Task typed decisionMetric accuracyComparison conditions not established | 0.79 | jaredpalmer Publisher reported |
Evaluated revision not stated | — |
| transfer-v4 development (764 records; six never-trained sources + held-out policy structures) | Task typed decisionMetric brier_scoreComparison conditions not established | 0.328 | jaredpalmer Publisher reported |
Evaluated revision not stated | — |
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 137.4 MB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Built on This Model
- Quantized fromKev-4B-MLX-Serve-8bit
- Derived fromKev-4B-MLX-Serve-8bit
- Adapter ofkev-4b-code-verify-v1
- Derived fromkev-4b-code-verify-v1
Questions About kev-4b
Can I use kev-4b commercially?
Yes. kev-4b is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.
Similar Models
This is a RoBERTa-base model trained on ~124M tweets from January 2018 to December 2021, and finetuned for sentiment analysis with the TweetEval benchmark. The original Twitter-based RoBERTa model can be found here and the original reference paper is TweetEval. This model is suitable for English. 0 -> Negative; 1 -> Neutral; 2 -> Positive This sentiment analysis model has been integrated into TweetNLP. You can access the demo here.
FinBERT is a BERT model pre-trained on financial communication text. The purpose is to enhance financial NLP research and practice. It is trained on the following three financial communication corpus. The total corpora size is 4.9B tokens. More technical details on FinBERT: Click Link This released finbert-tone model is the FinBERT model fine-tuned on 10,000 manually annotated (positive, negative, neutral) sentences from analyst reports. This model achieves superior performance on financial tone analysis task. If you are simply interested in using FinBERT for financial tone analysis, give it a try. If you use the model in your academic work, please cite the following paper: Huang, Allen H.…
https://huggingface.co/cross-encoder/ms-marco-MiniLM-L-6-v2 with ONNX weights to be compatible with Transformers.js. If you haven't already, you can install the Transformers.js JavaScript library from NPM using: Note: Having a separate repo for ONNX weights is intended to be a temporary solution until WebML gains more traction. If you would like to make your models web-ready, we recommend converting to ONNX using Optimum and structuring your repo like this one (with ONNX weights located in a subfolder named onnx).
This is a multilingual XLM-roBERTa-base model trained on ~198M tweets and finetuned for sentiment analysis. The sentiment fine-tuning was done on 8 languages (Ar, En, Fr, De, Hi, It, Sp, Pt) but it can be used for more languages (see paper for details). This model has been integrated into the TweetNLP library.
With this model, you can classify emotions in English text data. The model was trained on 6 diverse datasets (see Appendix below) and predicts Ekman's 6 basic emotions, plus a neutral class: 1) anger 2) disgust 3) fear 4) joy 5) neutral 6) sadness 7) surprise The model is a fine-tuned checkpoint of DistilRoBERTa-base. For a 'non-distilled' emotion model, please refer to the model card of the RoBERTa-large version. a) Run emotion model with 3 lines of code on single text example using Hugging Face's pipeline command on Google Colab: b) Run emotion model on multiple examples and full datasets (e.g.,.csv files) on Google Colab: Please reach out to [email protected] if you have any…
