SAVRN
Search Contact SAVRN

Open-weight model · Text classification

kev-0.8b-code-verify-v1

by James jtatman/kev-0.8b-code-verify-v1

kev-0.8b-code-verify-v1 is an open-weight model for text classification from James, released under Apache License 2.0. Its published files total 65.5 MB.

A fine-tune of jaredpalmer/kev-0.8b (first fine-tune): a Jev-style decision model that judges whether code written by a cheap coding model is correct, so a router can keep the answer local or escalate it.

Parameters—
Context—
Weights45.4 MB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Model Card

By James, published under apache-2.0, revision 34ac4c100405.

A fine-tune of jaredpalmer/kev-0.8b (first fine-tune): a Jev-style decision model that judges whether code written by a cheap coding model is correct, so a router can keep the answer local or escalate it. One forward pass, a calibrated probability, no generated text. Trained on execution-labelled data: coder attempts at HumanEvalPack (Python) tasks, labelled by running each task's hidden test suite. The model never sees the hidden tests; its input is what a router can compute itself. The fine-tune binds these exact strings. One noul question (probability that the statement holds): State (an object; Kev renders it as key: value lines): generatededgecasetests comes from a small model writing…

Read James's full model card

A fine-tune of jaredpalmer/kev-0.8b (first fine-tune): a Jev-style decision model that judges whether code written by a cheap coding model is correct, so a router can keep the answer local or escalate it. One forward pass, a calibrated probability, no generated text.

Trained on execution-labelled data: coder attempts at HumanEvalPack (Python) tasks, labelled by running each task's hidden test suite. The model never sees the hidden tests; its input is what a router can compute itself.

How to ask it

The fine-tune binds these exact strings. One noul question (probability that the statement holds):

{"type": "noul", "instructions": "Does the code correctly and completely implement the request, so it would pass a thorough hidden test suite including edge cases? Judge from the request, the code and the checks shown."}

State (an object; Kev renders it as key: value lines):

{"request": "<the task, verbatim>",
  "code": "<the model's code>",
  "checks": {"compiles": true, "defines_requested_function": true, "passes_examples_in_request": true,
             "generated_edge_case_tests": "7 of 8 passed",
             "note": "edge-case tests were written by a small model from the request alone; some may be wrong"}}

generated_edge_case_tests comes from a small model writing ~8 asserts from the request only, then running them. Serve with Kev's runtime (python -m kev.serve --run jtatman/kev-0.8b-code-verify-v1, a TypeSafe /v1/systemone endpoint) or kev.predictors.LocalPredictor. The checkpoint carries its fitted temperature.

Results

Unseen coder (llama-3.1-8b-instruct, never in training): 488 attempts, pass rate 0.57. "Kept local @X% err" = share of attempts a router can accept, highest score first, with at most X% of them wrong. evidence score = (request examples passed + generated-assert pass rate) / 2.

score AUROC Brier ECE kept local @2% err @5% @10%
baseline kev 0.935 0.096 0.12 0.23 0.28 0.48
fine-tuned kev 0.959 0.064 0.05 0.23 0.40 0.63
evidence score 0.944 0.091 0.09 0.00 0.21 0.62
mean(evidence, baseline) 0.952 0.088 0.10 0.11 0.38 0.61
mean(evidence, fine-tuned) 0.957 0.069 0.07 0.15 0.38 0.63

Development (coders seen in training, tasks not seen):

score AUROC Brier ECE kept local @2% err @5% @10%
baseline kev 0.925 0.029 0.05 0.73 0.89 0.95
fine-tuned kev 0.941 0.025 0.03 0.80 0.89 0.95
evidence score 0.960 0.036 0.07 0.79 0.89 0.92
mean(evidence, baseline) 0.957 0.029 0.06 0.84 0.89 0.95
mean(evidence, fine-tuned) 0.961 0.028 0.07 0.83 0.89 0.95

Numbers are on small sets; see the project repo for task-bootstrap confidence intervals.

Recommended use

Average this model's probability with the execution evidence score, and tune the accept threshold per cheap-tier coder on a few hundred of that coder's execution-labelled attempts: a fixed threshold's error rate depends on how often the coder fails.

Training

  • Init: jaredpalmer/kev-0.8b@9a45d25 (LoRA r=16 + pointer head, base Qwen/Qwen3.5-0.8B-Base), Kev's own trainer (kev.train, commit f2bb629d).
  • Data: 1416 execution-labelled records (task-grouped split; llama-3.1-8b held out) + 2000 public decision-v7 replay records against forgetting.
  • One epoch, lr 2e-05, batch 4 x accum 2, gradient checkpointing, bf16 autocast, state limit 1664 tokens.
  • Temperature fitted on a held-out calibration split (min NLL).
  • Hardware: one NVIDIA L4 (Google Colab).

Training data provenance

Code attempts whose correctness this model learned to judge (labels from executing hidden tests):

source model license training rows
or:mistral-small-3.2-24b mistralai/Mistral-Small-3.2-24B-Instruct-2506 Apache-2.0 297
or:qwen3-235b-a22b Qwen/Qwen3-235B-A22B-Instruct-2507 Apache-2.0 292
or:qwen3-coder-30b-a3b Qwen/Qwen3-Coder-30B-A3B-Instruct Apache-2.0 238
qwen2.5-coder:3b Qwen/Qwen2.5-Coder-3B-Instruct Qwen Research License 344
reference-buggy bigcode/humanevalpack buggy solutions MIT 115
reference-canonical bigcode/humanevalpack canonical solutions MIT 115
ternary-qwen3.8-27b prism-ml/Ternary-Bonsai-2-27B-gguf (PTQ1_0; base Qwen/Qwen3.8-27B) Apache-2.0 15

Built with Qwen. Training data includes outputs of Qwen2.5-Coder-3B-Instruct, used under the Qwen Research License Agreement (section 4b). The held-out test coder (Llama 3.1 8B Instruct) was never used for training.

Limitations

  • Python function-level tasks only (HumanEvalPack); other languages, repos and multi-file changes are untested.
  • The coder pool is 8-10 models of 3B-235B parameters; accept thresholds do not transfer across coders.
  • Trained on this question and state layout; other phrasings work less well.

Credits

Kev by Jared Palmer (jaredpalmer/kev, Apache-2.0); Qwen3.5 base (Apache-2.0); HumanEvalPack (bigcode). Project: https://github.com/jtatman/kev-code-verify

Identity and Version

Repository
jtatman/kev-0.8b-code-verify-v1
Publisher
James
Task
Text classification
Modality
Text
Library
peft
Parameters
Not stated by the source
Languages
kev, jev
Revision
34ac4c100405e6e07aaea4000f50dcef686f6de8
First published
2026-09-29
Last updated
2026-09-30

Files and Weights

13 files, 65.5 MB in total. The weights are 2 files totalling 45.4 MB in pt, safetensors.

Weights2 files · 45.4 MB
Configuration5 files · 12.9 KB
Tokenizer2 files · 20.0 MB
Documentation1 file · 5.5 KB
Other2 files · 11.6 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
adapter_model.safetensorsWeights43.3 MB f87f2b743634
head.ptWeights2.1 MB 0cd90db643fc
adapter_config.jsonConfiguration1.3 KB —
run_config.jsonConfiguration464 B —
run_eval.jsonConfiguration3.9 KB —
training_config.jsonConfiguration1.9 KB —
training_metrics.jsonConfiguration5.3 KB —
README.mdDocumentation5.5 KB —
chat_template.jinjaOther7.8 KB —
run_train.logOther3.8 KB —
.gitattributesRepository1.6 KB —
tokenizer.jsonTokenizer20.0 MB 06b9509352d2
tokenizer_config.jsonTokenizer1.1 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
45.4 MB
Download from James

Released by James through its official repository on Hugging Face. Read the license.

Built From

  • Adapter of jaredpalmer/kev-0.8b
  • Derived from jaredpalmer/kev-0.8b

Memory Requirements

PrecisionWeights in memory
As published45.4 MB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About kev-0.8b-code-verify-v1

Can I use kev-0.8b-code-verify-v1 commercially?

Yes. kev-0.8b-code-verify-v1 is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Text classification

finbert

Prosus AI

FinBERT is a pre-trained NLP model to analyze sentiment of financial text. It is built by further training the BERT language model in the finance domain, using a large financial corpus and thereby fine-tuning it for financial sentiment classification. Financial PhraseBank by Malo et al. (2014) is used for fine-tuning. For more details, please see the paper FinBERT: Financial Sentiment Analysis with Pre-trained Language Models and our related blog post on Medium. The model will give softmax outputs for three labels: positive, negative or neutral. About Prosus Prosus is a global consumer internet group and one of the largest technology investors in the world. Operating and investing globally…

Open weights 512 tokens transformers

Model · Text classification

twitter-roberta-base-sentiment-latest

Cardiff NLP

This is a RoBERTa-base model trained on ~124M tweets from January 2018 to December 2021, and finetuned for sentiment analysis with the TweetEval benchmark. The original Twitter-based RoBERTa model can be found here and the original reference paper is TweetEval. This model is suitable for English. 0 -> Negative; 1 -> Neutral; 2 -> Positive This sentiment analysis model has been integrated into TweetNLP. You can access the demo here.

Open weights cc-by-4.0 514 tokens transformers

Model · Text classification

finbert-tone

Yi

FinBERT is a BERT model pre-trained on financial communication text. The purpose is to enhance financial NLP research and practice. It is trained on the following three financial communication corpus. The total corpora size is 4.9B tokens. More technical details on FinBERT: Click Link This released finbert-tone model is the FinBERT model fine-tuned on 10,000 manually annotated (positive, negative, neutral) sentences from analyst reports. This model achieves superior performance on financial tone analysis task. If you are simply interested in using FinBERT for financial tone analysis, give it a try. If you use the model in your academic work, please cite the following paper: Huang, Allen H.…

Open weights 512 tokens transformers

Model · Text classification

ms-marco-MiniLM-L-6-v2

Joshua

https://huggingface.co/cross-encoder/ms-marco-MiniLM-L-6-v2 with ONNX weights to be compatible with Transformers.js. If you haven't already, you can install the Transformers.js JavaScript library from NPM using: Note: Having a separate repo for ONNX weights is intended to be a temporary solution until WebML gains more traction. If you would like to make your models web-ready, we recommend converting to ONNX using Optimum and structuring your repo like this one (with ONNX weights located in a subfolder named onnx).

Open weights 512 tokens transformers.js

Model · Text classification

twitter-xlm-roberta-base-sentiment

Cardiff NLP

This is a multilingual XLM-roBERTa-base model trained on ~198M tweets and finetuned for sentiment analysis. The sentiment fine-tuning was done on 8 languages (Ar, En, Fr, De, Hi, It, Sp, Pt) but it can be used for more languages (see paper for details). This model has been integrated into the TweetNLP library.

Open weights 514 tokens transformers

Model · Text classification

emotion-english-distilroberta-base

Hartmann

With this model, you can classify emotions in English text data. The model was trained on 6 diverse datasets (see Appendix below) and predicts Ekman's 6 basic emotions, plus a neutral class: 1) anger 2) disgust 3) fear 4) joy 5) neutral 6) sadness 7) surprise The model is a fine-tuned checkpoint of DistilRoBERTa-base. For a 'non-distilled' emotion model, please refer to the model card of the RoBERTa-large version. a) Run emotion model with 3 lines of code on single text example using Hugging Face's pipeline command on Google Colab: b) Run emotion model on multiple examples and full datasets (e.g.,.csv files) on Google Colab: Please reach out to [email protected] if you have any…

Open weights 514 tokens transformers