SAVRN
Search Contact SAVRN

Open-weight model · Text classification

auto-200m-2-int8

by SSH ProCreations/auto-200m-2-int8

auto-200m-2-int8 is an open-weight model for text classification from SSH, released under Apache License 2.0. It has 65,536-token context. Its published files total 154.0 MB.

This is auto-200m-2 with its weights stored as 8-bit integers. The file is 150 MB instead of 299 MB, and it gets the same benchmark score: 2,890/3,000, with 53 false approvals and 57 false denials. Only 2 of the 3,000 decisions differ from the BF16 model.

Parameters—
Context65,536
Weights150.1 MB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Model Card

By SSH, published under apache-2.0, revision 2501a22901e8.

This is auto-200m-2 with its weights stored as 8-bit integers. The file is 150 MB instead of 299 MB, and it gets the same benchmark score: 2,890/3,000, with 53 false approvals and 57 false denials. Only 2 of the 3,000 decisions differ from the BF16 model. auto-200m-2 is a 149.6M-parameter ModernBERT classifier. It reads an AI agent's proposed tool call, the user's request and the agent's history, then answers approve or deny. It takes up to 65,536 tokens of context. This version loads through one small Python file, autoquant.py, in plain PyTorch, without compiled kernels. It was checked on NVIDIA CUDA, the Apple M4 Max CPU and Apple MPS. For an even smaller file, see auto-200m-2-int4 (77…

Read SSH's full model card

This is auto-200m-2 with its weights stored as 8-bit integers. The file is 150 MB instead of 299 MB, and it gets the same benchmark score: 2,890/3,000, with 53 false approvals and 57 false denials. Only 2 of the 3,000 decisions differ from the BF16 model.

auto-200m-2 is a 149.6M-parameter ModernBERT classifier. It reads an AI agent's proposed tool call, the user's request and the agent's history, then answers approve or deny. It takes up to 65,536 tokens of context. This version loads through one small Python file, auto_quant.py, in plain PyTorch, without compiled kernels. It was checked on NVIDIA CUDA, the Apple M4 Max CPU and Apple MPS. For an even smaller file, see auto-200m-2-int4 (77 MB).

Results

These results use the pinned 3,000-item Approve-or-Deny benchmark (revision a38b6259), full input lengths, P(deny) >= 0.5, and the same evaluation code as the base model card. The quantized rows were computed on CUDA in BF16 from the stored integer weights. A false approval is an unsafe call that was approved. A false denial is an authorized call that was denied.

Model Weights file Accuracy False approvals False denials AUROC 16k–64k tokens Validation audit Decisions that differ from BF16
auto-200m-2 (BF16) 299 MB 96.33% (2890) 53/1401 57/1599 0.9937 94.14% 98.34% —
auto-200m-2-int8 150 MB 96.33% (2890) 53/1401 57/1599 0.9937 94.14% 98.34% 2
auto-200m-2-int4 77 MB 96.13% (2884) 60/1401 56/1599 0.9933 94.98% 98.15% 44

The 16k–64k column covers 239 benchmark items. The validation audit is 2,595 validation rows that were never used for training or selection. Paired with the BF16 model on the same items, accuracy differs by +0.00 points (95% interval -0.10 to +0.10; 1 items right only for BF16, 1 right only for int8; exact McNemar p = 1.00). On the tool probes it gets 24/24 published skills/MCP/custom-tool probes and 38/40 fresh scope/history/injection probes; the BF16 model gets 24/24 and 38/40. Breakdowns by category, language, difficulty and length are in eval_results.json, and per-item logits are in benchmark_predictions.npz.

Devices

Device Compute dtype Items Accuracy on those items Same items, CUDA Decisions that differ from CUDA Largest P(deny) difference
NVIDIA RTX PRO 6000 (CUDA) BF16 all 3,000 96.33% reference — —
NVIDIA RTX PRO 6000, public loader (both attention modes) BF16 300 under 2,048 tokens — — 0 0.018
Apple M4 Max CPU FP32 2,452 up to 2,048 tokens 96.82% 96.86% 1 0.090
Apple M4 Max GPU (MPS) FP32 all 3,000 (longest 56,176 tokens) 96.30% 96.33% 1 0.090
  • CUDA row: the benchmark run above.
  • Public loader row: 300 benchmark items were scored again with auto_quant.load on CUDA, once with memory-linear attention and once with PyTorch SDPA, and compared with the benchmark run.
  • Mac rows: auto_quant.load and score in FP32 on an M4 Max (macOS). The CPU pass used the 2,452 items up to 2,048 tokens. The MPS pass used all 3,000 items, up to 56,176 tokens.

The P(deny) differences come from BF16 on CUDA versus FP32 on the Mac. AMD ROCm should work, since the loader is plain PyTorch, but it wasn't tested. The details are in eval/mac_verify.json.

Memory and speed on the same Mac, one request at a time. All three use the same memory-linear attention:

Model Device Memory after load Median, under 1k tokens 7,449 tokens 29,233 tokens
auto-200m-2 (BF16 file, runs in FP32) CPU 806 MB 68 ms 1.65 s 15.7 s
int8 CPU 379 MB 70 ms 1.49 s 13.4 s
int4 CPU 305 MB 104 ms 1.51 s 12.5 s
auto-200m-2 (BF16 file, runs in FP32) MPS 570 MB 24 ms 0.46 s 4.4 s
int8 MPS 143 MB 30 ms 0.47 s 4.1 s
int4 MPS 74 MB 38 ms 0.47 s 4.3 s

Each model and device ran in its own process. Latency is the median over 110 short benchmark items. On CPU, memory is how much the process grew during loading. On MPS, it's the tensors held on the GPU. The BF16 checkpoint runs in FP32 on CPU and MPS, which is Transformers' default there.

On the GPU, the int8 and int4 weights take about 4× and 8× less memory than the FP32 copy. They don't make short requests faster, though: each layer turns its integer weights back into floats on every call, and unpacking 4-bit values costs extra, so int4 is the slowest on short inputs. On long inputs, activations dominate both time and memory; at 29,233 tokens the CPU peak was 2.1 / 1.8 / 1.7 GB (BF16 / int8 / int4). See eval/speed_mac.json.

How it was made

  • Format. Weights are stored as symmetric int8 in [-127, 127], with one FP16 scale per output channel (per row for the token embedding). The token embeddings, every attention and MLP linear layer, and head.dense are quantized. The LayerNorm weights and the final 768×2 classifier stay in FP16. Activations keep the device's float type, and each layer rebuilds its own float weights just before its matrix multiply. The details are in quant_config.json and the docstring of auto_quant.py.
  • Quantization-aware training (QAT), and why these weights aren't from it. QAT ran on the base model: fake quantization with straight-through rounding, learnable (LSQ) scales, loss 0.1·CE + 0.9·KL(base ‖ quantized) at T = 1, 150,000 training rows and 4 exports. A rule fixed before any benchmark run froze whichever candidate had the lowest loss (NLL) on the 7,824-row validation selection split. Plain post-training quantization (PTQ, MSE-clipped scales) was one of the candidates. PTQ won with an NLL of 0.07195, against 0.07194 for the BF16 model itself. Every QAT export came out slightly higher (0.0730–0.0733). So these weights are the PTQ export. At 8 bits, rounding alone loses almost nothing on this model. All candidates are in eval/qat_candidates.json.
  • One benchmark look. The frozen export was benchmarked once.
  • Base weights. This quantizes the original auto-200m-2 (revision 0bbb929f). In the same job, Auto 3B was distilled into auto-200m-2 as a candidate "iteration 1". It made fewer false approvals but more false denials, so it failed its gate and wasn't published, and both quants start from the original weights.

Usage

import os, sys
from huggingface_hub import hf_hub_download

repo = "ProCreations/auto-200m-2-int8"
sys.path.insert(0, os.path.dirname(hf_hub_download(repo, "auto_quant.py")))
import auto_quant

model, tokenizer = auto_quant.load(repo)   # picks CUDA/ROCm, then Apple MPS, then CPU
text = auto_quant.build_input(
    user_request="Clean up the build artifacts and reinstall dependencies.",
    history=[{"tool": "Bash", "args": "ls", "result": "node_modules dist package.json"}],
    call={"tool": "Bash", "args": "rm -rf node_modules dist && npm install"},
)
p_deny = auto_quant.score(model, tokenizer, text)
print("deny" if p_deny >= 0.5 else "approve", round(p_deny, 3))

You need torch, transformers 5.x, huggingface_hub and safetensors; nothing gets compiled. Labels are 0 = approve and 1 = deny. Keep the three section headers exactly as build_input writes them. Arguments are strings; for JSON arguments, compact JSON (json.dumps(args, separators=(",", ":"))) matches the training data.

score() takes one request at a time, up to 65,536 tokens, and its memory grows linearly with length. For padded batches, load with attention="sdpa" and call model(**inputs). AutoModelForSequenceClassification.from_pretrained can't read this format. The Auto runtime 0.2.0 runs the BF16 model and doesn't load these files yet.

Files

  • auto_quant_int8.safetensors: the quantized weights. quant_config.json describes the format.
  • config.json, tokenizer.json, tokenizer_config.json: from auto-200m-2.
  • auto_quant.py: the loader, about 280 lines of PyTorch.
  • eval_results.json and benchmark_predictions.npz: benchmark, audit and probe results, plus per-item logits.
  • eval/: the PTQ sweep, the QAT candidates and their validation scores, the probe results and the device checks.
  • training/: the QAT and evaluation code. It ran inside a temporary job, so paths refer to that job.

Limitations

This is a classifier, not a policy engine. It approves routine authorized work, and it denies consequential unauthorized actions and actions that follow injected instructions. It can't inspect hidden file contents, resolve opaque executables or know what a URL will do at runtime, so false approvals remain possible. The labels are synthetic, and the benchmark has been reused across Auto releases, so evaluate it on your own traffic before relying on it.

Configuration

Architecture
ModernBertForSequenceClassification
Context length (tokens)
65,536
Layers
22
Hidden size
768
Feed-forward size
1,152
Attention heads
12
Vocabulary size
50,368
Model type
modernbert

Identity and Version

Repository
ProCreations/auto-200m-2-int8
Publisher
SSH
Task
Text classification
Modality
Text
Library
pytorch
Parameters
Not stated by the source
Languages
qat
Revision
2501a22901e8cc520c746a86f3f9d04f7feaaefb
First published
2026-09-27
Last updated
2026-09-27

Files and Weights

34 files, 154.0 MB in total. The weights are 2 files totalling 150.1 MB in npz, safetensors.

Weights2 files · 150.1 MB
Configuration28 files · 267.6 KB
Tokenizer2 files · 3.6 MB
Documentation1 file · 10.2 KB
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
auto_quant_int8.safetensorsWeights150.0 MB 94d19fbe6389
benchmark_predictions.npzWeights72.8 KB fa14d5c87627
auto_quant.pyConfiguration14.3 KB —
config.jsonConfiguration2.0 KB —
eval/mac_verify.jsonConfiguration1.4 KB —
eval/probe_results.jsonConfiguration8.2 KB —
eval/ptq_sweep.jsonConfiguration1.2 KB —
eval/qat_candidates.jsonConfiguration3.1 KB —
eval/speed_mac.jsonConfiguration4.0 KB —
eval_results.jsonConfiguration89.9 KB —
quant_config.jsonConfiguration3.1 KB —
training/augment2.pyConfiguration5.1 KB —
training/auto3b_eval.pyConfiguration4.2 KB —
training/auto_quant.pyConfiguration14.3 KB —
training/chain.pyConfiguration1.1 KB —
training/choose.pyConfiguration4.4 KB —
training/common.pyConfiguration3.1 KB —
training/download.pyConfiguration2.2 KB —
training/engine.pyConfiguration18.1 KB —
training/evaluate.pyConfiguration11.9 KB —
training/label.pyConfiguration8.7 KB —
training/mac_verify.pyConfiguration4.4 KB —
training/plan.jsonConfiguration6.2 KB —
training/prepare.pyConfiguration10.1 KB —
training/qat.pyConfiguration21.4 KB —
training/run_tool_probes.pyConfiguration4.6 KB —
training/smoke.pyConfiguration3.4 KB —
training/speed_mac.pyConfiguration5.9 KB —
training/stage.pyConfiguration6.8 KB —
training/train.pyConfiguration4.6 KB —
README.mdDocumentation10.2 KB —
.gitattributesRepository1.5 KB —
tokenizer.jsonTokenizer3.6 MB —
tokenizer_config.jsonTokenizer310 B —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
150.1 MB
Download from SSH

Released by SSH through its official repository on Hugging Face. Read the license.

Built From

  • Derived from ProCreations/auto-200m-2
  • Quantized from ProCreations/auto-200m-2
  • Trained on (disclosed) ProCreations/auto-1b-data

Evaluations

Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.

BenchmarkConditionsResultReported byRevisionDate
Approve-or-Deny Task Agentic tool-call approve/deny classificationMetric F1 (deny)Comparison conditions not established 0.960798 ProCreations
Publisher reported
Evaluated revision not stated —
Approve-or-Deny Task Agentic tool-call approve/deny classificationMetric accuracyComparison conditions not established 0.963333 ProCreations
Publisher reported
Evaluated revision not stated —
Approve-or-Deny Task Agentic tool-call approve/deny classificationMetric false_approve_rateComparison conditions not established 0.0378301 ProCreations
Publisher reported
Evaluated revision not stated —
Approve-or-Deny Task Agentic tool-call approve/deny classificationMetric false_deny_rateComparison conditions not established 0.0356473 ProCreations
Publisher reported
Evaluated revision not stated —
Approve-or-Deny Task Agentic tool-call approve/deny classificationMetric roc_aucComparison conditions not established 0.993694 ProCreations
Publisher reported
Evaluated revision not stated —

Memory Requirements

PrecisionWeights in memory
As published150.1 MB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About auto-200m-2-int8

Can I use auto-200m-2-int8 commercially?

Yes. auto-200m-2-int8 is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is auto-200m-2-int8's context length?

65,536 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text classification

finbert

Prosus AI

FinBERT is a pre-trained NLP model to analyze sentiment of financial text. It is built by further training the BERT language model in the finance domain, using a large financial corpus and thereby fine-tuning it for financial sentiment classification. Financial PhraseBank by Malo et al. (2014) is used for fine-tuning. For more details, please see the paper FinBERT: Financial Sentiment Analysis with Pre-trained Language Models and our related blog post on Medium. The model will give softmax outputs for three labels: positive, negative or neutral. About Prosus Prosus is a global consumer internet group and one of the largest technology investors in the world. Operating and investing globally…

Open weights 512 tokens transformers

Model · Text classification

twitter-roberta-base-sentiment-latest

Cardiff NLP

This is a RoBERTa-base model trained on ~124M tweets from January 2018 to December 2021, and finetuned for sentiment analysis with the TweetEval benchmark. The original Twitter-based RoBERTa model can be found here and the original reference paper is TweetEval. This model is suitable for English. 0 -> Negative; 1 -> Neutral; 2 -> Positive This sentiment analysis model has been integrated into TweetNLP. You can access the demo here.

Open weights cc-by-4.0 514 tokens transformers

Model · Text classification

finbert-tone

Yi

FinBERT is a BERT model pre-trained on financial communication text. The purpose is to enhance financial NLP research and practice. It is trained on the following three financial communication corpus. The total corpora size is 4.9B tokens. More technical details on FinBERT: Click Link This released finbert-tone model is the FinBERT model fine-tuned on 10,000 manually annotated (positive, negative, neutral) sentences from analyst reports. This model achieves superior performance on financial tone analysis task. If you are simply interested in using FinBERT for financial tone analysis, give it a try. If you use the model in your academic work, please cite the following paper: Huang, Allen H.…

Open weights 512 tokens transformers

Model · Text classification

ms-marco-MiniLM-L-6-v2

Joshua

https://huggingface.co/cross-encoder/ms-marco-MiniLM-L-6-v2 with ONNX weights to be compatible with Transformers.js. If you haven't already, you can install the Transformers.js JavaScript library from NPM using: Note: Having a separate repo for ONNX weights is intended to be a temporary solution until WebML gains more traction. If you would like to make your models web-ready, we recommend converting to ONNX using Optimum and structuring your repo like this one (with ONNX weights located in a subfolder named onnx).

Open weights 512 tokens transformers.js

Model · Text classification

twitter-xlm-roberta-base-sentiment

Cardiff NLP

This is a multilingual XLM-roBERTa-base model trained on ~198M tweets and finetuned for sentiment analysis. The sentiment fine-tuning was done on 8 languages (Ar, En, Fr, De, Hi, It, Sp, Pt) but it can be used for more languages (see paper for details). This model has been integrated into the TweetNLP library.

Open weights 514 tokens transformers

Model · Text classification

emotion-english-distilroberta-base

Hartmann

With this model, you can classify emotions in English text data. The model was trained on 6 diverse datasets (see Appendix below) and predicts Ekman's 6 basic emotions, plus a neutral class: 1) anger 2) disgust 3) fear 4) joy 5) neutral 6) sadness 7) surprise The model is a fine-tuned checkpoint of DistilRoBERTa-base. For a 'non-distilled' emotion model, please refer to the model card of the RoBERTa-large version. a) Run emotion model with 3 lines of code on single text example using Hugging Face's pipeline command on Google Colab: b) Run emotion model on multiple examples and full datasets (e.g.,.csv files) on Google Colab: Please reach out to [email protected] if you have any…

Open weights 514 tokens transformers