credential-line-32m (v2)
A 32M-parameter classifier that reads one line of code, config, log or prose and scores whether
the line holds a literal credential: a password, API key, access token, private-key header,
connection string with a password, or an Authorization header value. It runs on a laptop CPU in
about 4 ms per line, so you can score every line of a trace, log or diff before you share it.
Use it next to a rule-based scanner, not in place of one. On 1,264 human-written lines it finds
about 4 in 5 credentials and is right about 9 times in 10 when it flags a line. It also misses
about 1 credential in 5, so it must never be the only check before you publish a file.
Run it now
pip install onnxruntime tokenizers numpy
python predict_onnx.py path/to/file.env # prints lines at or above the threshold
python predict_onnx.py path/to/file.env --all # prints every line with its score
predict_onnx.py applies the temperature (2.10) and threshold (0.74) in policy.json. Both were
fitted on human-written validation lines only. If you load the weights with transformers, apply
the same two numbers yourself; the raw softmax at 0.5 is not the calibrated decision.
import json, torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer
repo = "zaindanaharper/credential-line-32m"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForSequenceClassification.from_pretrained(repo).eval()
policy = {"temperature": 2.0954, "threshold": 0.74}
enc = tok(['stripe.api_key = "sk_live_51Hx9QeLkdT3pZr8XyWvA0bc"'], return_tensors="pt", truncation=True, max_length=128)
p = torch.softmax(model(**enc).logits / policy["temperature"], -1)[0, 1].item()
print(p, p >= policy["threshold"])
What it is good at
Human-written holdout: 1,264 lines, 472 credentials, 159 source groups. Lines from 75
open-source repositories in the CredData benchmark (reviewed real code, with credential values
scrambled) and from the Nosey Parker rule examples. No line, repository or rule file in this set
was used for training or tuning. Intervals are 95% bootstrap intervals that resample whole
repositories or rule files.
| System |
Precision |
Recall |
F1 |
| credential-line-32m v2 |
0.916 [0.866, 0.958] |
0.807 [0.723, 0.884] |
0.858 [0.802, 0.907] |
| credential-line-32m v1 |
0.848 [0.782, 0.908] |
0.731 [0.650, 0.807] |
0.785 [0.728, 0.842] |
| Char n-gram logistic regression, same training data |
0.761 [0.686, 0.828] |
0.761 [0.679, 0.836] |
0.761 [0.704, 0.812] |
| Zero-shot Qwen2.5-Coder-14B |
0.593 [0.542, 0.648] |
0.998 [0.993, 1.000] |
0.744 [0.703, 0.786] |
| Rule catalog (Flywheel trace redaction, credential rules) |
0.656 [0.595, 0.723] |
0.593 [0.539, 0.656] |
0.623 [0.576, 0.678] |
| detect-secrets 1.5.0, default plugins |
0.704 [0.647, 0.756] |
0.398 [0.342, 0.451] |
0.509 [0.453, 0.554] |
| Rule catalog OR this model |
0.707 [0.653, 0.765] |
0.888 [0.852, 0.924] |
0.787 [0.745, 0.832] |
| Rule catalog AND this model |
0.968 [0.942, 0.992] |
0.513 [0.439, 0.592] |
0.670 [0.603, 0.738] |
- Recall is higher than the rule catalog's with intervals that do not overlap, at higher precision.
Run it beside a rule scanner and flag a line if either one fires: that finds 89% of the
credentials in this set, against 59% for the catalog alone.
- Against the n-gram regression, the F1 intervals overlap slightly (0.802 against 0.812). A paired
bootstrap on the same lines puts the model ahead by 0.098 F1 [0.050, 0.144] and 0.155 precision
[0.108, 0.206].
- The zero-shot 14B model finds nearly every credential but flags 323 clean lines; this model
flags 35, at about one fifteenth of the per-line latency.
- Fixes since v1, measured on the holdout: keys standing alone on a line or in a list went from 10
of 47 found to 43 of 47; private-key header lines from 0 of 5 to 5 of 5; clean lines holding a
UUID flagged dropped from 17 to 2.
Synthetic regression set (6,207 lines, from v1). F1 0.994 [0.991, 0.997]: 372 of 379
credentials in formats never generated for training, 775 of 782 in trained formats, 0 of 409 hard
negatives flagged, 0 of 4,637 natural code lines flagged. This set is synthetic and was already
inspected during v1, so read it as a regression check, not a measure of real-world accuracy.
Speed. ONNX fp32 on an Intel desktop CPU, 4 threads, tokenization included, on the holdout
lines: 4.1 ms median and 7.5 ms 95th percentile per line at batch size 1 (500 lines); 219 lines per
second at batch size 32. ONNX matches PyTorch within 2.5e-6 on all 1,264 holdout lines, with
identical decisions at the threshold. The rule catalog takes 0.04 ms per line.
Calibration. Expected calibration error 0.037 and Brier score 0.072 on the holdout, after
temperature scaling fitted on validation lines. v1 scored 0.140 and 0.142 on the same lines.
Where it fails
- It misses about 1 credential in 5 on human-written lines (91 of 472). Misses cluster on
short or word-like passwords (
"thisismypassword", super$ecret, devpass), passwords inside
long connection strings, and keys whose only clue is the next line.
- It cannot reliably clear a rule scanner's hits. Among lines the rule catalog flagged, the
model's "clean" calls were right 139 of 177 times, 0.785 [0.637, 0.938]. Keep every hit a rule
scanner makes; do not use this model to suppress them.
- Password hashes and random-looking non-secrets draw false alarms. Of 35 false positives,
several are crypt-style hashes (
$6$..., $2y$...), idempotency tokens, and random-looking host
names.
- UUIDs remain hard. It found 10 of 16 credentials that are UUIDs (for example a UUID used as
an API key) on the holdout.
- One line at a time. A context-window variant that also reads the neighboring lines scored
F1 0.869 [0.817, 0.914] on the holdout, inside this model's interval. It did not clear the
validation margin set before training, so it is not the released model.
- It does not check whether a credential is live, and a high score is not proof of a leak.
- Mostly English identifiers; strongest on Python, JavaScript, Go, YAML, JSON, shell and env files.
Intended use
- A second pass beside a rule scanner when you redact traces, logs, transcripts or diffs before
sharing them. Flag a line when either system fires.
- Triage: ranking lines for a human reviewer.
Out of scope
- The only check before publishing or committing a file.
- Suppressing or overriding a rule scanner's findings.
- Access control, authentication or any decision that grants permission.
- Finding credentials in systems or data you are not authorized to handle.
- Tuning text so that it passes a secret scanner.
How the labels were made
- Holdout labels. Every holdout line got two blind labels, from Claude Haiku and from
Qwen3-8B, under a written rule (
training_code/v2/LABELING-v2.md). The file was hashed before
labeling and again before any system was scored. Raw agreement was 978 of 1,264, Cohen's kappa
0.55. All 286 disagreements, plus 23 lines where both labelers disagreed with the CredData
reviewers' own label, were adjudicated against the rule with a written reason per line.
- Agreement with human reviewers. On the 514 holdout lines CredData's reviewers had marked,
the final labels agree with the reviewers' labels on 493 (kappa 0.91), after mapping their
categories to this rule. The 21 differences are mostly UUIDs, salts and nonces, which the
CredData rules count as credentials and this rule does not.
- Corrections after scoring. A rule-based scan for filler values (runs such as
ABCDEFG,
0000, EXAMPLE) found 10 agreed labels that break the written rule; they were changed from
credential to clean (eval/post_scoring_label_corrections.tsv). The scan reads no model score.
The tables above use the corrected labels. On the labels frozen before scoring, the model scored
F1 0.855 [0.801, 0.905]; every system moved by 0.011 or less (eval/holdout_results_frozen_labels.json).
- Shared blind spots. One labeler and the independent label auditor are both Claude models
and may share blind spots. Apart from the CredData reviewer labels, no human labeled these lines.
Training data and provenance
87,707 training lines, 18,641 credentials. No real credential appears in the training data:
every value CredData marks as a credential was replaced by a random string of the same shape, and
every long random-looking token in every human-written line was replaced the same way.
- CredData (Samsung, Apache-2.0 dataset): 17,302 lines from 178 repositories, using only
repositories that GitHub reports as MIT, Apache-2.0, BSD, ISC, CC0, UPL or similar
(
eval/creddata_source_licenses.json; 306 of 337 repositories qualified). Reviewer labels were
mapped to this rule; 1,082 UUID lines and 58 salt or nonce lines the reviewers marked true were
relabeled clean, and 67 UUIDs in secret fields kept as credentials.
- Scanner test fixtures: 3,171 lines from secretlint (MIT), detect-secrets (Apache-2.0),
whispers (Apache-2.0) and CredSweeper (MIT). Lines with no credential signal are labeled clean;
the rest carry one label from Qwen3-8B under the written rule.
- v1 data: 60,653 lines: natural lines from permissively licensed Python packages and
synthetic credentials labeled by construction (see v1 card,
training_code/v1).
- New synthetic lines (v2): 6,581 lines labeled by construction: keys standing alone, in lists
and table cells; UUIDs in identifier fields such as
credentialsId; UUIDs as API key values;
private-key headers; public keys and END lines; CSRF tokens, nonces and salts; filler values.
- Validation (calibration and model choice): 4,564 human-written lines from 53 other CredData
repositories and the gitleaks rule examples (MIT), with four label corrections from an
independent audit.
Base model: jhu-clsp/ettin-encoder-32m, MIT,
revision 1b8ba06455dd44f80fc9c1ca9e22806157a57379. Fine-tuned for 2 epochs (2,742 steps, batch
64, learning rate 5e-5, max length 128 tokens, seed 1729) in 57 seconds on one RTX 4090. A control
run with shuffled training labels scored F1 0.565 on the holdout, the score of flagging nearly
every line, so the evaluation does not reward a model that learned nothing.
License
MIT for the fine-tuned weights, predict_onnx.py and the training code (LICENSE). The base
encoder is MIT-licensed by the Johns Hopkins University Center for Language and Speech
Processing; its notice is in NOTICE.
Reproduce
training_code/v2 holds the data builders, labeling, training, calibration and evaluation
scripts; training_code/v1 holds the v1 pipeline they extend. Paths are relative to a working
directory that holds ext/ (the cloned sources at the commits in eval/data_hashes.txt and the
scripts), v1/ and v2/. The rule-catalog baseline and the holdout lines are not redistributed,
so a rebuild reproduces the method rather than the exact bytes.
Files
| File |
What it is |
model.safetensors, config.json |
Fine-tuned weights, 32,031,746 parameters |
tokenizer.json, tokenizer_config.json |
Tokenizer from the base model |
onnx/model.onnx |
fp32 ONNX export |
policy.json |
Temperature and threshold, fitted on human-written validation lines |
predict_onnx.py |
Command-line scorer, ONNX Runtime only |
eval/holdout_results.json |
Every holdout number in this card, per slice, with intervals |
eval/label_agreement.json |
Labeler agreement and kappa |
eval/onnx_parity_latency.json |
Parity and timing |
LICENSE, NOTICE |
MIT license and the base model's notice |