GroundCheck v2 (ModernBERT-base)
A ~150M-parameter encoder that checks whether a RAG answer is supported by the source it
was given. Given an answer and a source (and optionally the question), it returns
grounded or hallucinated with P(grounded). It runs on CPU.
- Labels:
0 = grounded, 1 = hallucinated, and P(grounded) = softmax(logits)[0].
- Code, evaluation harness and provenance: https://github.com/Pranshurs/groundcheck
- Every number below is traceable to a committed report in that repository.
Which weights
|
|
| Recommended revision |
998cec35563d6b90947409d1c7510adac7f7c80c (the groundcheck-rag package pins this) |
| Weights committed in |
0c7dd0636c0d58b5e3865db8f1deee5a5acfeb6c (later commits change only this card) |
model.safetensors SHA-256 |
9ec331ba6d8a9d93236dd72b239df518b07e61241323876f8aafc223d959461a |
| Base model |
answerdotai/ModernBERT-base (Apache-2.0) |
Hashes for every file are in the repository's MODEL_PROVENANCE.json.
python -m eval.provenance --verify re-checks them.
Results
All measured on CPU against the pinned weights, using test sets rebuilt from pinned public
datasets. F1 is for the hallucinated class, and the 95% CIs come from a 2,000-resample
bootstrap.
| Suite |
n |
max_length 512 (training protocol) |
max_length 2048 (groundcheck-rag default) |
| RAGTruth test (first 2,500 of 2,700 responses) |
2,500 |
F1 0.682 [0.658, 0.705], acc 0.746 |
F1 0.696 [0.672, 0.718], acc 0.758 |
| VitaminC test |
2,000 |
acc 0.850, F1 0.845 |
identical (all inputs are short) |
| One-fact flips caught (regenerated holdout) |
500 |
78.0% |
87.2% |
| Same answers unflipped, kept grounded |
500 |
76.0% |
74.0% |
- 512 tokens reproduces the training run's published numbers (F1 0.6824, acc 0.7468;
VitaminC acc 0.8495) to within one prediction in 2,500.
- On the same rows, 2048 tokens scores a paired ΔF1 of +0.014 (95% CI −0.001 to +0.028).
That's a small gain, at about 2.5× the CPU time on long documents.
- The flipped-fact holdout is regenerated. The original run's sample depended on
Python's per-process hash seed and can't be rebuilt. The original sample reported 80.4%
caught and 76.2% kept.
- RAGTruth uses the first 2,500 of the 2,700 test responses, the same subset the
training run evaluated on.
External published reference (not a controlled comparison)
The RAGTruth paper (Niu et al., 2024; tabulated in LettuceDetect, 2025) reports F1 0.634 for
a zero-shot GPT-4-turbo prompt judge. That figure comes from a different protocol: a
prompted judge scored on all 2,700 test responses. It was not run head-to-head with
GroundCheck, so it's there for orientation and doesn't support a "beats GPT-4" claim.
Latency (one machine, not a guarantee)
Apple M1, CPU, 4 torch threads, single requests after warm-up, at max_length 2048:
| Pair length |
p50 |
| ≤ 128 tokens |
39 ms |
| 129–512 tokens |
172 ms |
| 513–2,048 tokens |
349 ms |
| > 2,048 tokens |
~1.45 s |
Other hardware will differ. Measure with python -m bench.latency.
Usage
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
name, rev = "Pranshurs/groundcheck-modernbert", "998cec35563d6b90947409d1c7510adac7f7c80c"
tok = AutoTokenizer.from_pretrained(name, revision=rev)
model = AutoModelForSequenceClassification.from_pretrained(name, revision=rev).eval()
source = "France's capital and largest city is Paris."
answer = "Paris is the capital of France."
enc = tok(source, answer, truncation="only_first", max_length=512, return_tensors="pt")
with torch.no_grad():
grounded = torch.softmax(model(**enc).logits, dim=-1)[0, 0].item()
print("grounded" if grounded >= 0.5 else "hallucinated", round(grounded, 3))
Or use the library (pip install "groundcheck-rag[model]"), which pins the revision,
truncates only the source, and raises an error rather than substituting a heuristic when the
model can't load:
from groundcheck import GroundCheck
print(GroundCheck().check(source="...", answer="..."))
Training
Fine-tuned from answerdotai/ModernBERT-base (pinned at 8949b909) as a sequence-pair
classifier: the premise is the optional question plus the source, and the hypothesis is the
answer. The training run was on a single Kaggle P100 with torch 2.4.1 and transformers 4.49.0:
3 epochs, batch 16, learning rate 2e-5, linear schedule, warmup 0.06, weight decay 0.01,
seed 42, fp16, sequence length 512.
The 28,500 training rows break down as:
- 10,000 from RAGTruth (wandb/RAGTruth-processed @ eb4f4b9d), with the question dropped
on ~50% of rows.
- 16,000 from VitaminC (tals/vitaminc @ be6febb7): SUPPORTS → grounded; REFUTES and NOT
ENOUGH INFO → hallucinated.
- 2,500 rule-based one-fact flips (a number, date, direction word or entity) of grounded
RAGTruth answers.
There is no LLM-generated augmentation and no private data. The recipe and data builders
are in the repository's training/ directory.
Intended use
Use it as a post-generation check in RAG pipelines: flag answers the retrieved source
doesn't support, so they can be routed to review, regeneration or a stronger checker. It
checks support against the provided source only and isn't a world-knowledge fact-checker.
Limitations
- English only. Verdicts are per answer, not per span.
- Long sources are truncated from the end. 76% of RAGTruth test pairs exceed 512 tokens and
19 of 2,500 exceed 2,048, so chunk long documents.
- RAGTruth precision is about 0.63, so roughly a third of
hallucinated verdicts on long RAG
answers are false alarms. Scores aren't calibrated probabilities; tune the threshold on
your own data.
- About a quarter of unedited grounded answers in the minimal-edit holdout are flagged.
- It hasn't been evaluated on adversarial or out-of-domain inputs.
License and data terms
| What |
Terms |
| These model weights |
MIT |
Base model, answerdotai/ModernBERT-base |
Apache-2.0 |
| The GroundCheck code |
Apache-2.0 |
| Training data |
Keeps its upstream terms, which the MIT license on the weights does not change |
The upstream data terms (detailed in the repository's DATA_LICENSES.md):
- RAGTruth is MIT. Its source passages come
from:
- MS MARCO: Microsoft's terms allow non-commercial research use only.
- The Yelp Open Dataset: academic and non-commercial use only.
- CNN/DailyMail: the articles are copyrighted by their publishers.
- VitaminC is CC BY-SA 3.0.
Whether a dataset's non-commercial terms extend to a model trained on it is legally
unsettled. Review the data terms before any commercial use. This card is not legal advice.