tasksource-jev-nano-v0 · Model Card
tasksource-jev-nano-v0: Model Card
Written by Tasksource, published under apache-2.0, revision ad12951788ca, read 2026-10-04. Shown as written; SAVRN's own facts about this model are on its page.
What is Tasksource-JEV-Nano?
Tasksource-JEV-Nano is a compact (~149M parameter) decision model that picks the optimal action, category, or verdict from a candidate list using token-level multi-vector late interaction rather than conventional classification heads.
Built on answerdotai/ModernBERT-base and lightonai/LateOn, it handles all three fundamental System-One decision types:
1. Choice: Selecting the best candidate among variable numbers of options ($K=2$ to $K=100+$).
2. NOUL: Yes/no uncertainty decisions.
3. Score: Quantitative ranking and graded scales.
Key Architectural Strengths
- Reusable State Caching: The situation context (state + question) is encoded once into token vectors. Evaluating 1, 10, or 100 candidate options reuses this cached state without re-encoding the context, delivering up to 5.45x speedup.
- Exact Candidate Permutation Equivariance: Unlike encoder-concatenation classifiers where candidate order biases logits, late-interaction candidate scoring is mathematically independent across candidates. Permuting the option list guarantees an identically permuted output (max diff: 0.00e+00).
- Certified Zero Contamination: Certified zero contamination against all 21 standard evaluation benchmarks. Every evaluation is strictly zero-shot out-of-distribution.
Benchmark Evaluations
All evaluations below are measured on official benchmark suites under zero-shot out-of-distribution conditions:
Community Benchmarks
| Benchmark Suite | Metric | Score | Key Takeaway |
|---|---|---|---|
| Decision Index (Official 0.2.1) | Balanced Raw Index | 1.21 | Evaluated on official decision-index suite (2,000 cases across 44 benchmarks) |
| JevBench (v1.3 Composite) | Geometric Composite Score | 63.14 / 100 | High overall balance across intelligence, calibration, speed |
| Calibration (ECE) | 0.2039 (Score: 59.22) | Low confidence calibration error | |
| Inference Latency (p50 / p95) | 37.6ms / 55.4ms | Sub-40ms execution on standard GPUs | |
| Classifier-Benchmark v2 | Macro Accuracy (49 tasks) | 55.78% | Evaluated on jabr/classifier-benchmark across 49 real production tasks |
| In-Domain Cardinality | Banking77 (K=77) | 50.9% | Zero-shot discrimination across 77 fine-grained intents |
| Fast Decisions Exact Match | 42.59% | High-precision candidate ranking across 17 diverse domains |
Authentic Out-of-Distribution Transfer Accuracy (Official Decision Index Suite)
| Task / Domain | Area | Metric / Score | Random Chance |
|---|---|---|---|
| Humicroedit | Arts & Humor | 64.1% | 50.0% |
| cfcolor | Color Perception | 64.1% | 50.0% |
| CLadder | Causal Reasoning | 64.1% | 50.0% |
| BPoMP | Arts & Poetry | 62.7% | 50.0% |
| WinoGrande | Commonsense Coreference | 61.5% | 50.0% |
| RouterBench | Model Routing Quality | 0.5897 | 0.5000 |
| PhishNChips | Phishing Detection | 56.4% | 50.0% |
| ARC-Easy | Science Knowledge | 53.8% | 25.0% |
| CLINC150+OOS | Intent Classification | 53.2% | 0.6% |
| HoVer | Fact Verification | 52.6% | 50.0% |
| BANKING77 | Banking Intent Classification | 48.2% | 1.3% |
| RAGTruth | Hallucination Detection (F1) | 47.8% | 25.0% |
| CRUXEval | Code Execution Reasoning | 46.2% | 25.0% |
| ARC-Challenge | Hard Science Reasoning | 43.6% | 25.0% |
| FinEntity | Financial Entity Extraction | 42.8% | 20.0% |
| MMLU | General Knowledge | 35.9% | 25.0% |
| When2Call | Tool Invocation Decision | 34.2% | 25.0% |
| ANLI | Adversarial NLI | 30.9% | 33.3% |
| GPQA Diamond | PhD-Level Science Reasoning | 30.8% | 25.0% |
| ForecastBench | Future Prediction (Brier) | 0.2552 | 0.3333 |
Quickstart
Method 1: Using pylate (Recommended)
pip install -U pylate torch
import torch
from pylate import models
# Load model
model = models.ColBERT("tasksource/tasksource-jev-nano-v0")
state = "Customer reports unauthorized international wire transfer of $4,500 from their checking account."
question = "Select the appropriate fraud mitigation routing:"
options = [
"approve and monitor silently",
"challenge with an in-app biometric notification",
"freeze account and trigger automated customer callback",
"decline transfer and file immediate SAR report",
]
# 1. Encode context ONCE (cached state)
context = [f"{state}\nQuestion: {question}"]
ctx_vecs = model.encode(context, is_query=False, convert_to_tensor=True)[0]
# 2. Encode candidate options independently
opt_vecs = model.encode(options, is_query=True, convert_to_tensor=True)
# 3. ColBERT MaxSim late interaction: sum_l max_j <opt_l, ctx_j>
raw_scores = torch.stack([(opt @ ctx_vecs.T).max(dim=1).values.sum() for opt in opt_vecs])
# 4. Optional Soft Length Normalization (gamma = 0.04)
opt_lens = torch.tensor([max(1, opt.shape[0]) for opt in opt_vecs], dtype=torch.float32)
norm_scores = raw_scores / (opt_lens ** 0.04)
# 5. Calibrated choice probabilities
temperature = 0.3236
probs = torch.softmax(norm_scores / temperature, dim=0)
for opt, p in zip(options, probs):
print(f" [{p*100:5.1f}%] {opt}")
Method 2: Pure Hugging Face transformers (Zero Extra Dependencies)
pip install -U transformers torch
import torch
import torch.nn.functional as F
from transformers import AutoModel, AutoTokenizer
model_id = "tasksource/tasksource-jev-nano-v0"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id)
model.eval()
context = "Alert: database replica lag exceeded 450 seconds on US-East-1.\nQuestion: Select incident severity:"
candidates = ["P4 - Low / Informational", "P3 - Moderate", "P2 - Major Degradation", "P1 - Critical Outage"]
with torch.no_grad():
# Encode context
tok_ctx = tokenizer([context], padding=True, return_tensors="pt")
h_ctx = F.normalize(model(**tok_ctx).last_hidden_state[0], p=2, dim=-1)
# Encode candidates
tok_opts = tokenizer(candidates, padding=True, return_tensors="pt")
h_opts = F.normalize(model(**tok_opts).last_hidden_state, p=2, dim=-1)
# MaxSim late-interaction scoring
scores = torch.stack([(h_opt @ h_ctx.T).max(dim=1).values.sum() for h_opt in h_opts])
probs = F.softmax(scores / 0.3236, dim=0)
for opt, p in zip(candidates, probs):
print(f" {opt:<30} -> {p.item()*100:5.2f}%")
Benchmark Firewall & Zero-Contamination Guarantees
Tasksource-JEV-Nano was trained with a strict, automated benchmark firewall:
- No Evaluation Overlap: All 21 evaluation benchmarks (banking77, clinc_oos/plus, anli, arc, winogrande, hellaswag, nli4ct, hover, etc.) were blacklisted and filtered from the candidate training stream with 0 overlap.
- Unseen Task Generalization: Checkpoint selection was performed exclusively on unseen held-out Tasksource decision tasks.
Citation
@misc{sileo2026jevnanov0,
title={Tasksource-JEV-Nano: Decoupled Multi-Vector Late Interaction for Typed Decisions},
author={Sileo, Damien},
year={2026},
howpublished={\url{https://huggingface.co/tasksource/tasksource-jev-nano-v0}},
}