SAVRN
Search Contact SAVRN

tasksource-jev-nano-v0 · Model Card

tasksource-jev-nano-v0: Model Card

Written by Tasksource, published under apache-2.0, revision ad12951788ca, read 2026-10-04. Shown as written; SAVRN's own facts about this model are on its page.

# Tasksource-JEV-Nano (`tasksource-jev-nano-v0`) **High-Throughput Native Decision Model (<200M Parameters)** *Decoupled Multi-Vector Late Interaction for Fast, Calibrated System-One Decisions* [![Decision Index](https://img.shields.io/badge/Decision%20Index%20Raw-1.21-blue)](https://github.com/sileod/decision-models) [![JevBench Composite](https://img.shields.io/badge/JevBench%20Composite-63.14%2F100-purple)](https://github.com/sileod/decision-models) [![Classifier Benchmark v2](https://img.shields.io/badge/Classifier--Benchmark%20v2-55.78%25-green)](https://github.com/sileod/decision-models) [![License](https://img.shields.io/badge/License-Apache%202.0-green.svg)](https://opensource.org/licenses/Apache-2.0) [![Parameters](https://img.shields.io/badge/Parameters-149M-orange.svg)](https://huggingface.co/answerdotai/ModernBERT-base) [![Context Length](https://img.shields.io/badge/Context-8192-purple.svg)](https://huggingface.co/answerdotai/ModernBERT-base)

What is Tasksource-JEV-Nano?

Tasksource-JEV-Nano is a compact (~149M parameter) decision model that picks the optimal action, category, or verdict from a candidate list using token-level multi-vector late interaction rather than conventional classification heads.

Built on answerdotai/ModernBERT-base and lightonai/LateOn, it handles all three fundamental System-One decision types: 1. Choice: Selecting the best candidate among variable numbers of options ($K=2$ to $K=100+$). 2. NOUL: Yes/no uncertainty decisions. 3. Score: Quantitative ranking and graded scales.


Key Architectural Strengths

  • Reusable State Caching: The situation context (state + question) is encoded once into token vectors. Evaluating 1, 10, or 100 candidate options reuses this cached state without re-encoding the context, delivering up to 5.45x speedup.
  • Exact Candidate Permutation Equivariance: Unlike encoder-concatenation classifiers where candidate order biases logits, late-interaction candidate scoring is mathematically independent across candidates. Permuting the option list guarantees an identically permuted output (max diff: 0.00e+00).
  • Certified Zero Contamination: Certified zero contamination against all 21 standard evaluation benchmarks. Every evaluation is strictly zero-shot out-of-distribution.

Benchmark Evaluations

All evaluations below are measured on official benchmark suites under zero-shot out-of-distribution conditions:

Community Benchmarks

Benchmark Suite Metric Score Key Takeaway
Decision Index (Official 0.2.1) Balanced Raw Index 1.21 Evaluated on official decision-index suite (2,000 cases across 44 benchmarks)
JevBench (v1.3 Composite) Geometric Composite Score 63.14 / 100 High overall balance across intelligence, calibration, speed
Calibration (ECE) 0.2039 (Score: 59.22) Low confidence calibration error
Inference Latency (p50 / p95) 37.6ms / 55.4ms Sub-40ms execution on standard GPUs
Classifier-Benchmark v2 Macro Accuracy (49 tasks) 55.78% Evaluated on jabr/classifier-benchmark across 49 real production tasks
In-Domain Cardinality Banking77 (K=77) 50.9% Zero-shot discrimination across 77 fine-grained intents
Fast Decisions Exact Match 42.59% High-precision candidate ranking across 17 diverse domains

Authentic Out-of-Distribution Transfer Accuracy (Official Decision Index Suite)

Task / Domain Area Metric / Score Random Chance
Humicroedit Arts & Humor 64.1% 50.0%
cfcolor Color Perception 64.1% 50.0%
CLadder Causal Reasoning 64.1% 50.0%
BPoMP Arts & Poetry 62.7% 50.0%
WinoGrande Commonsense Coreference 61.5% 50.0%
RouterBench Model Routing Quality 0.5897 0.5000
PhishNChips Phishing Detection 56.4% 50.0%
ARC-Easy Science Knowledge 53.8% 25.0%
CLINC150+OOS Intent Classification 53.2% 0.6%
HoVer Fact Verification 52.6% 50.0%
BANKING77 Banking Intent Classification 48.2% 1.3%
RAGTruth Hallucination Detection (F1) 47.8% 25.0%
CRUXEval Code Execution Reasoning 46.2% 25.0%
ARC-Challenge Hard Science Reasoning 43.6% 25.0%
FinEntity Financial Entity Extraction 42.8% 20.0%
MMLU General Knowledge 35.9% 25.0%
When2Call Tool Invocation Decision 34.2% 25.0%
ANLI Adversarial NLI 30.9% 33.3%
GPQA Diamond PhD-Level Science Reasoning 30.8% 25.0%
ForecastBench Future Prediction (Brier) 0.2552 0.3333

Quickstart

Method 1: Using pylate (Recommended)

pip install -U pylate torch
import torch
from pylate import models

# Load model
model = models.ColBERT("tasksource/tasksource-jev-nano-v0")

state = "Customer reports unauthorized international wire transfer of $4,500 from their checking account."
question = "Select the appropriate fraud mitigation routing:"
options = [
    "approve and monitor silently",
    "challenge with an in-app biometric notification",
    "freeze account and trigger automated customer callback",
    "decline transfer and file immediate SAR report",
]

# 1. Encode context ONCE (cached state)
context = [f"{state}\nQuestion: {question}"]
ctx_vecs = model.encode(context, is_query=False, convert_to_tensor=True)[0]

# 2. Encode candidate options independently
opt_vecs = model.encode(options, is_query=True, convert_to_tensor=True)

# 3. ColBERT MaxSim late interaction: sum_l max_j <opt_l, ctx_j>
raw_scores = torch.stack([(opt @ ctx_vecs.T).max(dim=1).values.sum() for opt in opt_vecs])

# 4. Optional Soft Length Normalization (gamma = 0.04)
opt_lens = torch.tensor([max(1, opt.shape[0]) for opt in opt_vecs], dtype=torch.float32)
norm_scores = raw_scores / (opt_lens ** 0.04)

# 5. Calibrated choice probabilities
temperature = 0.3236
probs = torch.softmax(norm_scores / temperature, dim=0)

for opt, p in zip(options, probs):
    print(f"  [{p*100:5.1f}%]  {opt}")

Method 2: Pure Hugging Face transformers (Zero Extra Dependencies)

pip install -U transformers torch
import torch
import torch.nn.functional as F
from transformers import AutoModel, AutoTokenizer

model_id = "tasksource/tasksource-jev-nano-v0"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id)
model.eval()

context = "Alert: database replica lag exceeded 450 seconds on US-East-1.\nQuestion: Select incident severity:"
candidates = ["P4 - Low / Informational", "P3 - Moderate", "P2 - Major Degradation", "P1 - Critical Outage"]

with torch.no_grad():
    # Encode context
    tok_ctx = tokenizer([context], padding=True, return_tensors="pt")
    h_ctx = F.normalize(model(**tok_ctx).last_hidden_state[0], p=2, dim=-1)

    # Encode candidates
    tok_opts = tokenizer(candidates, padding=True, return_tensors="pt")
    h_opts = F.normalize(model(**tok_opts).last_hidden_state, p=2, dim=-1)

    # MaxSim late-interaction scoring
    scores = torch.stack([(h_opt @ h_ctx.T).max(dim=1).values.sum() for h_opt in h_opts])
    probs = F.softmax(scores / 0.3236, dim=0)

for opt, p in zip(candidates, probs):
    print(f"  {opt:<30} -> {p.item()*100:5.2f}%")

Benchmark Firewall & Zero-Contamination Guarantees

Tasksource-JEV-Nano was trained with a strict, automated benchmark firewall: - No Evaluation Overlap: All 21 evaluation benchmarks (banking77, clinc_oos/plus, anli, arc, winogrande, hellaswag, nli4ct, hover, etc.) were blacklisted and filtered from the candidate training stream with 0 overlap. - Unseen Task Generalization: Checkpoint selection was performed exclusively on unseen held-out Tasksource decision tasks.


Citation

@misc{sileo2026jevnanov0,
  title={Tasksource-JEV-Nano: Decoupled Multi-Vector Late Interaction for Typed Decisions},
  author={Sileo, Damien},
  year={2026},
  howpublished={\url{https://huggingface.co/tasksource/tasksource-jev-nano-v0}},
}