Contrastive activation extractions for 17 semantic concepts across 46 language models, supporting cross-architecture mechanistic interpretability research. - papern250/ — N=250, C=17. The 40 standard models are complete:.npy + all 7 JSON analysis families (caz, gem, ablation, ablationgem, ablationglobalsweep, ablationrandom, patch) for every concept (ablationrandom ≤17 by design — see tree note). Use for paper reproducibility. 6 large models are caz-only (.npy + caz only), hardware-blocked — see below. - rcpv1/ — raw N=2000 activations (.npy + meta.json) for 40 models. Derived analysis (caz/gem/ablation/globalsweep/random) at N=2000 is not yet computed — it exists only at N=250 in…
Independent publisher
James R. A. Henry
james-ra-henry
Mechanistic Interpretability
Models in Library0
Datasets in Library1
Models on Hugging Face—
Followers—