IdeaLens-Qwen3.5-9B is an idea-level detector: it judges whose ideas a document contains, not who wrote its words, so a document whose ideas are a person's counts as human however much of its prose an AI wrote. It is one of the detectors released with IdeaLens and trained on the same data. The idealens package (PyPI) runs the whole pipeline: it assigns each document one of the eight formats, extracts the outline with the prompt, role vocabulary and worked examples the detectors were trained with, and scores it with this model and the thresholds in this repo. Input is JSONL with a text field per document. To score outlines you already have, use idealens score outlines.jsonl -o scores.jsonl…
Independent publisher
Rishanth Rajendhran
rishanthrajendhran
NLP, LLM, Prompting, Multimodality, RLHF, Local Explainability
Models
IdeaLens detects who came up with the ideas in a document, rather than who wrote its words. It reads a role-labelled outline of the document (an ordered list of items, each giving one idea and the discourse role it plays, such as Central Development or Open Question) and returns P(human), the probability that the ideas are human. IdeaLens is nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 fine-tuned with LoRA (rank 64) on the outlines of 1M English web documents (WildOutlines). The training outlines were paraphrased to remove the documents' wording, so the model has to fit its labels through the ideas. Try it in your browser: the IdeaLens & ProseLens demo scores your own text with both…
IdeaLens-NoParaphrase is an ablation of IdeaLens: the same backbone, documents and labels, trained on the outlines as extracted, without the paraphrasing step that removes the documents' wording. It reads a role-labelled outline and returns P(human), the probability that the document's ideas are human. It is nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 fine-tuned with LoRA (rank 64) on the outlines of 1M English web documents (WildOutlines, outline field). Try it in your browser: the IdeaLens & ProseLens demo scores your own text with both detectors, no installation or keys needed. This model is not reported in the paper. It is released for comparison with IdeaLens. Scoring a document…
IdeaLens-ModernBERT-L-NoParaphrase is an idea-level detector: it judges whose ideas a document contains, not who wrote its words, so a document whose ideas are a person's counts as human however much of its prose an AI wrote. It is one of the detectors released with IdeaLens and trained on the same data. The idealens package (PyPI) runs the whole pipeline: it assigns each document one of the eight formats, extracts the outline with the prompt, role vocabulary and worked examples the detectors were trained with, and scores it with this model and the thresholds in this repo. Input is JSONL with a text field per document. To score outlines you already have, use idealens score outlines.jsonl -o…
ProseLens detects who wrote the words of a document. It is IdeaLens's counterpart in the paper: the same backbone, training documents and labels, but it reads the raw document text instead of an outline, so it learns word-level provenance. It returns P(human), the probability that the document was written by a person. ProseLens is nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 fine-tuned with LoRA (rank 64) on 1M English web documents (WildOutlines). Try it in your browser: the IdeaLens & ProseLens demo scores your own text with both detectors, no installation or keys needed. From the paper: ProseLens is accurate when a document's ideas and words come from the same source (99.1%), but…
IdeaLens-Qwen3.5-9B-PerItem is an idea-level detector: it judges whose ideas a document contains, not who wrote its words, so a document whose ideas are a person's counts as human however much of its prose an AI wrote. It is one of the detectors released with IdeaLens and trained on the same data. The idealens package (PyPI) runs the whole pipeline: it assigns each document one of the eight formats, extracts the outline with the prompt, role vocabulary and worked examples the detectors were trained with, and scores it with this model and the thresholds in this repo. Input is JSONL with a text field per document. To score outlines you already have, use idealens score outlines.jsonl -o…
ProseLens-ModernBERT-L is a prose-level detector: it reads the document itself and judges who wrote the words. It is the ModernBERT counterpart of ProseLens and the comparison point for the idea-level detectors. Score documents directly; no outline is needed. The idealens package (PyPI) applies this model and the thresholds in this repo: Input is JSONL with a text field per document. thresholds.json holds this model's cuts at 0.1%, 0.5%, 1%, 2% and 5% false-positive rates, fitted on the 80,000 human documents of WildOutlines' calibration split: one global cut per rate, plus per-format cuts. The package applies them. A cut fitted for one model does not transfer to another model's scores. For…
IdeaLens-ModernBERT-L is an idea-level detector: it judges whose ideas a document contains, not who wrote its words, so a document whose ideas are a person's counts as human however much of its prose an AI wrote. It is one of the detectors released with IdeaLens and trained on the same data. The idealens package (PyPI) runs the whole pipeline: it assigns each document one of the eight formats, extracts the outline with the prompt, role vocabulary and worked examples the detectors were trained with, and scores it with this model and the thresholds in this repo. Input is JSONL with a text field per document. To score outlines you already have, use idealens score outlines.jsonl -o scores.jsonl…
IdeaLens-ModernBERT-L-RolesOnly is an idea-level detector: it judges whose ideas a document contains, not who wrote its words, so a document whose ideas are a person's counts as human however much of its prose an AI wrote. It is one of the detectors released with IdeaLens and trained on the same data. The idealens package (PyPI) runs the whole pipeline: it assigns each document one of the eight formats, extracts the outline with the prompt, role vocabulary and worked examples the detectors were trained with, and scores it with this model and the thresholds in this repo. Input is JSONL with a text field per document. To score outlines you already have, use idealens score outlines.jsonl -o…
IdeaLens-ModernBERT-L-PerItem is an idea-level detector: it judges whose ideas a document contains, not who wrote its words, so a document whose ideas are a person's counts as human however much of its prose an AI wrote. It is one of the detectors released with IdeaLens and trained on the same data. The idealens package (PyPI) runs the whole pipeline: it assigns each document one of the eight formats, extracts the outline with the prompt, role vocabulary and worked examples the detectors were trained with, and scores it with this model and the thresholds in this repo. Input is JSONL with a text field per document. To score outlines you already have, use idealens score outlines.jsonl -o…
IdeaLens-LogisticClassifier is an idea-level detector: it judges whose ideas a document contains, not who wrote its words, so a document whose ideas are a person's counts as human however much of its prose an AI wrote. It is one of the detectors released with IdeaLens and trained on the same data. The idealens package (PyPI) runs the whole pipeline: it assigns each document one of the eight formats, extracts the outline with the prompt, role vocabulary and worked examples the detectors were trained with, and scores it with this model and the thresholds in this repo. Input is JSONL with a text field per document. To score outlines you already have, use idealens score outlines.jsonl -o…
IdeaLens-LogisticClassifier-PerItem is an idea-level detector: it judges whose ideas a document contains, not who wrote its words, so a document whose ideas are a person's counts as human however much of its prose an AI wrote. It is one of the detectors released with IdeaLens and trained on the same data. The idealens package (PyPI) runs the whole pipeline: it assigns each document one of the eight formats, extracts the outline with the prompt, role vocabulary and worked examples the detectors were trained with, and scores it with this model and the thresholds in this repo. Input is JSONL with a text field per document. To score outlines you already have, use idealens score outlines.jsonl…