SAVRN
Search Contact SAVRN

Research paper · 2026-10-05

IdeaLens: Detecting AI Ideas in Long-form Writing

Rishanth Rajendhran, Minjoon Choi, Jenna Russell, Ramya Namuduri, Deniz Bölöni-Turgut, Marzena Karpinska, John Wieting, Mohit Iyyer

12 open models in the SAVRN Model Hub cite IdeaLens: Detecting AI Ideas in Long-form Writing (2026). Together they draw 224 downloads a month. The most downloaded is IdeaLens-Qwen3.5-9B by Rishanth Rajendhran (text classification, 7.9B parameters). They are used for text classification, text generation.

Published2026-10-05
Authors8
Citing Models12
arXiv2610.06778

Abstract

While modern AI detectors identify who wrote the words, emerging policies on AI use increasingly hinge on a different question: who came up with the ideas? We introduce IdeaLens, a detector that identifies whether a document's ideas came from a human or AI (idea provenance), regardless of who wrote its words. To focus IdeaLens on ideas rather than prose, we represent documents as outlines: lists of items that each pair a discourse role with a brief, paraphrased description of the content, minimizing word-level overlap with the raw text. We train IdeaLens on 1M FineWeb documents with silver labels from Pangram, a prose provenance detector. Since the outlines are largely stripped of surface-level information, the labels must be fit mainly through the ideas. In a controlled study, IdeaLens's AI flag rate drops from 95% to 7% as models write from increasingly detailed human plans, while Pangram 4 still flags 92%; from AI-derived plans, IdeaLens stays above 96%. Conversely, on a new dataset of 50 stories that human authors wrote from AI-generated plans, IdeaLens flags 68% of the stories as AI, compared to 8% for Pangram 4. On a comprehensive suite of 19 existing detection benchmarks, we show that IdeaLens maintains strong detection rates at low false positive rates, suggesting that ideas themselves provide a powerful discriminative signal, and its performance holds across domains, formats, and languages. Finally, we examine 90K predictions from IdeaLens to characterize systematic differences between human and AI ideation. We release our models and labeled datasets to facilitate future research on idea provenance detection.

Full paper on arXiv

Details

arXiv identifier
2610.06778
Published
2026-10-05
Authors
Rishanth Rajendhran, Minjoon Choi, Jenna Russell, Ramya Namuduri, Deniz Bölöni-Turgut, Marzena Karpinska, John Wieting, Mohit Iyyer

Open Models Built on This Paper

Every model in the SAVRN Model Hub whose card cites this paper, most downloaded first, with what it takes to run each one.

ModelTaskSizeLicenseMonthly downloadsCheapest setup at 16-bit
IdeaLens-Qwen3.5-9B
Rishanth Rajendhran
Text classification 7.9B cc-by-nc-sa-4.0 72 1x MI300X $1.85/hr
IdeaLens
Rishanth Rajendhran
Text generation 32.9B other 55 1x MI300X $1.85/hr
IdeaLens-NoParaphrase
Rishanth Rajendhran
Text generation 32.9B other 33 1x MI300X $1.85/hr
IdeaLens-ModernBERT-L-NoParaphrase
Rishanth Rajendhran
Text classification 396M cc-by-nc-sa-4.0 12 1x MI300X $1.85/hr
ProseLens
Rishanth Rajendhran
Text generation 32.9B other 11 1x MI300X $1.85/hr
IdeaLens-Qwen3.5-9B-PerItem
Rishanth Rajendhran
Text classification 7.9B cc-by-nc-sa-4.0 9 1x MI300X $1.85/hr
ProseLens-ModernBERT-L
Rishanth Rajendhran
Text classification 396M cc-by-nc-sa-4.0 8 1x MI300X $1.85/hr
IdeaLens-ModernBERT-L
Rishanth Rajendhran
Text classification 396M cc-by-nc-sa-4.0 8 1x MI300X $1.85/hr
IdeaLens-ModernBERT-L-RolesOnly
Rishanth Rajendhran
Text classification 396M cc-by-nc-sa-4.0 8 1x MI300X $1.85/hr
IdeaLens-ModernBERT-L-PerItem
Rishanth Rajendhran
Text classification 396M cc-by-nc-sa-4.0 8 1x MI300X $1.85/hr
IdeaLens-LogisticClassifier
Rishanth Rajendhran
— — cc-by-nc-sa-4.0 — —
IdeaLens-LogisticClassifier-PerItem
Rishanth Rajendhran
— — cc-by-nc-sa-4.0 — —

By task: Text classification (7) · Text generation (3)