Dataset · Text generation
singapore-legal-ai-benchmark
by Jonathan Su Yuntao JonathanSu/singapore-legal-ai-benchmark
Public research release of 102 Singapore legal research questions, model responses from 6 systems, and overlapping grades on five dimensions. Headline metrics are overlapping binary flags, not a ranking and not a partition of 100%.
Dataset Card
By Jonathan Su Yuntao, published under cc-by-4.0, revision 1bb16210d02c.
Public research release of 102 Singapore legal research questions, model responses from 6 systems, and overlapping grades on five dimensions. Headline metrics are overlapping binary flags, not a ranking and not a partition of 100%.
Interactive explorer
Open the explorer → — comparison table, category heatmap, per-question comparison, and every answer with its sources and grades. (Space page)
Overall (n = 612)
| Metric | Count | Rate | Wilson 95% |
|---|---|---|---|
| Hallucination | 447 | 73.0% | 69.4–76.4 |
| Incorrectness | 402 | 65.7% | 61.8–69.3 |
| Misgroundedness | 403 | 65.8% | 62.0–69.5 |
| Incompleteness | 413 | 67.5% | 63.7–71.1 |
| Substantial Correctness | 399 | 65.2% | 61.3–68.9 |
Hallucination is Incorrectness OR Misgroundedness. Incompleteness and Substantial Correctness are independent of Hallucination. A response may be both Substantially Correct and Hallucinated.
Hallucination × Substantial Correctness: H yes / SC yes = 269; H yes / SC no = 178; H no / SC yes = 130; H no / SC no = 35.
Comparison table
Each cell is count (percentage; Wilson 95%). Systems are listed in the paper's
fixed comparison order, not as a ranking.
| System | n | Hallucination | Incorrectness | Misgroundedness | Incompleteness | Substantial Correctness |
|---|---|---|---|---|---|---|
| LawNet AI | 102 | 74 (72.5%; 63–80) | 70 (68.6%; 59–77) | 59 (57.8%; 48–67) | 77 (75.5%; 66–83) | 34 (33.3%; 25–43) |
| Pair Search | 102 | 81 (79.4%; 71–86) | 74 (72.5%; 63–80) | 75 (73.5%; 64–81) | 62 (60.8%; 51–70) | 64 (62.7%; 53–72) |
| Lucio AI | 102 | 72 (70.6%; 61–79) | 63 (61.8%; 52–71) | 68 (66.7%; 57–75) | 65 (63.7%; 54–72) | 82 (80.4%; 72–87) |
| Harvey | 102 | 86 (84.3%; 76–90) | 84 (82.4%; 74–88) | 79 (77.5%; 68–84) | 88 (86.3%; 78–92) | 44 (43.1%; 34–53) |
| Perplexity | 102 | 76 (74.5%; 65–82) | 70 (68.6%; 59–77) | 67 (65.7%; 56–74) | 90 (88.2%; 81–93) | 82 (80.4%; 72–87) |
| GPT 5.6 Sol | 102 | 58 (56.9%; 47–66) | 41 (40.2%; 31–50) | 55 (53.9%; 44–63) | 31 (30.4%; 22–40) | 93 (91.2%; 84–95) |
Hallucination by question category
| Category | LawNet AI | Pair Search | Lucio AI | Harvey | Perplexity | GPT 5.6 Sol |
|---|---|---|---|---|---|---|
| 1. Tracing lines of authority | 10/10 (100.0%) | 8/10 (80.0%) | 7/10 (70.0%) | 7/10 (70.0%) | 9/10 (90.0%) | 6/10 (60.0%) |
| 2. Diverging authorities | 8/10 (80.0%) | 8/10 (80.0%) | 10/10 (100.0%) | 7/10 (70.0%) | 10/10 (100.0%) | 6/10 (60.0%) |
| 3. Ratio and obiter | 7/10 (70.0%) | 8/10 (80.0%) | 10/10 (100.0%) | 8/10 (80.0%) | 9/10 (90.0%) | 7/10 (70.0%) |
| 4. Material factual distinctions | 6/10 (60.0%) | 7/10 (70.0%) | 8/10 (80.0%) | 8/10 (80.0%) | 6/10 (60.0%) | 5/10 (50.0%) |
| 5. Statute tracing | 6/10 (60.0%) | 6/10 (60.0%) | 0/10 (0.0%) | 10/10 (100.0%) | 4/10 (40.0%) | 6/10 (60.0%) |
| 6. Historical statutory application | 7/10 (70.0%) | 10/10 (100.0%) | 10/10 (100.0%) | 10/10 (100.0%) | 8/10 (80.0%) | 2/10 (20.0%) |
| 7. Good Law Status | 12/12 (100.0%) | 12/12 (100.0%) | 11/12 (91.7%) | 12/12 (100.0%) | 11/12 (91.7%) | 8/12 (66.7%) |
| 8. Adapted Magesh-style baseline | 18/30 (60.0%) | 22/30 (73.3%) | 16/30 (53.3%) | 24/30 (80.0%) | 19/30 (63.3%) | 18/30 (60.0%) |
Files
| Responses table | graded_responses.csv |
| Paper aggregates | aggregates/ |
| Datasheet | DATASHEET.md |
The Dataset Viewer uses a single responses config and loads only
graded_responses.csv (612 rows: six systems × 102 questions, in the paper's
comparison order). Files under aggregates/ are summary tables for download,
not additional viewer configs.
graded_responses.csv contains the questions, answers, recorded sources, access
mode (API or Manual interface), the system configuration sentence, public
grades, error-pattern labels, and collection timestamps (started_at /
finished_at) where recorded. Hallucination is Incorrectness OR
Misgroundedness. Timestamps are collection wall-clock times, not grading time;
Harvey rows are blank.
Citation
@misc{singapore_legal_ai_benchmark,
title = {Singapore Legal AI Benchmark},
author = {Su, Jonathan Yuntao and Tan, Zhen Jie Adam and Tan, Mei Bin and Lu, Isaac Yang},
year = {2026},
howpublished = {\url{https://huggingface.co/datasets/JonathanSu/singapore-legal-ai-benchmark}},
note = {Questions, model responses, and hallucination / grounding scores}
}
Su, J. Y., Tan, Z. J. A., Tan, M. B., & Lu, I. Y. (2026). Singapore Legal AI Benchmark. Licensed for research use; model outputs remain subject to each provider's terms.
License
Intended for research citation and reuse. Model outputs remain subject to the respective providers' terms of use.
Details
- Repository
- JonathanSu/singapore-legal-ai-benchmark
- Publisher
- Jonathan Su Yuntao
- Task category
- Text generation
- Tags
- legal, singapore, benchmark
- Size category
- n<1K
- Languages
- en
- Revision
- 1bb16210d02c5d538d448048e278494c2ed21d6b
- Last updated
- 2026-09-18
Files
12 files, 5.8 MB in total.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| aggregates/01_provider_summary.csv | Data | 638 B | — |
| aggregates/02_category_summary.csv | Data | 4.3 KB | — |
| aggregates/03_category_overall.csv | Data | 767 B | — |
| aggregates/04_grounding_deficiencies.csv | Data | 583 B | — |
| aggregates/05_diagnostic_labels.csv | Data | 14.6 KB | — |
| aggregates/06_hallucination_substantial_overlap.csv | Data | 397 B | — |
| aggregates/07_metric_and_label_definitions.csv | Data | 7.7 KB | — |
| aggregates/08_metric_intervals.csv | Data | 26.3 KB | — |
| graded_responses.csv | Data | 5.7 MB | — |
| DATASHEET.md | Documentation | 3.5 KB | — |
| README.md | Documentation | 5.8 KB | — |
| .gitattributes | Repository | 2.5 KB | — |
License and Download
- License
- cc-by-4.0
- Access
- No access gate
Released by Jonathan Su Yuntao through its official repository on Hugging Face. Read the license.