SAVRN
Search Contact SAVRN

Dataset · Text generation

singapore-legal-ai-benchmark

by Jonathan Su Yuntao JonathanSu/singapore-legal-ai-benchmark

Public research release of 102 Singapore legal research questions, model responses from 6 systems, and overlapping grades on five dimensions. Headline metrics are overlapping binary flags, not a ranking and not a partition of 100%.

Rows
Configurations
Size5.8 MB
Licensecc-by-4.0
AccessPublicly accessible
Monthly Downloads156

Dataset Card

By Jonathan Su Yuntao, published under cc-by-4.0, revision 1bb16210d02c.

Public research release of 102 Singapore legal research questions, model responses from 6 systems, and overlapping grades on five dimensions. Headline metrics are overlapping binary flags, not a ranking and not a partition of 100%.

Interactive explorer

Open the explorer → — comparison table, category heatmap, per-question comparison, and every answer with its sources and grades. (Space page)

Overall (n = 612)

Metric Count Rate Wilson 95%
Hallucination 447 73.0% 69.4–76.4
Incorrectness 402 65.7% 61.8–69.3
Misgroundedness 403 65.8% 62.0–69.5
Incompleteness 413 67.5% 63.7–71.1
Substantial Correctness 399 65.2% 61.3–68.9

Hallucination is Incorrectness OR Misgroundedness. Incompleteness and Substantial Correctness are independent of Hallucination. A response may be both Substantially Correct and Hallucinated.

Hallucination × Substantial Correctness: H yes / SC yes = 269; H yes / SC no = 178; H no / SC yes = 130; H no / SC no = 35.

Comparison table

Each cell is count (percentage; Wilson 95%). Systems are listed in the paper's fixed comparison order, not as a ranking.

System n Hallucination Incorrectness Misgroundedness Incompleteness Substantial Correctness
LawNet AI 102 74 (72.5%; 63–80) 70 (68.6%; 59–77) 59 (57.8%; 48–67) 77 (75.5%; 66–83) 34 (33.3%; 25–43)
Pair Search 102 81 (79.4%; 71–86) 74 (72.5%; 63–80) 75 (73.5%; 64–81) 62 (60.8%; 51–70) 64 (62.7%; 53–72)
Lucio AI 102 72 (70.6%; 61–79) 63 (61.8%; 52–71) 68 (66.7%; 57–75) 65 (63.7%; 54–72) 82 (80.4%; 72–87)
Harvey 102 86 (84.3%; 76–90) 84 (82.4%; 74–88) 79 (77.5%; 68–84) 88 (86.3%; 78–92) 44 (43.1%; 34–53)
Perplexity 102 76 (74.5%; 65–82) 70 (68.6%; 59–77) 67 (65.7%; 56–74) 90 (88.2%; 81–93) 82 (80.4%; 72–87)
GPT 5.6 Sol 102 58 (56.9%; 47–66) 41 (40.2%; 31–50) 55 (53.9%; 44–63) 31 (30.4%; 22–40) 93 (91.2%; 84–95)

Hallucination by question category

Category LawNet AI Pair Search Lucio AI Harvey Perplexity GPT 5.6 Sol
1. Tracing lines of authority 10/10 (100.0%) 8/10 (80.0%) 7/10 (70.0%) 7/10 (70.0%) 9/10 (90.0%) 6/10 (60.0%)
2. Diverging authorities 8/10 (80.0%) 8/10 (80.0%) 10/10 (100.0%) 7/10 (70.0%) 10/10 (100.0%) 6/10 (60.0%)
3. Ratio and obiter 7/10 (70.0%) 8/10 (80.0%) 10/10 (100.0%) 8/10 (80.0%) 9/10 (90.0%) 7/10 (70.0%)
4. Material factual distinctions 6/10 (60.0%) 7/10 (70.0%) 8/10 (80.0%) 8/10 (80.0%) 6/10 (60.0%) 5/10 (50.0%)
5. Statute tracing 6/10 (60.0%) 6/10 (60.0%) 0/10 (0.0%) 10/10 (100.0%) 4/10 (40.0%) 6/10 (60.0%)
6. Historical statutory application 7/10 (70.0%) 10/10 (100.0%) 10/10 (100.0%) 10/10 (100.0%) 8/10 (80.0%) 2/10 (20.0%)
7. Good Law Status 12/12 (100.0%) 12/12 (100.0%) 11/12 (91.7%) 12/12 (100.0%) 11/12 (91.7%) 8/12 (66.7%)
8. Adapted Magesh-style baseline 18/30 (60.0%) 22/30 (73.3%) 16/30 (53.3%) 24/30 (80.0%) 19/30 (63.3%) 18/30 (60.0%)

Files

Responses table graded_responses.csv
Paper aggregates aggregates/
Datasheet DATASHEET.md

The Dataset Viewer uses a single responses config and loads only graded_responses.csv (612 rows: six systems × 102 questions, in the paper's comparison order). Files under aggregates/ are summary tables for download, not additional viewer configs.

graded_responses.csv contains the questions, answers, recorded sources, access mode (API or Manual interface), the system configuration sentence, public grades, error-pattern labels, and collection timestamps (started_at / finished_at) where recorded. Hallucination is Incorrectness OR Misgroundedness. Timestamps are collection wall-clock times, not grading time; Harvey rows are blank.

Citation

@misc{singapore_legal_ai_benchmark,
  title        = {Singapore Legal AI Benchmark},
  author       = {Su, Jonathan Yuntao and Tan, Zhen Jie Adam and Tan, Mei Bin and Lu, Isaac Yang},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/datasets/JonathanSu/singapore-legal-ai-benchmark}},
  note         = {Questions, model responses, and hallucination / grounding scores}
}

Su, J. Y., Tan, Z. J. A., Tan, M. B., & Lu, I. Y. (2026). Singapore Legal AI Benchmark. Licensed for research use; model outputs remain subject to each provider's terms.

License

CC BY 4.0

Intended for research citation and reuse. Model outputs remain subject to the respective providers' terms of use.

Details

Repository
JonathanSu/singapore-legal-ai-benchmark
Publisher
Jonathan Su Yuntao
Task category
Text generation
Tags
legal, singapore, benchmark
Size category
n<1K
Languages
en
Revision
1bb16210d02c5d538d448048e278494c2ed21d6b
Last updated
2026-09-18

Files

12 files, 5.8 MB in total.

Data9 files · 5.8 MB
Documentation2 files · 9.2 KB
Repository1 file · 2.5 KB
Every file
FileTypeSizeSHA-256
aggregates/01_provider_summary.csvData638 B
aggregates/02_category_summary.csvData4.3 KB
aggregates/03_category_overall.csvData767 B
aggregates/04_grounding_deficiencies.csvData583 B
aggregates/05_diagnostic_labels.csvData14.6 KB
aggregates/06_hallucination_substantial_overlap.csvData397 B
aggregates/07_metric_and_label_definitions.csvData7.7 KB
aggregates/08_metric_intervals.csvData26.3 KB
graded_responses.csvData5.7 MB
DATASHEET.mdDocumentation3.5 KB
README.mdDocumentation5.8 KB
.gitattributesRepository2.5 KB

License and Download

License
cc-by-4.0
Access
No access gate
Download from Jonathan Su Yuntao

Released by Jonathan Su Yuntao through its official repository on Hugging Face. Read the license.