Addressee-labeled dialogue acts extracted from the AMI Meeting Corpus manual annotations v1.6.2 (groups.inf.ed.ac.uk/ami/download/, CC BY 4.0) for training/evaluating addressee detection ("who is this utterance addressed to") as a candidate proposer for…
Dataset Card
By Josh Estrada, published under cc-by-4.0, revision c874199a3441.
AMI Addressee (stage-1 extraction)
Addressee-labeled dialogue acts extracted from the AMI Meeting Corpus manual annotations v1.6.2 (groups.inf.ed.ac.uk/ami/download/, CC BY 4.0) for training/evaluating addressee detection ("who is this utterance addressed to") as a candidate proposer for multi-party floor control.
Scale — important caveat
The manual annotations mark addressee only on a 22-meeting subset (76 speaker files).
Across all 139 meetings / 117,915 dialogue acts, exactly 8,876 dacts carry addressee labels;
the remaining ~109K have no addressee markup. Downstream papers reporting "~8.9K AMI DAs with
addressee" (e.g. Malik et al.) refer to this same subset.
Schema
data/records.jsonl— the 8,876 labeled dialogue acts.data/records_all.jsonl— all 110,795 dialogue acts that resolve to words (labeled or not;addressee_rawis null andlabeledis false for unlabeled ones). Use this for building conversational-context windows: the labeled subset is only ~8% of the dialogue flow, and windows need the backchannels/fragments in between.data/participants.json— meeting -> sorted participant channels; all 139 meetings are 4-participant.- Per-record fields:
meeting_id,speaker,da_id,da_type,addressee_raw,other,unexplained,text,start_time,end_time,n_words,labeled.
Label distribution (top of data/stats.json)
Multi-speaker addressees dominate (~6,190 of 8,876, e.g. D,C,B), single-speaker ~3,600.
This differs from the Jovanović addressing-behaviour subset (61.7% individual / 34.2% group) —
treat cross-subset comparability with care.
Provenance
build_ami_addressee.py in this repo reproduces data/* from the official zip.
Details
- Repository
- ACloudCenter/ami-addressee
- Publisher
- Josh Estrada
- Task category
- Text classification
- Tags
- addressee-detection, multiparty-dialogue, floor-control
- Size category
- n<100K
- Languages
- en
- Revision
- c874199a3441c69e9e17debc432db92dd8a25546
- Last updated
- 2026-09-21
Files
18 files, 345.2 MB in total.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| data/participants.json | Data | 4.6 KB | — |
| data/records.jsonl | Data | 2.6 MB | — |
| data/records_all.jsonl | Data | 31.4 MB | 358407494940 |
| data/regex_baseline.json | Data | 1.9 KB | — |
| data/report_metrics_consolidated_v2.json | Data | 3.3 KB | — |
| data/seed_stability.json | Data | 12.7 KB | — |
| data/split_v2.json | Data | 838 B | — |
| data/stats.json | Data | 1.1 KB | — |
| data/tau_dial_widesplit.json | Data | 8.3 KB | — |
| learning_curve.json | Data | 571 B | — |
| mix_degradation.json | Data | 1.4 KB | — |
| report_metrics_consolidated.json | Data | 5.3 KB | — |
| README.md | Documentation | 1.9 KB | — |
| audio_prep.py | Other | 14.0 KB | — |
| build_ami_addressee.py | Other | 9.3 KB | — |
| data/ihm_mels.npz | Other | 263.8 MB | 7a98b28310e1 |
| data/sdm_test5_mels.npz | Other | 47.3 MB | d80a09e28770 |
| .gitattributes | Repository | 2.6 KB | — |
License and Download
- License
- cc-by-4.0
- Access
- No access gate
Released by Josh Estrada through its official repository on Hugging Face. Read the license.