SAVRN
Search Contact SAVRN

Dataset · Text classification

ami-addressee

by Josh Estrada ACloudCenter/ami-addressee

Addressee-labeled dialogue acts extracted from the AMI Meeting Corpus manual annotations v1.6.2 (groups.inf.ed.ac.uk/ami/download/, CC BY 4.0) for training/evaluating addressee detection ("who is this utterance addressed to") as a candidate proposer for…

Rows—
Configurations—
Size345.2 MB
Licensecc-by-4.0
AccessPublicly accessible
Monthly Downloads—

Dataset Card

By Josh Estrada, published under cc-by-4.0, revision c874199a3441.

AMI Addressee (stage-1 extraction)

Addressee-labeled dialogue acts extracted from the AMI Meeting Corpus manual annotations v1.6.2 (groups.inf.ed.ac.uk/ami/download/, CC BY 4.0) for training/evaluating addressee detection ("who is this utterance addressed to") as a candidate proposer for multi-party floor control.

Scale — important caveat

The manual annotations mark addressee only on a 22-meeting subset (76 speaker files). Across all 139 meetings / 117,915 dialogue acts, exactly 8,876 dacts carry addressee labels; the remaining ~109K have no addressee markup. Downstream papers reporting "~8.9K AMI DAs with addressee" (e.g. Malik et al.) refer to this same subset.

Schema

  • data/records.jsonl — the 8,876 labeled dialogue acts.
  • data/records_all.jsonl — all 110,795 dialogue acts that resolve to words (labeled or not; addressee_raw is null and labeled is false for unlabeled ones). Use this for building conversational-context windows: the labeled subset is only ~8% of the dialogue flow, and windows need the backchannels/fragments in between.
  • data/participants.json — meeting -> sorted participant channels; all 139 meetings are 4-participant.
  • Per-record fields: meeting_id, speaker, da_id, da_type, addressee_raw, other, unexplained, text, start_time, end_time, n_words, labeled.

Label distribution (top of data/stats.json)

Multi-speaker addressees dominate (~6,190 of 8,876, e.g. D,C,B), single-speaker ~3,600. This differs from the Jovanović addressing-behaviour subset (61.7% individual / 34.2% group) — treat cross-subset comparability with care.

Provenance

build_ami_addressee.py in this repo reproduces data/* from the official zip.

Details

Repository
ACloudCenter/ami-addressee
Publisher
Josh Estrada
Task category
Text classification
Tags
addressee-detection, multiparty-dialogue, floor-control
Size category
n<100K
Languages
en
Revision
c874199a3441c69e9e17debc432db92dd8a25546
Last updated
2026-09-21

Files

18 files, 345.2 MB in total.

Data12 files · 34.0 MB
Documentation1 file · 1.9 KB
Other4 files · 311.2 MB
Repository1 file · 2.6 KB
Every file
FileTypeSizeSHA-256
data/participants.jsonData4.6 KB—
data/records.jsonlData2.6 MB—
data/records_all.jsonlData31.4 MB358407494940
data/regex_baseline.jsonData1.9 KB—
data/report_metrics_consolidated_v2.jsonData3.3 KB—
data/seed_stability.jsonData12.7 KB—
data/split_v2.jsonData838 B—
data/stats.jsonData1.1 KB—
data/tau_dial_widesplit.jsonData8.3 KB—
learning_curve.jsonData571 B—
mix_degradation.jsonData1.4 KB—
report_metrics_consolidated.jsonData5.3 KB—
README.mdDocumentation1.9 KB—
audio_prep.pyOther14.0 KB—
build_ami_addressee.pyOther9.3 KB—
data/ihm_mels.npzOther263.8 MB7a98b28310e1
data/sdm_test5_mels.npzOther47.3 MBd80a09e28770
.gitattributesRepository2.6 KB—

License and Download

License
cc-by-4.0
Access
No access gate
Download from Josh Estrada

Released by Josh Estrada through its official repository on Hugging Face. Read the license.