Datasets › AMI Meeting Corpus

AMI Meeting Corpus

archive 2025-07-28

The AMI Meeting Corpus is a multi-modal data set comprising 100 hours of meeting recordings. It has been meticulously curated for research purposes and includes various modes of data capture. Let me provide you with more details:

  1. Purpose and Context:
  2. The corpus was created in the context of a project that aims to develop meeting browsing technology.
  3. Eventually, it will be publicly released for use by researchers and practitioners.

  4. Data Composition:

  5. 100 hours of recorded meetings are included.
  6. The data is collected from various domains and scenarios.
  7. Around two-thirds of the data involves participants playing different roles in a design team, taking a design project from kick-off to completion over the course of a day.
  8. The remaining data consists of naturally occurring meetings.

  9. Modalities and Annotations:

  10. The corpus includes synchronized recording devices such as close-talking and far-field microphones, individual and room-view video cameras, projection, a whiteboard, and individual pens.
  11. Annotations cover various phenomena, including orthographic transcription, dialog acts, and head movement.

  12. Research Applications:

  13. Although initially designed for meeting browsing technology, the AMI Meeting Corpus is useful for a wide range of research areas.
  14. Researchers engaged in video processing can access higher resolution videos.

Source: Conversation with Bing, 3/16/2024 (1) AMI Corpus - University of Edinburgh. https://groups.inf.ed.ac.uk/ami/corpus/. (2) The AMI Meeting Corpus: A Pre-announcement. https://www.research.ed.ac.uk/en/publications/the-ami-meeting-corpus-a-pre-announcement. (3) The AMI Meeting Corpus: A Pre-announcement | SpringerLink. https://link.springer.com/chapter/10.1007/11677482_3. (4) The AMI Meeting Corpus: A Pre-announcement — University of Twente .... https://research.utwente.nl/en/publications/the-ami-meeting-corpus-a-pre-announcement.

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Meeting Summarization AMI Meeting Corpus UNS ROUGE-1 F1 37.53 Unsupervised Abstractive Meeting Summarization with... xcfcode/Summarization-Papers +3 1 Compare

Papers archive 2025-07-28

1 shown of 1 paper with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 6. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • AMI Meeting Corpus

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections