Home › Datasets › task › Speaker Diarization
Speaker Diarization datasets
archive 2025-07-28
11 datasets carry the task tag "Speaker Diarization" (the task itself: Speaker Diarization), ordered by the archive's paper count. Page 1 of 1: 11 shown of 11. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Speaker Diarization datasets 1–11 of 11
AVA (Atomic Visual Actions)
AVA is a project that provides audiovisual annotations of video for improving our understanding of human activity.
113 papers · 7 benchmarks
AliMeeting (Multi-Channel Multi-Party Meeting Transcription Challenge)
AliMeeting corpus consists of 120 hours of recorded Mandarin meeting data, including far-field data collected by 8-channel microphone array as well as near-field data collected by headset microphone.
46 papers · 1 benchmark
CHiME-5 (CHiME Speech Separation and Recognition Challenge)
The CHiME challenge series aims to advance robust automatic speech recognition (ASR) technology by promoting research at the interface of speech and language processing, signal processing , and machine learning.
42 papers · 0 benchmarks
The DIHARD II development and evaluation sets draw from a diverse set of sources exhibiting wide variation in recording equipment, recording environment, ambient noise, number of speakers, and speaker demographics.
34 papers · 1 benchmark
Contains temporally labeled face tracks in video, where each face instance is labeled as speaking or not, and whether the speech is audible.
22 papers · 1 benchmark
The MagicData-RAMC corpus contains 180 hours of conversational speech data recorded from native speakers of Mandarin Chinese over mobile phones with a sampling rate of 16 kHz.
12 papers · 0 benchmarks
The CALLHOME English Corpus is a collection of unscripted telephone conversations between native speakers of English.
11 papers · 7 benchmarks
Contains densely labeled speech activity in YouTube videos, with the goal of creating a shared, available dataset for this task.
10 papers · 1 benchmark
A Rich Annotated Mandarin Conversational (RAMC) Speech Dataset, including 180 hours of Mandarin Chinese dialogue, 150, 10 and 20 hours for the training set, development set and test set respectively.
1 paper · 0 benchmarks
FSC-P2 (Fearless Steps Challenge Phase2)
The Fearless Steps Initiative by UTDallas-CRSS led to the digitization, recovery, and diarization of 19,000 hours of original analog audio data, as well as the development of algorithms to extract meaningful information from this…
1 paper · 0 benchmarks
RadioTalk is a corpus of speech recognition transcripts sampled from talk radio broadcasts in the United States between October of 2018 and March of 2019.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.