Datasets › GRID Dataset
GRID Dataset
The QMUL underGround Re-IDentification (GRID) dataset contains 250 pedestrian image pairs. Each pair contains two images of the same individual seen from different camera views. All images are captured from 8 disjoint camera views installed in a busy underground station. The figures beside show a snapshot of each of the camera views of the station and sample images in the dataset. The dataset is challenging due to variations of pose, colours, lighting changes; as well as poor image quality caused by low spatial resolution.
Source: https://personal.ie.cuhk.edu.hk/~ccloy/downloads_qmul_underground_reid.html
Benchmarks archive 2025-07-28
All 5 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Lipreading | GRID corpus (mixed-speech) | CTC/Attention Word Error Rate (WER) 1.2 | Visual Speech Recognition for Multiple Languages in the Wild | mpc001/Visual_Speech_Recognition_for_Multiple_Languages +1 | 5 | Compare |
| Speaker-Specific Lip to Speech Synthesis | GRID corpus (mixed-speech) | Visual Voice Memory ESTOI 0.579 | Speech Reconstruction with Reminiscent Sound via Visual... | joannahong/Speech-Reconstruction-with-Reminiscent-Sound-via-Visual-Voice-Memory | 2 | Compare |
| Speech Separation | GRID corpus (mixed-speech) | SDR 9.6 | MIDI: Multi-Instance Diffusion for Single Image to 3D... | — | 2 | Compare |
| Lip Reading | GRID corpus (mixed-speech) | Lip2Wav WER 14.08 | Learning Individual Speaking Styles for Accurate Lip to... | Rudrabha/Lip2Wav | 1 | Compare |
| Speech Enhancement | GRID corpus (mixed-speech) | Audio-Visual concat-ref PESQ 2.70 | Face Landmark-based Speaker-Independent Audio-Visual... | dr-pato/audio_visual_speech_enhancement | 1 | Compare |
Papers archive 2025-07-28
9 shown of 9 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 10. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation | 0 | 1 | 4 Dec 2024 | not harvested |
| Visual Speech Recognition for Multiple Languages in the Wild | 2 | 1 | 26 Feb 2022 | not harvested |
| Speech Reconstruction with Reminiscent Sound via Visual Voice Memory | 1 | 1 | 17 Nov 2021 | not harvested |
| Learning Individual Speaking Styles for Accurate Lip to Speech Synthesis | 1 | 2 | 17 May 2020 | ran 1 of 9 samples (8 unverified) |
| Can We Read Speech Beyond the Lips? Rethinking RoI Selection for Deep Visual Speech Recognition | 1 | 1 | 6 Mar 2020 | not harvested |
| Face Landmark-based Speaker-Independent Audio-Visual Speech Enhancement in Multi-Talker Environments | 1 | 2 | 6 Nov 2018 | not harvested |
| LCANet: End-to-End Lipreading with Cascaded Attention-CTC | 0 | 1 | 13 Mar 2018 | not harvested |
| Lip Reading Sentences in the Wild | 0 | 1 | 16 Nov 2016 | not harvested |
| LipNet: End-to-End Sentence-level Lipreading | 13 | 1 | 5 Nov 2016 | ran 0 of 15 samples (15 unverified) |
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
No licence recorded in the archive. Absence here is not a statement about the dataset's terms.
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- GRID corpus (mixed-speech)
- GRID Dataset
2 variant names, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections