Datasets › CUHK-PEDES

CUHK-PEDES

Introduced by Shuang Li et al. in Person Search with Natural Language Description19 Feb 2017 archive 2025-07-28

The CUHK-PEDES dataset is a caption-annotated pedestrian dataset. It contains 40,206 images over 13,003 persons. Images are collected from five existing person re-identification datasets, CUHK03, Market-1501, SSM, VIPER, and CUHK01 while each image is annotated with 2 text descriptions by crowd-sourcing workers. Sentences incorporate rich details about person appearances, actions, poses.

Source: MGD-GAN: Text-to-Pedestrian generation through Multi-Grained Discrimination Image Source: https://www.researchgate.net/figure/Image-samples-in-three-datasets-For-MSCOCO-and-Flickr30k-dataset-we-view-every-image_fig2_321095980

Benchmarks archive 2025-07-28

All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

25 shown of 25 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 93. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
MARS: Paying more attention to visual attributes for text-based person search 1 1 5 Jul 2024 not harvested
From Data Deluge to Data Curation: A Filtering-WoRA Paradigm for Efficient Text-based Person Search 1 1 16 Apr 2024 not harvested
Cross-Modal Adaptive Dual Association for Text-to-Image Person Retrieval 0 1 4 Dec 2023 not harvested
Noisy-Correspondence Learning for Text-to-Image Person Re-identification 1 1 19 Aug 2023 ran 12 of 14 samples (2 unverified; 14 pointer-only for licence)
Towards Unified Text-based Person Retrieval: A Large-scale Multi-Attribute and Language Search Benchmark 1 1 5 Jun 2023 ran 6 of 12 samples (6 unverified)
RaSa: Relation and Sensitivity Aware Representation Learning for Text-based Person Search 1 1 23 May 2023 ran 3 of 7 samples (4 unverified)
Cross-Modal Implicit Relation Reasoning and Aligning for Text-to-Image Person Retrieval 1 1 22 Mar 2023 ran 5 of 9 samples (4 unverified)
A Simple and Robust Correlation Filtering Method for Text-based Person Search 1 1 4 Nov 2022 not harvested
Deep Evidential Learning with Noisy Correspondence for Cross-Modal Retrieval 1 1 10 Oct 2022 not harvested
See Finer, See More: Implicit Modality Alignment for Text-based Person Retrieval 1 1 18 Aug 2022 not harvested
Learning Semantic-Aligned Feature Representation for Text-based Person Search 1 1 13 Dec 2021 ran 1 of 1 samples (0 unverified)
Text-Based Person Search with Limited Data 1 1 20 Oct 2021 not harvested
DSSL: Deep Surroundings-person Separation Learning for Text-based Person Retrieval 1 1 12 Sep 2021 not harvested
Semantically Self-Aligned Network for Text-to-Image Part-aware Person Re-identification 1 2 27 Jul 2021 not harvested
TIPCB: A Simple but Effective Part-based Convolutional Baseline for Text-based Person Search 1 1 25 May 2021 not harvested
Learning Transferable Visual Models From Natural Language Supervision 82 1 26 Feb 2021 ran 16 of 20 samples (4 unverified; 16 pointer-only for licence)
AXM-Net: Implicit Cross-Modal Feature Alignment for Person Re-identification 0 1 19 Jan 2021 not harvested
Contextual Non-Local Alignment over Full-Scale Representation for Text-Based Person Search 2 1 8 Jan 2021 not harvested
Hierarchical Gumbel Attention Network for Text-based Person Search 0 1 10 Oct 2020 not harvested
ViTAA: Visual-Textual Attributes Alignment in Person Search by Natural Language 2 1 15 May 2020 not harvested
Text-based Person Search via Attribute-aided Matching 0 1 14 Mar 2020 not harvested
Improving Description-based Person Re-identification by Multi-granularity Image-text Alignments 0 1 23 Jun 2019 not harvested
Deep Cross-Modal Projection Learning for Image-Text Matching 1 1 1 Sep 2018 not harvested
Improving Deep Visual Representation for Person Re-identification by Global and Local Image-language Association 0 1 5 Aug 2018 not harvested
Dual-Path Convolutional Image-Text Embeddings with Instance Loss 2 2 15 Nov 2017 not harvested

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Unknown

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • CUHK-PEDES

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections