Browse State-of-the-Art › Text based Person Retrieval
Text based Person Retrieval
30 papers with code · 3 benchmarks · 3 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
3 leaderboard tables shown for this task, 3 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| CUHK-PEDES (21 rows) | MARS | MARS: Paying more attention to visual attributes for text-based... | code | — | Compare |
| ICFG-PEDES (12 rows) | APTM | Towards Unified Text-based Person Retrieval: A Large-scale... | code | Syntology ran 6 of 12 samples · 6 unverified | Compare |
| RSTPReid (9 rows) | MARS | MARS: Paying more attention to visual attributes for text-based... | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
3 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 30 papers with code (49 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
9 Dec 2023 2 repositories listedPerson Re-Identification (Re-ID) task seeks to enhance the tracking of multiple individuals by surveillance cameras.
-
16 Jul 2022 2 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)In PGU, we adopt a set of shared and learnable prototypes as the queries to extract diverse and semantically aligned features for both modalities in the granularity-unified feature space, which further promotes the ReID…
-
8 Jan 2021 2 repositories listedSecondly, a BERT with locality-constrained attention is proposed to obtain representations of descriptions at different scales.
-
15 May 2020 2 repositories listedPerson search by natural language aims at retrieving a specific person in a large-scale image pool that matches the given textual descriptions.
-
15 Nov 2017 2 repositories listedIn this paper, we propose a new system to discriminatively embed the image and text to a shared visual-textual space.
-
26 Apr 2025 1 repository listedDue to the high cost of annotation and privacy protection, researchers resort to synthesized data for the paradigm of pretraining and fine-tuning.
-
14 Apr 2025 1 repository listedText-based Person Retrieval (TPR) as a multi-modal task, which aims to retrieve the target person from a pool of candidate images given a text description, has recently garnered considerable attention due to the…
-
28 Mar 2025 1 repository listed(1) We propose an inter-class image generation pipeline, in which an automatic prompt construction strategy is introduced to guide generative Artificial Intelligence (AI) models in generating various inter-class images…
-
5 Jul 2024 1 repository listedThe task is that of retrieving one or more images of a specific individual based on a textual description.
-
16 Apr 2024 1 repository listedIn text-based person search endeavors, data generation has emerged as a prevailing practice, addressing concerns over privacy preservation and the arduous task of manual annotation.
-
6 Dec 2023 1 repository listedFirstly, we construct a new \textbf{dataset} named UFine6926.
-
25 Nov 2023 1 repository listedTo address the above limitations, we propose a new Composed Person Retrieval (CPR) task, which combines visual and textual queries to identify individuals of interest from large-scale person image databases.
-
18 Sep 2023 1 repository listedText-based Person Retrieval (TPR) aims to retrieve the target person images given a textual query.
-
19 Aug 2023 1 repository listedTPBS, as a fine-grained cross-modal retrieval task, is also facing the rise of research on the CLIP-based TBPS.
-
19 Aug 2023 1 repository listed Syntology ran 12 of 14 samples · 2 unverified · 14 pointer-only (licence)Text-to-image person re-identification (TIReID) is a compelling topic in the cross-modal community, which aims to retrieve the target person based on a textual query.
-
5 Jun 2023 1 repository listed Syntology ran 6 of 12 samples · 6 unverifiedTo verify the feasibility of learning from the generated data, we develop a new joint Attribute Prompt Learning and Text Matching Learning (APTM) framework, considering the shared knowledge between attribute and text.
-
23 May 2023 1 repository listed Syntology ran 3 of 7 samples · 4 unverifiedRA offsets the overfitting risk by introducing a novel positive relation detection task (i.
-
15 May 2023 1 repository listed Syntology ran 0 of 9 samples · 9 unverifiedTo address this issue, we propose a novel language-image pre-training framework for person representation learning, termed PLIP.
-
22 Mar 2023 1 repository listed Syntology ran 5 of 9 samples · 4 unverifiedTo alleviate these issues, we present IRRA: a cross-modal Implicit Relation Reasoning and Aligning framework that learns relations between local visual-textual tokens and enhances global image-text matching without…
-
4 Nov 2022 1 repository listedText-based person search aims to associate pedestrian images with natural language descriptions.
-
19 Oct 2022 1 repository listed Syntology ran 10 of 12 samples · 2 unverifiedSecondly, cross-grained feature refinement (CFR) and fine-grained correspondence discovery (FCD) modules are proposed to establish the cross-grained and fine-grained interactions between modalities, which can filter out…
-
18 Aug 2022 1 repository listedTo explore the fine-grained alignment, we further propose two implicit semantic alignment paradigms: multi-level alignment (MLA) and bidirectional mask modeling (BMM).
-
13 Dec 2021 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedIn this paper, we propose a semantic-aligned embedding method for text-based person search, in which the feature alignment across modalities is achieved by automatically learning the semantic-aligned visual features and…
-
20 Oct 2021 1 repository listedFirstly, to fully utilize the existing small-scale benchmarking datasets for more discriminative feature learning, we introduce a cross-modal momentum contrastive learning framework to enrich the training data for a…
-
27 Sep 2021 1 repository listedFinding target persons in full scene images with a query of text description has important practical applications in intelligent video surveillance.
-
12 Sep 2021 1 repository listedMany previous methods on text-based person retrieval tasks are devoted to learning a latent common space mapping, with the purpose of extracting modality-invariant features from both visual and textual modality.
-
27 Jul 2021 1 repository listedThird, we introduce a Compound Ranking (CR) loss that makes use of textual descriptions for other images of the same identity to provide extra supervision, thereby effectively reducing the intra-class variance in…
-
25 May 2021 1 repository listedText-based person search is a sub-task in the field of image retrieval, which aims to retrieve target person images according to a given textual description.
-
1 Sep 2018 1 repository listedThe key point of image-text matching is how to accurately measure the similarity between visual and textual inputs.
-
19 Feb 2017 1 repository listedSearching persons in large-scale image databases with the query of natural language description has important applications in video surveillance.
Syntology lines on 8 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections