Papers › AttriVision: Advancing Generalization in Pedestrian Attribute Recognition using CLIP
AttriVision: Advancing Generalization in Pedestrian Attribute Recognition using CLIP
Mehran ADIBI SEDEH, Assia Benbihi, Romain MARTIN, Marianne Clausel, Cédric Pradalier
Pedestrian Attribute Recognition (PAR) is a critical task in computer vision that identifies semantic attributes such as gender, age, clothing, and accessories from images of individuals. This task is essential in applications such as surveillance, smart city infrastructure, and security systems. Despite significant advances in deep learning, PAR remains challenging due to strong imbalances in the attribute classes and the need for robust generalization across different datasets and environments. In this work, we address these two limitations with AttriVision, a novel approach that adopts the generic CLIP features to make PAR better generalize and introduces a new Focal CrossEntropy (FCE) loss function to handle the inherent class imbalance in PAR datasets. FCE improves the model’s robustness by giving more weight to difficult-to-classify samples. Our method also demonstrates remarkable transferability to other attribute recognition tasks, such as vehicle attributes, without any architectural modifications. This transferability makes AttriVision a powerful and versatile tool for attribute recognition. We validate our approach on the Unified Pedestrian Attribute Recognition (UPAR) dataset that integrates data from several sources including PA100K, PETA, RAPv2, and Market1501. AttriVision achieves new state-of-the-art results on UPAR, with a mean accuracy of 89.4% and an F1 score of 91.9%. These results demonstrate the model’s effectiveness in handling real-world variability, including differences in image sensors, viewing conditions, and person densities, making it highly suitable for a wide range of real-world applications.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Pedestrian Attribute Recognition | Market1501-Attributes | Attrivision | Accuracy | 83.8 | #1 of 1 | Archive leaderboard | report |
| Pedestrian Attribute Recognition | Market1501-Attributes | Attrivision | F1 score | 88.8 | #1 of 1 | Archive leaderboard | report |
| Pedestrian Attribute Recognition | PA-100K | Attrivision | Accuracy | 89.8 | #12 of 13 | Archive leaderboard | report |
| Pedestrian Attribute Recognition | PA-100K | Attivision | F1 score | 93.1 | #13 of 13 | Archive leaderboard | report |
| Pedestrian Attribute Recognition | RAPv2 | Attrivision | Accuracy | 84.2 | #4 of 4 | Archive leaderboard | report |
| Pedestrian Attribute Recognition | RAPv2 | Attrivision | F1 score | 89.1 | #4 of 4 | Archive leaderboard | report |
| Pedestrian Attribute Recognition | UPAR | Attrivision | Accuracy | 89.4 | #1 of 1 | Archive leaderboard | report |
| Pedestrian Attribute Recognition | UPAR | Attrivision | F1 score | 91.9 | #1 of 1 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections