{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/attrivision-advancing-generalization-in","title":"AttriVision: Advancing Generalization in Pedestrian Attribute Recognition using CLIP","arxiv_id":null,"date":"2025-02-18","proceeding":"WACV 2025 2","authors":["Mehran ADIBI SEDEH","Assia Benbihi","Romain MARTIN","Marianne Clausel","Cédric Pradalier"],"abstract":"Pedestrian Attribute Recognition (PAR) is a critical task\r\nin computer vision that identifies semantic attributes such\r\nas gender, age, clothing, and accessories from images of\r\nindividuals. This task is essential in applications such as\r\nsurveillance, smart city infrastructure, and security systems. Despite significant advances in deep learning, PAR\r\nremains challenging due to strong imbalances in the attribute classes and the need for robust generalization across\r\ndifferent datasets and environments. In this work, we address these two limitations with AttriVision, a novel approach that adopts the generic CLIP features to make\r\nPAR better generalize and introduces a new Focal CrossEntropy (FCE) loss function to handle the inherent class\r\nimbalance in PAR datasets. FCE improves the model’s robustness by giving more weight to difficult-to-classify samples. Our method also demonstrates remarkable transferability to other attribute recognition tasks, such as vehicle attributes, without any architectural modifications. This\r\ntransferability makes AttriVision a powerful and versatile\r\ntool for attribute recognition. We validate our approach\r\non the Unified Pedestrian Attribute Recognition (UPAR)\r\ndataset that integrates data from several sources including PA100K, PETA, RAPv2, and Market1501. AttriVision\r\nachieves new state-of-the-art results on UPAR, with a mean\r\naccuracy of 89.4% and an F1 score of 91.9%. These\r\nresults demonstrate the model’s effectiveness in handling\r\nreal-world variability, including differences in image sensors, viewing conditions, and person densities, making it\r\nhighly suitable for a wide range of real-world applications.","url_abs":"https://openaccess.thecvf.com/content/WACV2025W/V3SC/html/Sedeh_AttriVision_Advancing_Generalization_in_Pedestrian_Attribute_Recognition_using_CLIP_WACVW_2025_paper.html","url_pdf":"https://openaccess.thecvf.com/content/WACV2025W/V3SC/papers/Sedeh_AttriVision_Advancing_Generalization_in_Pedestrian_Attribute_Recognition_using_CLIP_WACVW_2025_paper.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"attribute","task_name":"Attribute"},{"task_slug":"pedestrian-attribute-recognition","task_name":"Pedestrian Attribute Recognition"}],"methods":[{"method_slug":"clip","method_name":"CLIP"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/pedestrian-attribute-recognition-on","task":"Pedestrian Attribute Recognition","dataset":"Market1501-Attributes","model":"Attrivision","rank_in_archive_order":1,"of":1,"metrics":{"Accuracy ":"83.8","F1 score":"88.8"},"uses_additional_data":false},{"leaderboard":"/sota/pedestrian-attribute-recognition-on-pa-100k","task":"Pedestrian Attribute Recognition","dataset":"PA-100K","model":"Attrivision","rank_in_archive_order":12,"of":13,"metrics":{"Accuracy ":"89.8"},"uses_additional_data":false},{"leaderboard":"/sota/pedestrian-attribute-recognition-on-pa-100k","task":"Pedestrian Attribute Recognition","dataset":"PA-100K","model":"Attivision","rank_in_archive_order":13,"of":13,"metrics":{"F1 score":"93.1"},"uses_additional_data":false},{"leaderboard":"/sota/pedestrian-attribute-recognition-on-rapv2","task":"Pedestrian Attribute Recognition","dataset":"RAPv2","model":"Attrivision","rank_in_archive_order":4,"of":4,"metrics":{"Accuracy ":"84.2","F1 score":"89.1"},"uses_additional_data":false},{"leaderboard":"/sota/pedestrian-attribute-recognition-on-upar","task":"Pedestrian Attribute Recognition","dataset":"UPAR","model":"Attrivision","rank_in_archive_order":1,"of":1,"metrics":{"Accuracy ":"89.4","F1 score":"91.9"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}