{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-deep-context-aware-features-over","title":"Learning Deep Context-aware Features over Body and Latent Parts for Person Re-identification","arxiv_id":"1710.06555","date":"2017-10-18","proceeding":"CVPR 2017 7","authors":["Dangwei Li","Xiaotang Chen","Zhang Zhang","Kaiqi Huang"],"abstract":"Person Re-identification (ReID) is to identify the same person across\ndifferent cameras. It is a challenging task due to the large variations in\nperson pose, occlusion, background clutter, etc How to extract powerful\nfeatures is a fundamental problem in ReID and is still an open problem today.\nIn this paper, we design a Multi-Scale Context-Aware Network (MSCAN) to learn\npowerful features over full body and body parts, which can well capture the\nlocal context knowledge by stacking multi-scale convolutions in each layer.\nMoreover, instead of using predefined rigid parts, we propose to learn and\nlocalize deformable pedestrian parts using Spatial Transformer Networks (STN)\nwith novel spatial constraints. The learned body parts can release some\ndifficulties, eg pose variations and background clutters, in part-based\nrepresentation. Finally, we integrate the representation learning processes of\nfull body and body parts into a unified framework for person ReID through\nmulti-class person identification tasks. Extensive evaluations on current\nchallenging large-scale person ReID datasets, including the image-based\nMarket1501, CUHK03 and sequence-based MARS datasets, show that the proposed\nmethod achieves the state-of-the-art results.","url_abs":"http://arxiv.org/abs/1710.06555v1","url_pdf":"http://arxiv.org/pdf/1710.06555v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"person-identification","task_name":"Person Identification"},{"task_slug":"person-re-identification","task_name":"Person Re-Identification"},{"task_slug":"representation-learning","task_name":"Representation Learning"}],"methods":[{"method_slug":"absolute-position-encodings","method_name":"Absolute Position Encodings"},{"method_slug":"adam","method_name":"Adam"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"bpe","method_name":"BPE"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"label-smoothing","method_name":"Label Smoothing"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"position-wise-feed-forward-layer","method_name":"Position-Wise Feed-Forward Layer"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"spatial-transformer","method_name":"Spatial Transformer"},{"method_slug":"transformer","method_name":"Transformer"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/person-re-identification-on-market-1501","task":"Person Re-Identification","dataset":"Market-1501","model":"MSCAN","rank_in_archive_order":113,"of":135,"metrics":{"Rank-1":"80.31","mAP":"57.53"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1710.06555","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}