{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/gaitmixer-skeleton-based-gait-representation","title":"GaitMixer: Skeleton-based Gait Representation Learning via Wide-spectrum Multi-axial Mixer","arxiv_id":"2210.15491","date":"2022-10-27","proceeding":null,"authors":["Ekkasit Pinyoanuntapong","Ayman Ali","Pu Wang","Minwoo Lee","Chen Chen"],"abstract":"Most existing gait recognition methods are appearance-based, which rely on the silhouettes extracted from the video data of human walking activities. The less-investigated skeleton-based gait recognition methods directly learn the gait dynamics from 2D/3D human skeleton sequences, which are theoretically more robust solutions in the presence of appearance changes caused by clothes, hairstyles, and carrying objects. However, the performance of skeleton-based solutions is still largely behind the appearance-based ones. This paper aims to close such performance gap by proposing a novel network model, GaitMixer, to learn more discriminative gait representation from skeleton sequence data. In particular, GaitMixer follows a heterogeneous multi-axial mixer architecture, which exploits the spatial self-attention mixer followed by the temporal large-kernel convolution mixer to learn rich multi-frequency signals in the gait feature maps. Experiments on the widely used gait database, CASIA-B, demonstrate that GaitMixer outperforms the previous SOTA skeleton-based methods by a large margin while achieving a competitive performance compared with the representative appearance-based solutions. Code will be available at https://github.com/exitudio/gaitmixer","url_abs":"https://arxiv.org/abs/2210.15491v2","url_pdf":"https://arxiv.org/pdf/2210.15491v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"gaitmixer-skeleton-based-gait-representation","repo_url":"https://github.com/exitudio/gaitmixer","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"gait-recognition","task_name":"Gait Recognition"},{"task_slug":"multiview-gait-recognition","task_name":"Multiview Gait Recognition"},{"task_slug":"representation-learning","task_name":"Representation Learning"}],"methods":[{"method_slug":"absolute-position-encodings","method_name":"Absolute Position Encodings"},{"method_slug":"adam","method_name":"Adam"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"bpe","method_name":"BPE"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"label-smoothing","method_name":"Label Smoothing"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"position-wise-feed-forward-layer","method_name":"Position-Wise Feed-Forward Layer"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"transformer","method_name":"Transformer"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/multiview-gait-recognition-on-casia-b","task":"Multiview Gait Recognition","dataset":"CASIA-B","model":"GaitMixer","rank_in_archive_order":9,"of":12,"metrics":{"Accuracy (Cross-View, Avg)":"88.3","BG#1-2":"85.6","CL#1-2":"84.5","NM#5-6 ":"94.9"},"uses_additional_data":false},{"leaderboard":"/sota/multiview-gait-recognition-on-casia-b","task":"Multiview Gait Recognition","dataset":"CASIA-B","model":"GaitFormer","rank_in_archive_order":11,"of":12,"metrics":{"Accuracy (Cross-View, Avg)":"83.4","BG#1-2":"81.4","CL#1-2":"77.2","NM#5-6 ":"91.5"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2210.15491","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}