{"url":"/method/ger","slug":"ger","name":"GER","full_name":"Gait Emotion Recognition","full_name_withheld":false,"description_markdown":"We present a novel classifier network called STEP, to classify perceived human emotion from gaits, based on a Spatial Temporal Graph Convolutional Network (ST-[GCN](https://paperswithcode.com/method/gcn)) architecture. Given an RGB video of an individual walking, our formulation implicitly exploits the gait features to classify the perceived emotion of the human into one of four emotions: happy, sad, angry, or neutral. We train STEP on annotated real-world gait videos, augmented with annotated synthetic gaits generated using a novel generative network called STEP-Gen, built on an ST-GCN based Conditional Variational Autoencoder (CVAE). We incorporate a novel push-pull regularization loss in the CVAE formulation of STEP-Gen to generate realistic gaits and improve the classification accuracy of STEP.\r\nWe also release a novel dataset (E-Gait), which consists of 4,227 human gaits annotated with perceived emotions along with thousands of synthetic gaits. In practice, STEP can learn the affective features and exhibits classification accuracy of 88\\% on E-Gait, which is 14--30\\% more accurate over prior methods.","description_state":"present","introduced_year":null,"introduced_by":{"title":"STEP: Spatial Temporal Graph Convolutional Networks for Emotion Perception from Gaits","paper":"/paper/step-spatial-temporal-graph-convolutional","first_author":"Uttaran Bhattacharya","n_authors":6,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/step-spatial-temporal-graph-convolutional"},"source":{"url":"https://arxiv.org/abs/1910.12906v3","title":"STEP: Spatial Temporal Graph Convolutional Networks for Emotion Perception from Gaits","url_on_a_paper_host":true},"code_snippet_url":"https://github.com/UttaranB127/STEP","code_snippet_url_on_a_code_host":true,"categories":[],"n_papers_tagged":18,"archive_num_papers":18,"papers_newest_first":[{"paper":"/paper/llm-based-generative-error-correction-for","title":"LLM-based Generative Error Correction for Rare Words with Synthetic Data and Phonetic Context","date":"2025-05-23","arxiv_id":"2505.17410","n_code_links":1,"syntology":null},{"paper":"/paper/listening-and-seeing-again-generative-error","title":"Listening and Seeing Again: Generative Error Correction for Audio-Visual Speech Recognition","date":"2025-01-03","arxiv_id":"2501.04038","n_code_links":1,"syntology":null},{"paper":null,"title":"Benchmarking Japanese Speech Recognition on ASR-LLM Setups with Multi-Pass Augmented Generative Error Correction","date":"2024-08-29","arxiv_id":"2408.16180","n_code_links":0,"syntology":null},{"paper":null,"title":"A Survey of Deep Learning for Group-level Emotion Recognition","date":"2024-08-13","arxiv_id":"2408.15276","n_code_links":0,"syntology":null},{"paper":"/paper/a-discrete-perspective-towards-the","title":"A Discrete Perspective Towards the Construction of Sparse Probabilistic Boolean Networks","date":"2024-07-16","arxiv_id":"2407.11543","n_code_links":1,"syntology":null},{"paper":null,"title":"Listen Again and Choose the Right Answer: A New Paradigm for Automatic Speech Recognition with Large Language Models","date":"2024-05-16","arxiv_id":"2405.10025","n_code_links":0,"syntology":null},{"paper":null,"title":"MMGER: Multi-modal and Multi-granularity Generative Error Correction with LLM for Joint Accent and Speech Recognition","date":"2024-05-06","arxiv_id":"2405.03152","n_code_links":0,"syntology":null},{"paper":"/paper/a-generative-approach-for-wikipedia-scale","title":"A Generative Approach for Wikipedia-Scale Visual Entity Recognition","date":"2024-03-04","arxiv_id":"2403.02041","n_code_links":2,"syntology":null},{"paper":"/paper/it-s-never-too-late-fusing-acoustic","title":"It's Never Too Late: Fusing Acoustic Information into Large Language Models for Automatic Speech Recognition","date":"2024-02-08","arxiv_id":"2402.05457","n_code_links":1,"syntology":{"ran":2,"of":6,"unverified":4,"pointer_only":0}},{"paper":"/paper/large-language-models-are-efficient-learners","title":"Large Language Models are Efficient Learners of Noise-Robust Speech Recognition","date":"2024-01-19","arxiv_id":"2401.10446","n_code_links":1,"syntology":{"ran":1,"of":1,"unverified":0,"pointer_only":0}},{"paper":"/paper/gpt-4v-with-emotion-a-zero-shot-benchmark-for","title":"GPT-4V with Emotion: A Zero-shot Benchmark for Generalized Emotion Recognition","date":"2023-12-07","arxiv_id":"2312.04293","n_code_links":1,"syntology":{"ran":1,"of":1,"unverified":0,"pointer_only":1}},{"paper":null,"title":"Generative error correction for code-switching speech recognition using large language models","date":"2023-10-17","arxiv_id":"2310.13013","n_code_links":0,"syntology":null},{"paper":null,"title":"Towards A Robust Group-level Emotion Recognition via Uncertainty-Aware Learning","date":"2023-10-06","arxiv_id":"2310.04306","n_code_links":0,"syntology":null},{"paper":null,"title":"Design and Planning of Flexible Mobile Micro-Grids Using Deep Reinforcement Learning","date":"2022-12-08","arxiv_id":"2212.04136","n_code_links":0,"syntology":null},{"paper":"/paper/modeling-fine-grained-information-via","title":"Modeling Fine-grained Information via Knowledge-aware Hierarchical Graph for Zero-shot Entity Retrieval","date":"2022-11-20","arxiv_id":"2211.10991","n_code_links":1,"syntology":null},{"paper":null,"title":"Distributed Privacy-Preserving Electric Vehicle Charging Control Based on Secret Sharing","date":"2021-10-05","arxiv_id":"2110.02280","n_code_links":0,"syntology":null},{"paper":null,"title":"Generative Ensemble Regression: Learning Particle Dynamics from Observations of Ensembles with Physics-Informed Deep Generative Models","date":"2020-08-05","arxiv_id":"2008.01915","n_code_links":0,"syntology":null},{"paper":"/paper/step-spatial-temporal-graph-convolutional","title":"STEP: Spatial Temporal Graph Convolutional Networks for Emotion Perception from Gaits","date":"2019-10-28","arxiv_id":"1910.12906","n_code_links":1,"syntology":null}],"papers_shown":18,"tasks":[{"task":"/task/automatic-speech-recognition-2","name":"Automatic Speech Recognition","papers":8},{"task":"/task/automatic-speech-recognition","name":"Automatic Speech Recognition (ASR)","papers":8},{"task":"/task/speech-recognition","name":"Speech Recognition","papers":8},{"task":"/task/speech-recognition-1","name":"speech-recognition","papers":8},{"task":"/task/emotion-recognition","name":"Emotion Recognition","papers":3},{"task":"/task/language-modelling","name":"Language Modelling","papers":3},{"task":"/task/audio-visual-speech-recognition","name":"Audio-Visual Speech Recognition","papers":2},{"task":"/task/language-modeling","name":"Language Modeling","papers":2},{"task":"/task/sentence","name":"Sentence","papers":2},{"task":"/task/specificity","name":"Specificity","papers":2},{"task":"/task/visual-speech-recognition","name":"Visual Speech Recognition","papers":2},{"task":"/task/benchmarking","name":"Benchmarking","papers":1},{"task":"/task/cloze-test","name":"Cloze Test","papers":1},{"task":"/task/deep-learning","name":"Deep Learning","papers":1},{"task":"/task/deep-reinforcement-learning","name":"Deep Reinforcement Learning","papers":1},{"task":"/task/denoising","name":"Denoising","papers":1},{"task":"/task/distributed-optimization","name":"Distributed Optimization","papers":1},{"task":"/task/entity-retrieval","name":"Entity Retrieval","papers":1},{"task":"/task/facial-emotion-recognition","name":"Facial Emotion Recognition","papers":1},{"task":"/task/classification","name":"General Classification","papers":1}],"tasks_shown":20,"n_tasks":41,"usage_by_year":[{"year":"2019","papers":1},{"year":"2020","papers":1},{"year":"2021","papers":1},{"year":"2022","papers":2},{"year":"2023","papers":3},{"year":"2024","papers":8},{"year":"2025","papers":2}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/ger"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}