{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/sequence-to-sequence-modeling-for-action-1","title":"Sequence-to-Sequence Modeling for Action Identification at High Temporal Resolution","arxiv_id":"2111.02521","date":"2021-11-03","proceeding":null,"authors":["Aakash Kaku","Kangning Liu","Avinash Parnandi","Haresh Rengaraj Rajamohan","Kannan Venkataramanan","Anita Venkatesan","Audre Wirtanen","Natasha Pandit","Heidi Schambra","Carlos Fernandez-Granda"],"abstract":"Automatic action identification from video and kinematic data is an important machine learning problem with applications ranging from robotics to smart health. Most existing works focus on identifying coarse actions such as running, climbing, or cutting a vegetable, which have relatively long durations. This is an important limitation for applications that require the identification of subtle motions at high temporal resolution. For example, in stroke recovery, quantifying rehabilitation dose requires differentiating motions with sub-second durations. Our goal is to bridge this gap. To this end, we introduce a large-scale, multimodal dataset, StrokeRehab, as a new action-recognition benchmark that includes subtle short-duration actions labeled at a high temporal resolution. These short-duration actions are called functional primitives, and consist of reaches, transports, repositions, stabilizations, and idles. The dataset consists of high-quality Inertial Measurement Unit sensors and video data of 41 stroke-impaired patients performing activities of daily living like feeding, brushing teeth, etc. We show that current state-of-the-art models based on segmentation produce noisy predictions when applied to these data, which often leads to overcounting of actions. To address this, we propose a novel approach for high-resolution action identification, inspired by speech-recognition techniques, which is based on a sequence-to-sequence model that directly predicts the sequence of actions. This approach outperforms current state-of-the-art methods on the StrokeRehab dataset, as well as on the standard benchmark datasets 50Salads, Breakfast, and Jigsaws.","url_abs":"https://arxiv.org/abs/2111.02521v1","url_pdf":"https://arxiv.org/pdf/2111.02521v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"sequence-to-sequence-modeling-for-action-1","repo_url":"https://github.com/aakashrkaku/seq2seq_hrar","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"high","task_name":"Vocal Bursts Intensity Prediction"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2111.02521","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2111.02521"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/aakashrkaku/seq2seq_hrar","reach":null}],"summary":{"ran":1,"ran_draft_wrong":2,"unverified":1},"by_repo_kind":{"listed":{"samples":4,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":4,"samples":[{"code_sha256_prefix":"32349d8900b86ecd","entry":"Attention","repo":"aakashrkaku/seq2seq_hrar","repo_kind":"listed","path":"imu_data/las_model.py","file_url":"https://github.com/aakashrkaku/seq2seq_hrar/blob/HEAD/imu_data/las_model.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"GPL-3.0","inline_ok":false,"mcp_get_code":{"code_sha256":"32349d8900b86ecd"}},{"code_sha256_prefix":"c8d8031c6643de74","entry":"CreateOnehotVariable","repo":"aakashrkaku/seq2seq_hrar","repo_kind":"listed","path":"imu_data/las_model.py","file_url":"https://github.com/aakashrkaku/seq2seq_hrar/blob/HEAD/imu_data/las_model.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"GPL-3.0","inline_ok":false,"mcp_get_code":{"code_sha256":"c8d8031c6643de74"}},{"code_sha256_prefix":"8da5fe9b71998bfe","entry":"TimeDistributed","repo":"aakashrkaku/seq2seq_hrar","repo_kind":"listed","path":"imu_data/las_model.py","file_url":"https://github.com/aakashrkaku/seq2seq_hrar/blob/HEAD/imu_data/las_model.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"GPL-3.0","inline_ok":false,"mcp_get_code":{"code_sha256":"8da5fe9b71998bfe"}},{"code_sha256_prefix":"847ad682c7e34f92","entry":"Speller","repo":"aakashrkaku/seq2seq_hrar","repo_kind":"listed","path":"imu_data/las_model.py","file_url":"https://github.com/aakashrkaku/seq2seq_hrar/blob/HEAD/imu_data/las_model.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"GPL-3.0","inline_ok":false,"mcp_get_code":{"code_sha256":"847ad682c7e34f92"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}