{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/depthwise-separable-temporal-convolutional","title":"Depthwise Separable Temporal Convolutional Network for Action Segmentation","arxiv_id":null,"date":"2021-01-19","proceeding":"2020 International Conference on 3D Vision (3DV) 2021 1","authors":["Basavaraj Hampiholi","Christian Jarvers","Wolfgang Mader","Heiko Neumann"],"abstract":"Fine-grained temporal action segmentation in long,\r\nuntrimmed RGB videos is a key topic in visual human-\r\nmachine interaction. Recent temporal convolution based\r\napproaches either use encoder-decoder(ED) architecture or\r\ndilations with doubling factor in consecutive convolution\r\nlayers to segment actions in videos. However ED networks\r\noperate on low temporal resolution and the dilations in suc-\r\ncessive layers cause gridding artifacts problem. We propose\r\ndepthwise separable temporal convolution network (DS-\r\nTCN) that operates on full temporal resolution and with re-\r\nduced gridding effects. The basic component of DS-TCN\r\nis residual depthwise dilated block (RDDB). We explore the\r\ntrade-off between large kernels and small dilation rates us-\r\ning RDDB. We show that our DS-TCN is capable of captur-\r\ning long-term dependencies as well as local temporal cues\r\nefficiently. Our evaluation on three benchmark datasets,\r\nGTEA, 50Salads, and Breakfast demonstrates that DS-TCN\r\noutperforms the existing ED-TCN and dilation based TCN\r\nbaselines even with comparatively fewer parameters.","url_abs":"https://ieeexplore.ieee.org/document/9320343","url_pdf":"https://ieeexplore.ieee.org/document/9320343","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-segmentation","task_name":"Action Segmentation"},{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"temporal-action-segmentation","task_name":"Temporal Action Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-segmentation-on-50-salads-1","task":"Action Segmentation","dataset":"50 Salads","model":"DS-TCN","rank_in_archive_order":26,"of":28,"metrics":{"Acc":"80.0","Edit":"70.0","F1@10%":"77.0","F1@25%":"74.43","F1@50%":"65.78"},"uses_additional_data":false},{"leaderboard":"/sota/action-segmentation-on-breakfast-1","task":"Action Segmentation","dataset":"Breakfast","model":"DS-TCN","rank_in_archive_order":25,"of":37,"metrics":{"Acc":"70.75","Average F1":"59.6","Edit":"69.02","F1@10%":"67.70","F1@25%":"62.05","F1@50%":"49.18"},"uses_additional_data":false},{"leaderboard":"/sota/action-segmentation-on-gtea-1","task":"Action Segmentation","dataset":"GTEA","model":"DS-TCN","rank_in_archive_order":24,"of":28,"metrics":{"Acc":"78.10","Edit":"84.05","F1@10%":"88.30","F1@25%":"85.44","F1@50%":"72.84"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}