{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/maximization-and-restoration-action","title":"Maximization and restoration: Action segmentation through dilation passing and temporal reconstruction","arxiv_id":null,"date":"2022-05-02","proceeding":"Pattern Recognition 2022 5","authors":["Junyong Park","Daekyum Kim","Sejoon Huh","Sungho Jo"],"abstract":"Action segmentation aims to split videos into segments of different actions. Recent work focuses on dealing with long-range dependencies of long, untrimmed videos, but still suffers from over-segmentation and performance saturation due to increased model complexity. This paper addresses the aforementioned issues through a divide-and-conquer strategy that first maximizes the frame-wise classification accuracy of the model and then reduces the over-segmentation errors. This strategy is implemented with the Dilation Passing and Reconstruction Network, composed of the Dilation Passing Network, which primarily aims to increase accuracy by propagating information of different dilations, and the Temporal Reconstruction Network, which reduces over-segmentation errors by temporally encoding and decoding the output features from the Dilation Passing Network. We also propose a weighted temporal mean squared error loss that further reduces over-segmentation. Through evaluations on the 50Salads, GTEA, and Breakfast datasets, we show that our model achieves significant results compared to existing state-of-the-art models.","url_abs":"https://www.sciencedirect.com/science/article/pii/S003132032200245X","url_pdf":"http://nmail.kaist.ac.kr/paper/pr2022.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-segmentation","task_name":"Action Segmentation"},{"task_slug":"segmentation","task_name":"Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-segmentation-on-50-salads-1","task":"Action Segmentation","dataset":"50 Salads","model":"DPRN","rank_in_archive_order":11,"of":28,"metrics":{"Acc":"87.2","Edit":"82.0","F1@10%":"87.8","F1@25%":"86.3","F1@50%":"79.4"},"uses_additional_data":false},{"leaderboard":"/sota/action-segmentation-on-breakfast-1","task":"Action Segmentation","dataset":"Breakfast","model":"DPRN","rank_in_archive_order":14,"of":37,"metrics":{"Acc":"71.7","Average F1":"67.9","Edit":"75.1","F1@10%":"75.6","F1@25%":"70.5","F1@50%":"57.6"},"uses_additional_data":false},{"leaderboard":"/sota/action-segmentation-on-gtea-1","task":"Action Segmentation","dataset":"GTEA","model":"DPRN","rank_in_archive_order":7,"of":28,"metrics":{"Acc":"82.0","Edit":"90.9","F1@10%":"92.9","F1@25%":"92.0","F1@50%":"82.9"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}