{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/temporally-consistent-unbalanced-optimal","title":"Temporally Consistent Unbalanced Optimal Transport for Unsupervised Action Segmentation","arxiv_id":"2404.01518","date":"2024-04-01","proceeding":"CVPR 2024 1","authors":["Ming Xu","Stephen Gould"],"abstract":"We propose a novel approach to the action segmentation task for long, untrimmed videos, based on solving an optimal transport problem. By encoding a temporal consistency prior into a Gromov-Wasserstein problem, we are able to decode a temporally consistent segmentation from a noisy affinity/matching cost matrix between video frames and action classes. Unlike previous approaches, our method does not require knowing the action order for a video to attain temporal consistency. Furthermore, our resulting (fused) Gromov-Wasserstein problem can be efficiently solved on GPUs using a few iterations of projected mirror descent. We demonstrate the effectiveness of our method in an unsupervised learning setting, where our method is used to generate pseudo-labels for self-training. We evaluate our segmentation approach and unsupervised learning pipeline on the Breakfast, 50-Salads, YouTube Instructions and Desktop Assembly datasets, yielding state-of-the-art results for the unsupervised video action segmentation task.","url_abs":"https://arxiv.org/abs/2404.01518v3","url_pdf":"https://arxiv.org/pdf/2404.01518v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"temporally-consistent-unbalanced-optimal","repo_url":"https://github.com/mingu6/action_seg_ot","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"action-segmentation","task_name":"Action Segmentation"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"unsupervised-action-segmentation","task_name":"Unsupervised Action Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/unsupervised-action-segmentation-on-breakfast","task":"Unsupervised Action Segmentation","dataset":"Breakfast","model":"ASOT","rank_in_archive_order":2,"of":8,"metrics":{"Acc":"56.1","F1":"38.3","JSD":"94.9","Precision":"36.7","Recall":"40.1","mIoU":"18.6"},"uses_additional_data":false},{"leaderboard":"/sota/unsupervised-action-segmentation-on-ikea-asm","task":"Unsupervised Action Segmentation","dataset":"IKEA ASM","model":"ASOT","rank_in_archive_order":2,"of":5,"metrics":{"Accuracy":"34.0","F1":"27.9","JSD":"88.7","Precision":"21.1","Recall":"24.0"},"uses_additional_data":false},{"leaderboard":"/sota/unsupervised-action-segmentation-on-youtube","task":"Unsupervised Action Segmentation","dataset":"Youtube INRIA Instructional","model":"ASOT","rank_in_archive_order":2,"of":8,"metrics":{"Acc":"52.9","F1":"35.1","Precision":"47.6","Recall":"27.8","mIoU":"24.7"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2404.01518","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}