{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unsupervised-learning-of-action-classes-with","title":"Unsupervised learning of action classes with continuous temporal embedding","arxiv_id":"1904.04189","date":"2019-04-08","proceeding":"CVPR 2019 6","authors":["Anna Kukleva","Hilde Kuehne","Fadime Sener","Juergen Gall"],"abstract":"The task of temporally detecting and segmenting actions in untrimmed videos\nhas seen an increased attention recently. One problem in this context arises\nfrom the need to define and label action boundaries to create annotations for\ntraining which is very time and cost intensive. To address this issue, we\npropose an unsupervised approach for learning action classes from untrimmed\nvideo sequences. To this end, we use a continuous temporal embedding of\nframewise features to benefit from the sequential nature of activities. Based\non the latent space created by the embedding, we identify clusters of temporal\nsegments across all videos that correspond to semantic meaningful action\nclasses. The approach is evaluated on three challenging datasets, namely the\nBreakfast dataset, YouTube Instructions, and the 50Salads dataset. While\nprevious works assumed that the videos contain the same high level activity, we\nfurthermore show that the proposed approach can also be applied to a more\ngeneral setting where the content of the videos is unknown.","url_abs":"http://arxiv.org/abs/1904.04189v1","url_pdf":"http://arxiv.org/pdf/1904.04189v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"unsupervised-learning-of-action-classes-with","repo_url":"https://github.com/annusha/unsup_temp_embed","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"unsupervised-learning-of-action-classes-with","repo_url":"https://github.com/frans-db/progress-prediction","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"action-segmentation","task_name":"Action Segmentation"},{"task_slug":"unsupervised-action-segmentation","task_name":"Unsupervised Action Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/unsupervised-action-segmentation-on-breakfast","task":"Unsupervised Action Segmentation","dataset":"Breakfast","model":"CTE","rank_in_archive_order":7,"of":8,"metrics":{"Acc":"41.8","F1":"26.4","JSD":"87.4","Precision":"25.8","Recall":"27.0"},"uses_additional_data":false},{"leaderboard":"/sota/unsupervised-action-segmentation-on-ikea-asm","task":"Unsupervised Action Segmentation","dataset":"IKEA ASM","model":"CTE","rank_in_archive_order":3,"of":5,"metrics":{"Accuracy":"23.1","F1":"22.6","JSD":"73.7","Precision":"28.1","Recall":"18.9"},"uses_additional_data":false},{"leaderboard":"/sota/unsupervised-action-segmentation-on-youtube","task":"Unsupervised Action Segmentation","dataset":"Youtube INRIA Instructional","model":"CTE","rank_in_archive_order":8,"of":8,"metrics":{"Acc":"39","F1":"28.3","Precision":"39.3","Recall":"22.1"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1904.04189","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}