{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/segmental-spatiotemporal-cnns-for-fine","title":"Segmental Spatiotemporal CNNs for Fine-grained Action Segmentation","arxiv_id":"1602.02995","date":"2016-02-09","proceeding":null,"authors":["Colin Lea","Austin Reiter","Rene Vidal","Gregory D. Hager"],"abstract":"Joint segmentation and classification of fine-grained actions is important\nfor applications of human-robot interaction, video surveillance, and human\nskill evaluation. However, despite substantial recent progress in large-scale\naction classification, the performance of state-of-the-art fine-grained action\nrecognition approaches remains low. We propose a model for action segmentation\nwhich combines low-level spatiotemporal features with a high-level segmental\nclassifier. Our spatiotemporal CNN is comprised of a spatial component that\nuses convolutional filters to capture information about objects and their\nrelationships, and a temporal component that uses large 1D convolutional\nfilters to capture information about how object relationships change across\ntime. These features are used in tandem with a semi-Markov model that models\ntransitions from one action to another. We introduce an efficient constrained\nsegmental inference algorithm for this model that is orders of magnitude faster\nthan the current approach. We highlight the effectiveness of our Segmental\nSpatiotemporal CNN on cooking and surgical action datasets for which we observe\nsubstantially improved performance relative to recent baseline methods.","url_abs":"http://arxiv.org/abs/1602.02995v4","url_pdf":"http://arxiv.org/pdf/1602.02995v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-classification","task_name":"Action Classification"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"action-segmentation","task_name":"Action Segmentation"},{"task_slug":"fine-grained-action-recognition","task_name":"Fine-grained Action Recognition"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-segmentation-on-gtea-1","task":"Action Segmentation","dataset":"GTEA","model":"ST-CNN","rank_in_archive_order":28,"of":28,"metrics":{"Acc":"60.6","Edit":"-","F1@10%":"58.7","F1@25%":"54.4","F1@50%":"41.9"},"uses_additional_data":false},{"leaderboard":"/sota/action-segmentation-on-jigsaws","task":"Action Segmentation","dataset":"JIGSAWS","model":"ST-CNN+Seg","rank_in_archive_order":7,"of":7,"metrics":{"Accuracy":"74.22","Edit Distance":"66.56"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1602.02995","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}