{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/assembly101-a-large-scale-multi-view-video","title":"Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural Activities","arxiv_id":"2203.14712","date":"2022-03-28","proceeding":"CVPR 2022 1","authors":["Fadime Sener","Dibyadip Chatterjee","Daniel Shelepov","Kun He","Dipika Singhania","Robert Wang","Angela Yao"],"abstract":"Assembly101 is a new procedural activity dataset featuring 4321 videos of people assembling and disassembling 101 \"take-apart\" toy vehicles. Participants work without fixed instructions, and the sequences feature rich and natural variations in action ordering, mistakes, and corrections. Assembly101 is the first multi-view action dataset, with simultaneous static (8) and egocentric (4) recordings. Sequences are annotated with more than 100K coarse and 1M fine-grained action segments, and 18M 3D hand poses. We benchmark on three action understanding tasks: recognition, anticipation and temporal segmentation. Additionally, we propose a novel task of detecting mistakes. The unique recording format and rich set of annotations allow us to investigate generalization to new toys, cross-view transfer, long-tailed distributions, and pose vs. appearance. We envision that Assembly101 will serve as a new challenge to investigate various activity understanding problems.","url_abs":"https://arxiv.org/abs/2203.14712v2","url_pdf":"https://arxiv.org/pdf/2203.14712v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"assembly101-a-large-scale-multi-view-video","repo_url":"https://github.com/assembly101/assembly101.github.io","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"3d-human-action-recognition","task_name":"3D Action Recognition"},{"task_slug":"action-anticipation","task_name":"Action Anticipation"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"action-segmentation","task_name":"Action Segmentation"},{"task_slug":"mistake-detection","task_name":"Mistake Detection"}],"methods":[],"datasets_introduced":[{"slug":"assembly101","name":"Assembly101","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2203.14712","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}