{"url":"/dataset/assembly101","name":"Assembly101","full_name":null,"description_markdown":"Assembly101 is a new procedural activity dataset featuring 4321 videos of people assembling and disassembling 101 \"take-apart\" toy vehicles. Participants work without fixed instructions, and the sequences feature rich and natural variations in action ordering, mistakes, and corrections. Assembly101 is the first multi-view action dataset, with simultaneous static (8) and egocentric (4) recordings. Sequences are annotated with more than 100K coarse and 1M fine-grained action segments, and 18M 3D hand poses. We benchmark on three action understanding tasks: recognition, anticipation and temporal segmentation. Additionally, we propose a novel task of detecting mistakes. The unique recording format and rich set of annotations allow us to investigate generalization to new toys, cross-view transfer, long-tailed distributions, and pose vs. appearance. We envision that Assembly101 will serve as a new challenge to investigate various activity understanding problems.\r\n\r\nImage Source: [https://assembly-101.github.io/](https://assembly-101.github.io/)","description_withheld":null,"homepage":"https://assembly-101.github.io/","introduced_date":"2022-03-28","introduced_date_note":null,"introduced_by":{"paper":"/paper/assembly101-a-large-scale-multi-view-video","title":"Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural Activities","first_author":"Fadime Sener","url":null},"license":null,"modalities":[{"name":"Videos","url":"/datasets/modality/videos"}],"tasks":[{"name":"Action Recognition","url":"/task/action-recognition-in-videos","datasets_with_task":"/datasets/task/action-recognition-in-videos"},{"name":"3D Action Recognition","url":"/task/3d-human-action-recognition","datasets_with_task":"/datasets/task/3d-human-action-recognition"},{"name":"Action Segmentation","url":"/task/action-segmentation","datasets_with_task":"/datasets/task/action-segmentation"},{"name":"Action Anticipation","url":"/task/action-anticipation","datasets_with_task":"/datasets/task/action-anticipation"},{"name":"Open Vocabulary Action Recognition","url":"/task/open-vocabulary-action-recognition","datasets_with_task":"/datasets/task/open-vocabulary-action-recognition"},{"name":"Mistake Detection","url":"/task/mistake-detection","datasets_with_task":"/datasets/task/mistake-detection"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["Assembly101"],"data_loaders":[{"repo":"https://github.com/assembly-101/assembly101-download-scripts","url":"https://github.com/assembly-101/assembly101-download-scripts","frameworks":["pytorch"]}],"num_papers_in_archive":57,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/3d-action-recognition-on-assembly101","task":"3D Action Recognition","dataset_variant":"Assembly101","rows":7,"metrics":["Actions Top-1","Verbs Top-1","Object Top-1"],"first_row_in_archive_order":{"model":"HandFormer-B/21","paper":"/paper/on-the-utility-of-3d-hand-poses-for-action","metrics":{"Actions Top-1":"41.06","Object Top-1":"51.17","Verbs Top-1":"69.23"},"code_links":[{"title":"s-shamil/HandFormer","url":"https://github.com/s-shamil/HandFormer"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/action-segmentation-on-assembly101","task":"Action Segmentation","dataset_variant":"Assembly101","rows":7,"metrics":["F1@10%","F1@25%","F1@50%","Edit","MoF"],"first_row_in_archive_order":{"model":"ASQuery","paper":"/paper/asquery-a-query-based-model-for-action","metrics":{"Edit":"35.3","F1@10%":"37.8","F1@25%":"35.6","F1@50%":"29.4","MoF":"40.4"},"code_links":[{"title":"zlngan/ASQuery","url":"https://github.com/zlngan/ASQuery"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/action-anticipation-on-assembly101","task":"Action Anticipation","dataset_variant":"Assembly101","rows":2,"metrics":["Verbs Recall@5","Objects Recall@5","Actions Recall@5"],"first_row_in_archive_order":{"model":"Goal Consistency","paper":"/paper/action-anticipation-with-goal-consistency","metrics":{"Actions Recall@5":"12.07","Objects Recall@5":"28.38","Verbs Recall@5":"60.04"},"code_links":[{"title":"olga-zats/goal_consistency","url":"https://github.com/olga-zats/goal_consistency"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/open-vocabulary-action-recognition-on","task":"Open Vocabulary Action Recognition","dataset_variant":"Assembly101","rows":1,"metrics":["HM"],"first_row_in_archive_order":{"model":"OAP+AOP","paper":"/paper/opening-the-vocabulary-of-egocentric-actions-1","metrics":{"HM":"6.6"},"code_links":[{"title":"dibschat/openvocab-egoAR","url":"https://github.com/dibschat/openvocab-egoAR"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/chase-learning-convex-hull-adaptive-shift-for","title":"CHASE: Learning Convex Hull Adaptive Shift for Skeleton-based Multi-Entity Action Recognition","date":"2024-10-09","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":2,"samples_ran":2,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/asquery-a-query-based-model-for-action","title":"ASQuery: A Query-based Model for Action Segmentation","date":"2024-09-30","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/on-the-utility-of-3d-hand-poses-for-action","title":"On the Utility of 3D Hand Poses for Action Recognition","date":"2024-03-14","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":13,"samples_ran":8,"samples_unverified":5,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/progress-aware-online-action-segmentation-for","title":"Progress-Aware Online Action Segmentation for Egocentric Procedural Task Videos","date":"2024-01-01","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/opening-the-vocabulary-of-egocentric-actions-1","title":"Opening the Vocabulary of Egocentric Actions","date":"2023-08-22","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/how-much-temporal-long-term-context-is-needed","title":"How Much Temporal Long-Term Context is Needed for Action Segmentation?","date":"2023-08-22","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":6,"samples_ran":6,"samples_unverified":0,"pointer_only_for_licence":6,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/interactive-spatiotemporal-token-attention","title":"Interactive Spatiotemporal Token Attention Network for Skeleton-based General Interactive Action Recognition","date":"2023-07-14","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/action-anticipation-with-goal-consistency","title":"Action Anticipation with Goal Consistency","date":"2023-06-26","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/unified-fully-and-timestamp-supervised","title":"Unified Fully and Timestamp Supervised Temporal Action Segmentation via Sequence to Sequence Translation","date":"2022-09-01","rows_on_this_dataset":1,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":2,"samples_ran":2,"samples_unverified":0,"pointer_only_for_licence":2,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/asformer-transformer-for-action-segmentation","title":"ASFormer: Transformer for Action Segmentation","date":"2021-10-16","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/coarse-to-fine-multi-resolution-temporal","title":"Coarse to Fine Multi-Resolution Temporal Convolutional Network","date":"2021-05-23","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/revisiting-skeleton-based-action-recognition","title":"Revisiting Skeleton-based Action Recognition","date":"2021-04-28","rows_on_this_dataset":1,"code_links":4,"syntology":null},{"paper":"/paper/ms-tcn-multi-stage-temporal-convolutional-2","title":"MS-TCN++: Multi-Stage Temporal Convolutional Network for Action Segmentation","date":"2020-06-16","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/temporal-aggregate-representations-for-long","title":"Temporal Aggregate Representations for Long-Range Video Understanding","date":"2020-06-01","rows_on_this_dataset":1,"code_links":2,"syntology":null},{"paper":"/paper/disentangling-and-unifying-graph-convolutions","title":"Disentangling and Unifying Graph Convolutions for Skeleton-Based Action Recognition","date":"2020-03-31","rows_on_this_dataset":1,"code_links":3,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":3,"samples_ran":3,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/temporal-shift-module-for-efficient-video","title":"TSM: Temporal Shift Module for Efficient Video Understanding","date":"2018-11-20","rows_on_this_dataset":1,"code_links":13,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":16,"samples_ran":6,"samples_unverified":10,"pointer_only_for_licence":4,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/non-local-graph-convolutional-networks-for","title":"Two-Stream Adaptive Graph Convolutional Networks for Skeleton-Based Action Recognition","date":"2018-05-20","rows_on_this_dataset":1,"code_links":4,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":8,"samples_harvested":44,"samples_ran":29,"samples_unverified":15,"pointer_only_for_licence":12,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}