{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/coin-a-large-scale-dataset-for-comprehensive","title":"COIN: A Large-scale Dataset for Comprehensive Instructional Video Analysis","arxiv_id":"1903.02874","date":"2019-03-07","proceeding":"CVPR 2019 6","authors":["Yansong Tang","Dajun Ding","Yongming Rao","Yu Zheng","Danyang Zhang","Lili Zhao","Jiwen Lu","Jie zhou"],"abstract":"There are substantial instructional videos on the Internet, which enables us\nto acquire knowledge for completing various tasks. However, most existing\ndatasets for instructional video analysis have the limitations in diversity and\nscale,which makes them far from many real-world applications where more diverse\nactivities occur. Moreover, it still remains a great challenge to organize and\nharness such data. To address these problems, we introduce a large-scale\ndataset called \"COIN\" for COmprehensive INstructional video analysis. Organized\nwith a hierarchical structure, the COIN dataset contains 11,827 videos of 180\ntasks in 12 domains (e.g., vehicles, gadgets, etc.) related to our daily life.\nWith a new developed toolbox, all the videos are annotated effectively with a\nseries of step descriptions and the corresponding temporal boundaries.\nFurthermore, we propose a simple yet effective method to capture the\ndependencies among different steps, which can be easily plugged into\nconventional proposal-based action detection methods for localizing important\nsteps in instructional videos. In order to provide a benchmark for\ninstructional video analysis, we evaluate plenty of approaches on the COIN\ndataset under different evaluation criteria. We expect the introduction of the\nCOIN dataset will promote the future in-depth research on instructional video\nanalysis for the community.","url_abs":"http://arxiv.org/abs/1903.02874v1","url_pdf":"http://arxiv.org/pdf/1903.02874v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-detection","task_name":"Action Detection"}],"methods":[],"datasets_introduced":[{"slug":"coin","name":"COIN","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1903.02874","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}