{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/vidchapters-7m-video-chapters-at-scale","title":"VidChapters-7M: Video Chapters at Scale","arxiv_id":"2309.13952","date":"2023-09-25","proceeding":"NeurIPS 2023 11","authors":["Antoine Yang","Arsha Nagrani","Ivan Laptev","Josef Sivic","Cordelia Schmid"],"abstract":"Segmenting long videos into chapters enables users to quickly navigate to the information of their interest. This important topic has been understudied due to the lack of publicly released datasets. To address this issue, we present VidChapters-7M, a dataset of 817K user-chaptered videos including 7M chapters in total. VidChapters-7M is automatically created from videos online in a scalable manner by scraping user-annotated chapters and hence without any additional manual annotation. We introduce the following three tasks based on this data. First, the video chapter generation task consists of temporally segmenting the video and generating a chapter title for each segment. To further dissect the problem, we also define two variants of this task: video chapter generation given ground-truth boundaries, which requires generating a chapter title given an annotated video segment, and video chapter grounding, which requires temporally localizing a chapter given its annotated title. We benchmark both simple baselines and state-of-the-art video-language models for these three tasks. We also show that pretraining on VidChapters-7M transfers well to dense video captioning tasks in both zero-shot and finetuning settings, largely improving the state of the art on the YouCook2 and ViTT benchmarks. Finally, our experiments reveal that downstream performance scales well with the size of the pretraining dataset. Our dataset, code, and models are publicly available at https://antoyang.github.io/vidchapters.html.","url_abs":"https://arxiv.org/abs/2309.13952v1","url_pdf":"https://arxiv.org/pdf/2309.13952v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"vidchapters-7m-video-chapters-at-scale","repo_url":"https://github.com/antoyang/VidChapters","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"dense-video-captioning","task_name":"Dense Video Captioning"},{"task_slug":"navigate","task_name":"Navigate"},{"task_slug":"video-captioning","task_name":"Video Captioning"},{"task_slug":"video-chaptering","task_name":"Video Chaptering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/video-chaptering-on-vidchapters-7m","task":"Video Chaptering","dataset":"VidChapters-7M","model":"Vid2Seq","rank_in_archive_order":2,"of":2,"metrics":{"CIDEr":"55.7","P@0.5":"43.1","P@0.7":"26.4","P@3s":"24.0","P@5s":"30.3","R@0.5":"48.2","R@0.7":"28.5","R@3s":"28.5","R@5s":"36.4","SODA":"0.114"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2309.13952","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2309.13952"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/antoyang/VidChapters","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"unverified":9},"by_repo_kind":{"listed":{"samples":9,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"c057f6996d2a67d6","entry":"custom_collate_fn","repo":"antoyang/VidChapters","repo_kind":"listed","path":"zs_speechvcg.py","file_url":"https://github.com/antoyang/VidChapters/blob/HEAD/zs_speechvcg.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"c057f6996d2a67d6"}},{"code_sha256_prefix":"3729c17b4ed1c6a4","entry":"custom_collate_fn","repo":"antoyang/VidChapters","repo_kind":"listed","path":"zs_visualvcg.py","file_url":"https://github.com/antoyang/VidChapters/blob/HEAD/zs_visualvcg.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"3729c17b4ed1c6a4"}},{"code_sha256_prefix":"981f35dd0fe6edf3","entry":"evaluate_detection","repo":"antoyang/VidChapters","repo_kind":"listed","path":"zs_vcgr.py","file_url":"https://github.com/antoyang/VidChapters/blob/HEAD/zs_vcgr.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"981f35dd0fe6edf3"}},{"code_sha256_prefix":"3c58dab2a68d13df","entry":"evaluate_navigation","repo":"antoyang/VidChapters","repo_kind":"listed","path":"zs_vcgr.py","file_url":"https://github.com/antoyang/VidChapters/blob/HEAD/zs_vcgr.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"3c58dab2a68d13df"}},{"code_sha256_prefix":"7a6ba965c7ce56ef","entry":"extract_boundaries_from_ffprobe_output","repo":"antoyang/VidChapters","repo_kind":"listed","path":"zs_visualvcg.py","file_url":"https://github.com/antoyang/VidChapters/blob/HEAD/zs_visualvcg.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"7a6ba965c7ce56ef"}},{"code_sha256_prefix":"90aac3e2130f1dba","entry":"extract_shots_with_ffprobe","repo":"antoyang/VidChapters","repo_kind":"listed","path":"zs_visualvcg.py","file_url":"https://github.com/antoyang/VidChapters/blob/HEAD/zs_visualvcg.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"90aac3e2130f1dba"}},{"code_sha256_prefix":"23951a2f859364ee","entry":"iou","repo":"antoyang/VidChapters","repo_kind":"listed","path":"zs_vcgr.py","file_url":"https://github.com/antoyang/VidChapters/blob/HEAD/zs_vcgr.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"23951a2f859364ee"}},{"code_sha256_prefix":"bfbc7746415f3426","entry":"load_tf_weights_in_t5","repo":"antoyang/VidChapters","repo_kind":"listed","path":"model/modeling_t5.py","file_url":"https://github.com/antoyang/VidChapters/blob/HEAD/model/modeling_t5.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"bfbc7746415f3426"}},{"code_sha256_prefix":"fb51a172ea598d60","entry":"smooth","repo":"antoyang/VidChapters","repo_kind":"listed","path":"model/texttitling.py","file_url":"https://github.com/antoyang/VidChapters/blob/HEAD/model/texttitling.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"fb51a172ea598d60"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}