{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/lvd-2m-a-long-take-video-dataset-with","title":"LVD-2M: A Long-take Video Dataset with Temporally Dense Captions","arxiv_id":"2410.10816","date":"2024-10-14","proceeding":null,"authors":["Tianwei Xiong","Yuqing Wang","Daquan Zhou","Zhijie Lin","Jiashi Feng","Xihui Liu"],"abstract":"The efficacy of video generation models heavily depends on the quality of their training datasets. Most previous video generation models are trained on short video clips, while recently there has been increasing interest in training long video generation models directly on longer videos. However, the lack of such high-quality long videos impedes the advancement of long video generation. To promote research in long video generation, we desire a new dataset with four key features essential for training long video generation models: (1) long videos covering at least 10 seconds, (2) long-take videos without cuts, (3) large motion and diverse contents, and (4) temporally dense captions. To achieve this, we introduce a new pipeline for selecting high-quality long-take videos and generating temporally dense captions. Specifically, we define a set of metrics to quantitatively assess video quality including scene cuts, dynamic degrees, and semantic-level quality, enabling us to filter high-quality long-take videos from a large amount of source videos. Subsequently, we develop a hierarchical video captioning pipeline to annotate long videos with temporally-dense captions. With this pipeline, we curate the first long-take video dataset, LVD-2M, comprising 2 million long-take videos, each covering more than 10 seconds and annotated with temporally dense captions. We further validate the effectiveness of LVD-2M by fine-tuning video generation models to generate long videos with dynamic motions. We believe our work will significantly contribute to future research in long video generation.","url_abs":"https://arxiv.org/abs/2410.10816v1","url_pdf":"https://arxiv.org/pdf/2410.10816v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"lvd-2m-a-long-take-video-dataset-with","repo_url":"https://github.com/silentview/lvd-2m","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"video-captioning","task_name":"Video Captioning"},{"task_slug":"video-generation","task_name":"Video Generation"}],"methods":[{"method_slug":"set","method_name":"SET"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2410.10816","atlas_url":"https://app.syntology.ai/?focus=2410.10816","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2410.10816"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/SilentView/LVD-2M","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/silentview/lvd-2m","reach":{"status":"ok"}}],"summary":{"ran":5,"ran_honours":1,"ran_draft_wrong":1},"by_repo_kind":{"official":{"samples":7,"ran":7,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":7,"samples":[{"code_sha256_prefix":"fd8738bf3f18e932","entry":"cache","repo":"silentview/lvd-2m","repo_kind":"official","path":"pytube/helpers.py","file_url":"https://github.com/silentview/lvd-2m/blob/HEAD/pytube/helpers.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"fd8738bf3f18e932"}},{"code_sha256_prefix":"ed810f1befb0a75c","entry":"cal_span_time","repo":"SilentView/LVD-2M","repo_kind":"official","path":"download_videos_release.py","file_url":"https://github.com/SilentView/LVD-2M/blob/HEAD/download_videos_release.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"ed810f1befb0a75c"}},{"code_sha256_prefix":"288cd7ed350c4225","entry":"get_format_profile","repo":"silentview/lvd-2m","repo_kind":"official","path":"pytube/itags.py","file_url":"https://github.com/silentview/lvd-2m/blob/HEAD/pytube/itags.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"288cd7ed350c4225"}},{"code_sha256_prefix":"bbc530b455c98aa3","entry":"is_private","repo":"silentview/lvd-2m","repo_kind":"official","path":"pytube/extract.py","file_url":"https://github.com/silentview/lvd-2m/blob/HEAD/pytube/extract.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"bbc530b455c98aa3"}},{"code_sha256_prefix":"177afae9e319dff7","entry":"parse_time","repo":"SilentView/LVD-2M","repo_kind":"official","path":"download_videos_release.py","file_url":"https://github.com/SilentView/LVD-2M/blob/HEAD/download_videos_release.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"177afae9e319dff7"}},{"code_sha256_prefix":"90dfd7518fe80b47","entry":"recording_available","repo":"silentview/lvd-2m","repo_kind":"official","path":"pytube/extract.py","file_url":"https://github.com/silentview/lvd-2m/blob/HEAD/pytube/extract.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"90dfd7518fe80b47"}},{"code_sha256_prefix":"65bcff147b37d623","entry":"safe_filename","repo":"silentview/lvd-2m","repo_kind":"official","path":"pytube/helpers.py","file_url":"https://github.com/silentview/lvd-2m/blob/HEAD/pytube/helpers.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"65bcff147b37d623"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}