{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/cinepile-a-long-video-question-answering","title":"CinePile: A Long Video Question Answering Dataset and Benchmark","arxiv_id":"2405.08813","date":"2024-05-14","proceeding":null,"authors":["Ruchit Rawal","Khalid Saifullah","Miquel Farré","Ronen Basri","David Jacobs","Gowthami Somepalli","Tom Goldstein"],"abstract":"Current datasets for long-form video understanding often fall short of providing genuine long-form comprehension challenges, as many tasks derived from these datasets can be successfully tackled by analyzing just one or a few random frames from a video. To address this issue, we present a novel dataset and benchmark, CinePile, specifically designed for authentic long-form video understanding. This paper details our innovative approach for creating a question-answer dataset, utilizing advanced LLMs with human-in-the-loop and building upon human-generated raw data. Our comprehensive dataset comprises 305,000 multiple-choice questions (MCQs), covering various visual and multimodal aspects, including temporal comprehension, understanding human-object interactions, and reasoning about events or actions within a scene. Additionally, we fine-tuned open-source Video-LLMs on the training split and evaluated both open-source and proprietary video-centric LLMs on the test split of our dataset. The findings indicate that although current models underperform compared to humans, fine-tuning these models can lead to significant improvements in their performance.","url_abs":"https://arxiv.org/abs/2405.08813v3","url_pdf":"https://arxiv.org/pdf/2405.08813v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"form","task_name":"Form"},{"task_slug":"human-object-interaction-detection","task_name":"Human-Object Interaction Detection"},{"task_slug":"multiple-choice","task_name":"Multiple-choice"},{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"video-question-answering","task_name":"Video Question Answering"},{"task_slug":"video-understanding","task_name":"Video Understanding"},{"task_slug":"zeroshot-video-question-answer","task_name":"Zero-Shot Video Question Answer"}],"methods":[],"datasets_introduced":[{"slug":"cinepile","name":"CinePile: A Long Video Question Answering Dataset and Benchmark","full_name":""}],"methods_introduced":[],"results":[{"leaderboard":"/sota/zero-shot-video-question-answer-on-cinepile-a","task":"Zero-Shot Video Question Answer","dataset":"CinePile: A Long Video Question Answering Dataset and Benchmark","model":"Human","rank_in_archive_order":1,"of":1,"metrics":{"Accuracy":"86"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2405.08813","atlas_url":"https://app.syntology.ai/?focus=2405.08813","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}