{"url":"/dataset/video-mme","name":"Video-MME","full_name":null,"description_markdown":"Video-MME stands for **Video Multi-Modal Evaluation**. It is the first-ever comprehensive evaluation benchmark specifically designed for Multi-modal Large Language Models (MLLMs) in video analysis¹. This benchmark is significant because it addresses the need for a high-quality assessment of MLLMs' performance in processing sequential visual data, which has been less explored compared to their capabilities in static image understanding.\r\n\r\nThe Video-MME benchmark is characterized by its:\r\n1. **Diversity** in video types, covering 6 primary visual domains with 30 subfields for broad scenario generalizability.\r\n2. **Duration** in the temporal dimension, including short-, medium-, and long-term videos ranging from 11 seconds to 1 hour, to assess robust contextual dynamics.\r\n3. **Breadth** in data modalities, integrating multi-modal inputs such as video frames, subtitles, and audios.\r\n4. **Quality** in annotations, with rigorous manual labeling by expert annotators for precise and reliable model assessment¹.\r\n\r\nThe benchmark includes 900 videos totaling 256 hours, manually selected and annotated, resulting in 2,700 question-answer pairs. It has been used to evaluate various state-of-the-art MLLMs, including the GPT-4 series and Gemini 1.5 Pro, as well as open-source image and video models¹. The findings from Video-MME highlight the need for further improvements in handling longer sequences and multi-modal data, which is crucial for the advancement of MLLMs¹.\r\n\r\n(1) [2405.21075] Video-MME: The First-Ever Comprehensive Evaluation .... https://arxiv.org/abs/2405.21075.\r\n(2) Video-MME. https://video-mme.github.io/home_page.html.\r\n(3) Video-MME: Welcome. https://video-mme.github.io/.\r\n(4) undefined. https://doi.org/10.48550/arXiv.2405.21075.","description_withheld":null,"homepage":"https://video-mme.github.io","introduced_date":"2024-05-31","introduced_date_note":null,"introduced_by":{"paper":"/paper/video-mme-the-first-ever-comprehensive","title":"Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis","first_author":"Chaoyou Fu","url":null},"license":null,"modalities":[{"name":"Videos","url":"/datasets/modality/videos"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Zero-Shot Video Question Answer","url":"/task/zeroshot-video-question-answer","datasets_with_task":"/datasets/task/zeroshot-video-question-answer"}],"languages":[],"variants":["Video-MME","Video-MME (w/o subs)"],"data_loaders":[],"num_papers_in_archive":152,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/zero-shot-video-question-answer-on-video-mme-1","task":"Zero-Shot Video Question Answer","dataset_variant":"Video-MME","rows":11,"metrics":["Accuracy (%)"],"first_row_in_archive_order":{"model":"Gemini 1.5 Pro","paper":"/paper/gemini-1-5-unlocking-multimodal-understanding","metrics":{"Accuracy (%)":"81.3"},"code_links":[{"title":"dlvuldet/primevul","url":"https://github.com/dlvuldet/primevul"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/zero-shot-video-question-answer-on-video-mme","task":"Zero-Shot Video Question Answer","dataset_variant":"Video-MME (w/o subs)","rows":9,"metrics":["Accuracy (%)"],"first_row_in_archive_order":{"model":"Video-RAG (based on LLaVA-Video)","paper":"/paper/video-rag-visually-aligned-retrieval","metrics":{"Accuracy (%)":"77.4"},"code_links":[{"title":"leon1207/video-rag-master","url":"https://github.com/leon1207/video-rag-master"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/bimba-selective-scan-compression-for-long","title":"BIMBA: Selective-Scan Compression for Long-Range Video Question Answering","date":"2025-03-12","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/video-rag-visually-aligned-retrieval","title":"Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension","date":"2024-11-20","rows_on_this_dataset":2,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":1,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/timesuite-improving-mllms-for-long-video","title":"TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning","date":"2024-10-25","rows_on_this_dataset":2,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":6,"samples_ran":3,"samples_unverified":3,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/longvu-spatiotemporal-adaptive-compression","title":"LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding","date":"2024-10-22","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":11,"samples_ran":6,"samples_unverified":5,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/2408-01800","title":"MiniCPM-V: A GPT-4V Level MLLM on Your Phone","date":"2024-08-03","rows_on_this_dataset":1,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":14,"samples_ran":9,"samples_unverified":5,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/gpt-4o-visual-perception-performance-of","title":"GPT-4o: Visual perception performance of multimodal large language models in piglet activity understanding","date":"2024-06-14","rows_on_this_dataset":4,"code_links":0,"syntology":null},{"paper":"/paper/videollama-2-advancing-spatial-temporal","title":"VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs","date":"2024-06-11","rows_on_this_dataset":2,"code_links":3,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":17,"samples_ran":7,"samples_unverified":10,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/gemini-1-5-unlocking-multimodal-understanding","title":"Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context","date":"2024-03-08","rows_on_this_dataset":4,"code_links":1,"syntology":null},{"paper":"/paper/vila-on-pre-training-for-visual-language","title":"VILA: On Pre-training for Visual Language Models","date":"2023-12-12","rows_on_this_dataset":2,"code_links":3,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":5,"samples_harvested":49,"samples_ran":26,"samples_unverified":23,"pointer_only_for_licence":1,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}