{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/verified-a-video-corpus-moment-retrieval","title":"VERIFIED: A Video Corpus Moment Retrieval Benchmark for Fine-Grained Video Understanding","arxiv_id":"2410.08593","date":"2024-10-11","proceeding":null,"authors":["Houlun Chen","Xin Wang","Hong Chen","Zeyang Zhang","Wei Feng","Bin Huang","Jia Jia","Wenwu Zhu"],"abstract":"Existing Video Corpus Moment Retrieval (VCMR) is limited to coarse-grained understanding, which hinders precise video moment localization when given fine-grained queries. In this paper, we propose a more challenging fine-grained VCMR benchmark requiring methods to localize the best-matched moment from the corpus with other partially matched candidates. To improve the dataset construction efficiency and guarantee high-quality data annotations, we propose VERIFIED, an automatic \\underline{V}id\\underline{E}o-text annotation pipeline to generate captions with \\underline{R}el\\underline{I}able \\underline{FI}n\\underline{E}-grained statics and \\underline{D}ynamics. Specifically, we resort to large language models (LLM) and large multimodal models (LMM) with our proposed Statics and Dynamics Enhanced Captioning modules to generate diverse fine-grained captions for each video. To filter out the inaccurate annotations caused by the LLM hallucination, we propose a Fine-Granularity Aware Noise Evaluator where we fine-tune a video foundation model with disturbed hard-negatives augmented contrastive and matching losses. With VERIFIED, we construct a more challenging fine-grained VCMR benchmark containing Charades-FIG, DiDeMo-FIG, and ActivityNet-FIG which demonstrate a high level of annotation quality. We evaluate several state-of-the-art VCMR models on the proposed dataset, revealing that there is still significant scope for fine-grained video understanding in VCMR. Code and Datasets are in \\href{https://github.com/hlchen23/VERIFIED}{https://github.com/hlchen23/VERIFIED}.","url_abs":"https://arxiv.org/abs/2410.08593v1","url_pdf":"https://arxiv.org/pdf/2410.08593v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"verified-a-video-corpus-moment-retrieval","repo_url":"https://github.com/hlchen23/verified","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"hallucination","task_name":"Hallucination"},{"task_slug":"moment-retrieval","task_name":"Moment Retrieval"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"video-corpus-moment-retrieval","task_name":"Video Corpus Moment Retrieval"},{"task_slug":"video-understanding","task_name":"Video Understanding"},{"task_slug":"text-annotation","task_name":"text annotation"}],"methods":[{"method_slug":"aware","method_name":"AWARE"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2410.08593","atlas_url":"https://app.syntology.ai/?focus=2410.08593","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}