{"url":"/task/video-grounding","name":"Video Grounding","slug":"video-grounding","description_markdown":"**Video grounding** is the task of linking spoken language descriptions to specific video segments. In video grounding, the model is given a video and a natural language description, such as a sentence or a caption, and its goal is to identify the specific segment of the video that corresponds to the description. This can involve tasks such as localizing the objects or actions mentioned in the description within the video, or associating a specific time interval with the description.","categories":[{"name":"Computer Vision","url":"/area/computer-vision"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":114,"papers_with_code":56,"benchmarks":2,"benchmark_tables_in_archive":2,"benchmark_tables_shown":2,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":10,"subtasks":2,"parent_tasks":1},"benchmarks":[{"leaderboard":"/sota/video-grounding-on-qvhighlights","slug":"video-grounding-on-qvhighlights","dataset":"QVHighlights","dataset_url":"/dataset/qvhighlights","rows_in_archive":7,"metrics":["R@1,IoU=0.7","R@1,IoU=0.5"],"first_row_in_archive_order":{"model":"InternVideo2-6B","paper_title":"InternVideo2: Scaling Foundation Models for Multimodal Video Understanding","paper_url":"/paper/internvideo2-scaling-video-foundation-models","paper_date":"2024-03-22","arxiv_id":"2403.15377","code_links":[{"title":"opengvlab/internvideo","url":"https://github.com/opengvlab/internvideo"},{"title":"opengvlab/internvideo2","url":"https://github.com/opengvlab/internvideo2"}],"syntology":null}},{"leaderboard":"/sota/video-grounding-on-mad","slug":"video-grounding-on-mad","dataset":"MAD","dataset_url":"/dataset/mad","rows_in_archive":2,"metrics":["R@1,IoU=0.1","R@5,IoU=0.1","R@10,IoU=0.1","R@100,IoU=0.1","R@50,IoU=0.1","R@1,IoU=0.3","R@5,IoU=0.3"],"first_row_in_archive_order":{"model":"DeCafNet","paper_title":"DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long Videos","paper_url":"/paper/decafnet-delegate-and-conquer-for-efficient","paper_date":"2025-05-22","arxiv_id":"2505.16376","code_links":[{"title":"zijialewislu/cvpr2025-decafnet","url":"https://github.com/zijialewislu/cvpr2025-decafnet"}],"syntology":null}}],"datasets":[{"url":"/dataset/kinetics","name":"Kinetics","full_name":"Kinetics Human Action Video Dataset","num_papers_in_archive":1341},{"url":"/dataset/qvhighlights","name":"QVHighlights","full_name":"Query-based Video Highlights","num_papers_in_archive":41},{"url":"/dataset/mad","name":"MAD","full_name":"","num_papers_in_archive":36},{"url":"/dataset/animal-kingdom","name":"Animal Kingdom","full_name":"","num_papers_in_archive":26},{"url":"/dataset/situated-reasoning-star","name":"STAR Benchmark","full_name":"Situated Reasoning","num_papers_in_archive":17},{"url":"/dataset/dttd2","name":"DTTD-Mobile","full_name":"","num_papers_in_archive":8},{"url":"/dataset/longvale","name":"LongVALE","full_name":"","num_papers_in_archive":2},{"url":"/dataset/kinetic-gebc","name":"Kinetics-GEB+","full_name":"","num_papers_in_archive":1},{"url":"/dataset/youwikihow","name":"YouwikiHow","full_name":"","num_papers_in_archive":1},{"url":"/dataset/vript","name":"Vript","full_name":"🎬 Vript: A Video Is Worth Thousands of Words","num_papers_in_archive":0}],"subtasks":[{"url":"/task/boundary-grounding","name":"Boundary Grounding"},{"url":"/task/video-narrative-grounding","name":"Video Narrative Grounding"}],"parent_tasks":[{"url":"/task/video-retrieval","name":"Video Retrieval"}],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":30,"of":56,"tagged_in_all":114,"items":[{"url":"/paper/umt-unified-multi-modal-transformers-for","title":"UMT: Unified Multi-modal Transformers for Joint Video Moment Retrieval and Highlight Detection","date":"2022-03-23","arxiv_id":"2203.12745","repositories_listed":3,"syntology":null},{"url":"/paper/internvideo2-scaling-video-foundation-models","title":"InternVideo2: Scaling Foundation Models for Multimodal Video Understanding","date":"2024-03-22","arxiv_id":"2403.15377","repositories_listed":2,"syntology":null},{"url":"/paper/context-guided-spatio-temporal-video","title":"Context-Guided Spatio-Temporal Video Grounding","date":"2024-01-03","arxiv_id":"2401.01578","repositories_listed":2,"syntology":{"n":34,"n_ran":21,"n_unverified":13,"n_pointer_only":34}},{"url":"/paper/negative-sample-matters-a-renaissance-of","title":"Negative Sample Matters: A Renaissance of Metric Learning for Temporal Grounding","date":"2021-09-10","arxiv_id":"2109.04872","repositories_listed":2,"syntology":null},{"url":"/paper/reinforcement-learning-tuning-for-videollms","title":"Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency","date":"2025-06-02","arxiv_id":"2506.01908","repositories_listed":1,"syntology":{"n":17,"n_ran":2,"n_unverified":15,"n_pointer_only":0}},{"url":"/paper/decafnet-delegate-and-conquer-for-efficient","title":"DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long Videos","date":"2025-05-22","arxiv_id":"2505.16376","repositories_listed":1,"syntology":null},{"url":"/paper/object-shot-enhanced-grounding-network-for","title":"Object-Shot Enhanced Grounding Network for Egocentric Video","date":"2025-05-07","arxiv_id":"2505.04270","repositories_listed":1,"syntology":{"n":17,"n_ran":4,"n_unverified":13,"n_pointer_only":0}},{"url":"/paper/timezero-temporal-video-grounding-with","title":"TimeZero: Temporal Video Grounding with Reasoning-Guided LVLM","date":"2025-03-17","arxiv_id":"2503.13377","repositories_listed":1,"syntology":{"n":7,"n_ran":3,"n_unverified":4,"n_pointer_only":0}},{"url":"/paper/omnistvg-toward-spatio-temporal-omni-object","title":"OmniSTVG: Toward Spatio-Temporal Omni-Object Video Grounding","date":"2025-03-13","arxiv_id":"2503.10500","repositories_listed":1,"syntology":null},{"url":"/paper/timeloc-a-unified-end-to-end-framework-for","title":"TimeLoc: A Unified End-to-End Framework for Precise Timestamp Localization in Long Videos","date":"2025-03-09","arxiv_id":"2503.06526","repositories_listed":1,"syntology":null},{"url":"/paper/knowing-your-target-target-aware-transformer","title":"Knowing Your Target: Target-Aware Transformer Makes Better Spatio-Temporal Video Grounding","date":"2025-02-16","arxiv_id":"2502.11168","repositories_listed":1,"syntology":null},{"url":"/paper/tarsier2-advancing-large-vision-language","title":"Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding","date":"2025-01-14","arxiv_id":"2501.07888","repositories_listed":1,"syntology":{"n":4,"n_ran":4,"n_unverified":0,"n_pointer_only":4}},{"url":"/paper/llava-st-a-multimodal-large-language-model","title":"LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding","date":"2025-01-14","arxiv_id":"2501.08282","repositories_listed":1,"syntology":{"n":7,"n_ran":3,"n_unverified":4,"n_pointer_only":0}},{"url":"/paper/vidchain-chain-of-tasks-with-metric-based","title":"VidChain: Chain-of-Tasks with Metric-based Direct Preference Optimization for Dense Video Captioning","date":"2025-01-12","arxiv_id":"2501.06761","repositories_listed":1,"syntology":null},{"url":"/paper/consistency-of-compositional-generalization","title":"Consistency of Compositional Generalization across Multiple Levels","date":"2024-12-18","arxiv_id":"2412.13636","repositories_listed":1,"syntology":null},{"url":"/paper/videollm-knows-when-to-speak-enhancing-time","title":"VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Format","date":"2024-11-27","arxiv_id":"2411.17991","repositories_listed":1,"syntology":{"n":2,"n_ran":0,"n_unverified":2,"n_pointer_only":0}},{"url":"/paper/prior-knowledge-integration-via-llm-encoding","title":"Prior Knowledge Integration via LLM Encoding and Pseudo Event Regulation for Video Moment Retrieval","date":"2024-07-21","arxiv_id":"2407.15051","repositories_listed":1,"syntology":{"n":11,"n_ran":5,"n_unverified":6,"n_pointer_only":0}},{"url":"/paper/artemis-towards-referential-understanding-in","title":"Artemis: Towards Referential Understanding in Complex Videos","date":"2024-06-01","arxiv_id":"2406.00258","repositories_listed":1,"syntology":null},{"url":"/paper/snag-scalable-and-accurate-video-grounding","title":"SnAG: Scalable and Accurate Video Grounding","date":"2024-04-02","arxiv_id":"2404.02257","repositories_listed":1,"syntology":{"n":20,"n_ran":5,"n_unverified":15,"n_pointer_only":5}},{"url":"/paper/unified-static-and-dynamic-network-efficient","title":"Unified Static and Dynamic Network: Efficient Temporal Filtering for Video Grounding","date":"2024-03-21","arxiv_id":"2403.14174","repositories_listed":1,"syntology":null},{"url":"/paper/hawkeye-training-video-text-llms-for","title":"HawkEye: Training Video-Text LLMs for Grounding Text in Videos","date":"2024-03-15","arxiv_id":"2403.10228","repositories_listed":1,"syntology":{"n":6,"n_ran":6,"n_unverified":0,"n_pointer_only":6}},{"url":"/paper/gaussian-mixture-proposals-with-pull-push","title":"Gaussian Mixture Proposals with Pull-Push Learning Scheme to Capture Diverse Events for Weakly Supervised Temporal Video Grounding","date":"2023-12-27","arxiv_id":"2312.16388","repositories_listed":1,"syntology":null},{"url":"/paper/cross-modal-contrastive-learning-with","title":"Cross-modal Contrastive Learning with Asymmetric Co-attention Network for Video Moment Retrieval","date":"2023-12-12","arxiv_id":"2312.07435","repositories_listed":1,"syntology":null},{"url":"/paper/grounded-question-answering-in-long","title":"Grounded Question-Answering in Long Egocentric Videos","date":"2023-12-11","arxiv_id":"2312.06505","repositories_listed":1,"syntology":{"n":6,"n_ran":5,"n_unverified":1,"n_pointer_only":0}},{"url":"/paper/vtimellm-empower-llm-to-grasp-video-moments","title":"VTimeLLM: Empower LLM to Grasp Video Moments","date":"2023-11-30","arxiv_id":"2311.18445","repositories_listed":1,"syntology":{"n":11,"n_ran":5,"n_unverified":6,"n_pointer_only":11}},{"url":"/paper/bridging-the-gap-a-unified-video","title":"Bridging the Gap: A Unified Video Comprehension Framework for Moment Retrieval and Highlight Detection","date":"2023-11-28","arxiv_id":"2311.16464","repositories_listed":1,"syntology":{"n":12,"n_ran":9,"n_unverified":3,"n_pointer_only":0}},{"url":"/paper/pg-video-llava-pixel-grounding-large-video","title":"PG-Video-LLaVA: Pixel Grounding Large Video-Language Models","date":"2023-11-22","arxiv_id":"2311.13435","repositories_listed":1,"syntology":{"n":5,"n_ran":1,"n_unverified":4,"n_pointer_only":5}},{"url":"/paper/exploring-iterative-refinement-with-diffusion","title":"Exploring Iterative Refinement with Diffusion Models for Video Grounding","date":"2023-10-26","arxiv_id":"2310.17189","repositories_listed":1,"syntology":null},{"url":"/paper/dual-path-temporal-map-optimization-for-make","title":"Dual-Path Temporal Map Optimization for Make-up Temporal Video Grounding","date":"2023-09-12","arxiv_id":"2309.06176","repositories_listed":1,"syntology":null},{"url":"/paper/can-i-trust-your-answer-visually-grounded","title":"Can I Trust Your Answer? Visually Grounded Video Question Answering","date":"2023-09-04","arxiv_id":"2309.01327","repositories_listed":1,"syntology":{"n":7,"n_ran":5,"n_unverified":2,"n_pointer_only":0}}],"syntology_records":15,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-25T09:33:49+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}