{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/mad-a-scalable-dataset-for-language-grounding","title":"MAD: A Scalable Dataset for Language Grounding in Videos from Movie Audio Descriptions","arxiv_id":"2112.00431","date":"2021-12-01","proceeding":"CVPR 2022 1","authors":["Mattia Soldan","Alejandro Pardo","Juan León Alcázar","Fabian Caba Heilbron","Chen Zhao","Silvio Giancola","Bernard Ghanem"],"abstract":"The recent and increasing interest in video-language research has driven the development of large-scale datasets that enable data-intensive machine learning techniques. In comparison, limited effort has been made at assessing the fitness of these datasets for the video-language grounding task. Recent works have begun to discover significant limitations in these datasets, suggesting that state-of-the-art techniques commonly overfit to hidden dataset biases. In this work, we present MAD (Movie Audio Descriptions), a novel benchmark that departs from the paradigm of augmenting existing video datasets with text annotations and focuses on crawling and aligning available audio descriptions of mainstream movies. MAD contains over 384,000 natural language sentences grounded in over 1,200 hours of videos and exhibits a significant reduction in the currently diagnosed biases for video-language grounding datasets. MAD's collection strategy enables a novel and more challenging version of video-language grounding, where short temporal moments (typically seconds long) must be accurately grounded in diverse long-form videos that can last up to three hours. We have released MAD's data and baselines code at https://github.com/Soldelli/MAD.","url_abs":"https://arxiv.org/abs/2112.00431v2","url_pdf":"https://arxiv.org/pdf/2112.00431v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"mad-a-scalable-dataset-for-language-grounding","repo_url":"https://github.com/Soldelli/MAD","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"moment-retrieval","task_name":"Moment Retrieval"},{"task_slug":"natural-language-moment-retrieval","task_name":"Natural Language Moment Retrieval"}],"methods":[{"method_slug":"clip","method_name":"CLIP"},{"method_slug":"vlg-net","method_name":"VLG-Net"}],"datasets_introduced":[{"slug":"mad","name":"MAD","full_name":""}],"methods_introduced":[],"results":[{"leaderboard":"/sota/natural-language-moment-retrieval-on-mad","task":"Natural Language Moment Retrieval","dataset":"MAD","model":"CLIP","rank_in_archive_order":5,"of":8,"metrics":{"R@1,IoU=0.1":"6.57","R@1,IoU=0.3":"3.13","R@1,IoU=0.5":"1.39","R@10,IoU=0.1":"20.26","R@10,IoU=0.3":"14.13","R@10,IoU=0.5":"8.38","R@100,IoU=0.1":"47.73","R@100,IoU=0.3":"36.98","R@100,IoU=0.5":"24.99","R@5,IoU=0.1":"15.05","R@5,IoU=0.3":"9.85","R@5,IoU=0.5":"5.44","R@50,IoU=0.1":"37.92","R@50,IoU=0.3":"28.71","R@50,IoU=0.5":"18.80"},"uses_additional_data":false},{"leaderboard":"/sota/natural-language-moment-retrieval-on-mad","task":"Natural Language Moment Retrieval","dataset":"MAD","model":"VLG-Net","rank_in_archive_order":7,"of":8,"metrics":{"R@1,IoU=0.1":"3.50","R@1,IoU=0.3":"2.63","R@1,IoU=0.5":"1.61","R@10,IoU=0.1":"18.32","R@10,IoU=0.3":"15.2","R@10,IoU=0.5":"10.18","R@100,IoU=0.1":"49.65","R@100,IoU=0.3":"43.95","R@100,IoU=0.5":"34.18","R@5,IoU=0.1":"11.74","R@5,IoU=0.3":"9.49","R@5,IoU=0.5":"6.23","R@50,IoU=0.1":"38.41","R@50,IoU=0.3":"33.68","R@50,IoU=0.5":"25.33"},"uses_additional_data":false},{"leaderboard":"/sota/natural-language-moment-retrieval-on-mad","task":"Natural Language Moment Retrieval","dataset":"MAD","model":"Random Chance","rank_in_archive_order":8,"of":8,"metrics":{"R@1,IoU=0.1":"0.09","R@1,IoU=0.3":"0.04","R@1,IoU=0.5":"0.01","R@10,IoU=0.1":"0.88","R@10,IoU=0.3":"0.39","R@10,IoU=0.5":"0.14","R@100,IoU=0.1":"8.47","R@100,IoU=0.3":"3.80","R@100,IoU=0.5":"1.40","R@5,IoU=0.1":"0.44","R@5,IoU=0.3":"0.19","R@5,IoU=0.5":"0.07","R@50,IoU=0.1":"4.33","R@50,IoU=0.3":"1.92","R@50,IoU=0.5":"0.71"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2112.00431","atlas_url":"https://app.syntology.ai/?focus=2112.00431","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2112.00431"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Soldelli/MAD","reach":null}],"summary":{"ran_draft_wrong":1},"by_repo_kind":{},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":1,"samples":[{"code_sha256_prefix":"7a734d85428d248e","entry":"load_pretrained_graph_weights","repo":null,"repo_kind":null,"path":null,"file_url":null,"link_basis":"identical_code_first_harvested_elsewhere","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":null,"inline_ok":false,"mcp_get_code":{"code_sha256":"7a734d85428d248e"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}