{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/localizing-moments-in-video-with-natural","title":"Localizing Moments in Video with Natural Language","arxiv_id":"1708.01641","date":"2017-08-04","proceeding":"ICCV 2017 10","authors":["Lisa Anne Hendricks","Oliver Wang","Eli Shechtman","Josef Sivic","Trevor Darrell","Bryan Russell"],"abstract":"We consider retrieving a specific temporal segment, or moment, from a video\ngiven a natural language text description. Methods designed to retrieve whole\nvideo clips with natural language determine what occurs in a video but not\nwhen. To address this issue, we propose the Moment Context Network (MCN) which\neffectively localizes natural language queries in videos by integrating local\nand global video features over time. A key obstacle to training our MCN model\nis that current video datasets do not include pairs of localized video segments\nand referring expressions, or text descriptions which uniquely identify a\ncorresponding moment. Therefore, we collect the Distinct Describable Moments\n(DiDeMo) dataset which consists of over 10,000 unedited, personal videos in\ndiverse visual settings with pairs of localized video segments and referring\nexpressions. We demonstrate that MCN outperforms several baseline methods and\nbelieve that our initial results together with the release of DiDeMo will\ninspire further research on localizing video moments with natural language.","url_abs":"http://arxiv.org/abs/1708.01641v1","url_pdf":"http://arxiv.org/pdf/1708.01641v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"localizing-moments-in-video-with-natural","repo_url":"https://github.com/lisaanne/localizingmoments","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"caffe2","reach":null},{"paper_slug":"localizing-moments-in-video-with-natural","repo_url":"https://github.com/mrsalehi/ground-sentence-video","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"natural-language-queries","task_name":"Natural Language Queries"}],"methods":[],"datasets_introduced":[{"slug":"didemo","name":"DiDeMo","full_name":"Distinct Describable Moments"}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1708.01641","atlas_url":"https://app.syntology.ai/?focus=1708.01641","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1708.01641"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/mrsalehi/ground-sentence-video","reach":{"status":"unanswered"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/lisaanne/localizingmoments","reach":null}],"summary":{"ran_draft_wrong":1},"by_repo_kind":{"listed":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":1,"samples":[{"code_sha256_prefix":"d315ef366b9b9bb6","entry":"add_dict_values","repo":"lisaanne/localizingmoments","repo_kind":"listed","path":"build_net.py","file_url":"https://github.com/lisaanne/localizingmoments/blob/HEAD/build_net.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"d315ef366b9b9bb6"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}