{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/w2vv-fully-deep-learning-for-ad-hoc-video","title":"W2VV++: Fully Deep Learning for Ad-hoc Video Search","arxiv_id":null,"date":"2019-10-21","proceeding":"ACM Multimedia 2019 2019 10","authors":["Xirong Li; Chaoxi Xu; Gang Yang; Zhineng Chen; Jianfeng Dong"],"abstract":"Ad-hoc video search (AVS) is an important yet challenging problem in multimedia retrieval. Different from previous concept-based methods, we propose an end-to-end deep learning method for query representation learning. The proposed method requires no concept modeling, matching and selection. The backbone of our method is the proposed W2VV++ model, a super version of Word2VisualVec (W2VV) previously developed for visual-to-text matching. W2VV++ is obtained by tweaking W2VV with a better sentence encoding strategy and an improved triplet ranking loss. With these simple changes, W2VV++ brings in a substantial improvement in performance. As our participation in the TRECVID 2018 AVS task and retrospective experiments on the TRECVID 2016 and 2017 data show, our best single model, with an overall inferred average precision (infAP) of 0.157, outperforms the state-of-the-art. The performance can be further boosted by model ensemble using late average fusion, reaching a higher infAP of 0.163. With W2VV++, we establish a new baseline for ad-hoc video search.","url_abs":"https://dl.acm.org/doi/pdf/10.1145/3343031.3350906?download=true","url_pdf":"http://lixirong.net/pub/mm2019-w2vvpp.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"w2vv-fully-deep-learning-for-ad-hoc-video","repo_url":"https://github.com/li-xirong/w2vvpp","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"ad-hoc-video-search","task_name":"Ad-hoc video search"},{"task_slug":"deep-learning","task_name":"Deep Learning"},{"task_slug":"representation-learning","task_name":"Representation Learning"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"sentence","task_name":"Sentence"},{"task_slug":"text-matching","task_name":"Text Matching"},{"task_slug":null,"task_name":"Triplet"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/ad-hoc-video-search-on-trecvid-avs16-iacc-3","task":"Ad-hoc video search","dataset":"TRECVID-AVS16 (IACC.3)","model":"W2VV++","rank_in_archive_order":4,"of":4,"metrics":{"infAP":"0.151"},"uses_additional_data":true},{"leaderboard":"/sota/ad-hoc-video-search-on-trecvid-avs17-iacc-3","task":"Ad-hoc video search","dataset":"TRECVID-AVS17 (IACC.3)","model":"W2VV++","rank_in_archive_order":4,"of":4,"metrics":{"infAP":"0.220"},"uses_additional_data":true},{"leaderboard":"/sota/ad-hoc-video-search-on-trecvid-avs18-iacc-3","task":"Ad-hoc video search","dataset":"TRECVID-AVS18 (IACC.3)","model":"W2VV++","rank_in_archive_order":3,"of":4,"metrics":{"infAP":"0.121"},"uses_additional_data":true}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}