{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/vrag-region-attention-graphs-for-content","title":"VRAG: Region Attention Graphs for Content-Based Video Retrieval","arxiv_id":"2205.09068","date":"2022-05-18","proceeding":null,"authors":["Kennard Ng","Ser-Nam Lim","Gim Hee Lee"],"abstract":"Content-based Video Retrieval (CBVR) is used on media-sharing platforms for applications such as video recommendation and filtering. To manage databases that scale to billions of videos, video-level approaches that use fixed-size embeddings are preferred due to their efficiency. In this paper, we introduce Video Region Attention Graph Networks (VRAG) that improves the state-of-the-art of video-level methods. We represent videos at a finer granularity via region-level features and encode video spatio-temporal dynamics through region-level relations. Our VRAG captures the relationships between regions based on their semantic content via self-attention and the permutation invariant aggregation of Graph Convolution. In addition, we show that the performance gap between video-level and frame-level methods can be reduced by segmenting videos into shots and using shot embeddings for video retrieval. We evaluate our VRAG over several video retrieval tasks and achieve a new state-of-the-art for video-level retrieval. Furthermore, our shot-level VRAG shows higher retrieval precision than other existing video-level methods, and closer performance to frame-level methods at faster evaluation speeds. Finally, our code will be made publicly available.","url_abs":"https://arxiv.org/abs/2205.09068v1","url_pdf":"https://arxiv.org/pdf/2205.09068v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"video-retrieval","task_name":"Video Retrieval"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/video-retrieval-on-fivr-200k","task":"Video Retrieval","dataset":"FIVR-200K","model":"VRAG (CS)","rank_in_archive_order":13,"of":17,"metrics":{"mAP (CSVR)":"0.678","mAP (DSVR)":"0.723","mAP (ISVR)":"0.554"},"uses_additional_data":false},{"leaderboard":"/sota/video-retrieval-on-fivr-200k","task":"Video Retrieval","dataset":"FIVR-200K","model":"VRAG (video)","rank_in_archive_order":16,"of":17,"metrics":{"mAP (CSVR)":"0.470","mAP (DSVR)":"0.484","mAP (ISVR)":"0.399"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2205.09068","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}