{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/multi-modal-transformers-excel-at-class","title":"Class-agnostic Object Detection with Multi-modal Transformer","arxiv_id":"2111.11430","date":"2021-11-22","proceeding":null,"authors":["Muhammad Maaz","Hanoona Rasheed","Salman Khan","Fahad Shahbaz Khan","Rao Muhammad Anwer","Ming-Hsuan Yang"],"abstract":"What constitutes an object? This has been a long-standing question in computer vision. Towards this goal, numerous learning-free and learning-based approaches have been developed to score objectness. However, they generally do not scale well across new domains and novel objects. In this paper, we advocate that existing methods lack a top-down supervision signal governed by human-understandable semantics. For the first time in literature, we demonstrate that Multi-modal Vision Transformers (MViT) trained with aligned image-text pairs can effectively bridge this gap. Our extensive experiments across various domains and novel objects show the state-of-the-art performance of MViTs to localize generic objects in images. Based on the observation that existing MViTs do not include multi-scale feature processing and usually require longer training schedules, we develop an efficient MViT architecture using multi-scale deformable attention and late vision-language fusion. We show the significance of MViT proposals in a diverse range of applications including open-world object detection, salient and camouflage object detection, supervised and self-supervised detection tasks. Further, MViTs can adaptively generate proposals given a specific language query and thus offer enhanced interactability. Code: \\url{https://git.io/J1HPY}.","url_abs":"https://arxiv.org/abs/2111.11430v6","url_pdf":"https://arxiv.org/pdf/2111.11430v6.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"multi-modal-transformers-excel-at-class","repo_url":"https://github.com/mmaaz60/mvits_for_class_agnostic_od","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"class-agnostic-object-detection","task_name":"Class-agnostic Object Detection"},{"task_slug":"object","task_name":"Object"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"object-proposal-generation","task_name":"Object Proposal Generation"},{"task_slug":"open-world-object-detection","task_name":"Open World Object Detection"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[{"method_slug":"mavl","method_name":"MAVL"},{"method_slug":"mvit","method_name":"MViT"}],"datasets_introduced":[],"methods_introduced":[{"slug":"mavl","name":"MAVL","full_name":"Multiscale Attention ViT with Late fusion"}],"results":[{"leaderboard":"/sota/object-detection-on-pascal-voc-10","task":"Object Detection","dataset":"PASCAL VOC 10%","model":"DETReg (MDef-DETR)","rank_in_archive_order":1,"of":2,"metrics":{"AP":"58.78","AP50":"80.46","AP75":"65.65"},"uses_additional_data":false},{"leaderboard":"/sota/object-detection-on-pascal-voc-2007","task":"Object Detection","dataset":"PASCAL VOC 2007","model":"DETReg (MDef-DETR)","rank_in_archive_order":3,"of":30,"metrics":{"AP50":"84.16","MAP":"84.16%"},"uses_additional_data":true},{"leaderboard":"/sota/object-proposal-generation-on-coco","task":"Object Proposal Generation","dataset":"COCO (Common Objects in Context)","model":"MDef-DETR (Off-the-shelf evaluation)","rank_in_archive_order":1,"of":1,"metrics":{"Average Recall":"0.6503"},"uses_additional_data":false},{"leaderboard":"/sota/object-proposal-generation-on-pascal-voc-2012","task":"Object Proposal Generation","dataset":"PASCAL VOC 2012, 60 proposals per image","model":"MDef-DETR","rank_in_archive_order":1,"of":3,"metrics":{"Average Recall":"0.9126"},"uses_additional_data":false},{"leaderboard":"/sota/open-world-object-detection-on-coco-2017-2","task":"Open World Object Detection","dataset":"COCO 2017 (Electronic, Indoor, Kitchen, Furniture)","model":"ORE (MDef-DETR)","rank_in_archive_order":1,"of":2,"metrics":{"MAP":"31.66"},"uses_additional_data":false},{"leaderboard":"/sota/open-world-object-detection-on-coco-2017","task":"Open World Object Detection","dataset":"COCO 2017 (Outdoor, Accessories, Appliance, Truck)","model":"ORE (MDef-DETR)","rank_in_archive_order":1,"of":2,"metrics":{"A-OSE":"5212","MAP":"46.19","Unknown Recall":"49.54","WI":"0.0251"},"uses_additional_data":false},{"leaderboard":"/sota/open-world-object-detection-on-coco-2017-1","task":"Open World Object Detection","dataset":"COCO 2017 (Sports, Food)","model":"ORE (MDef-DETR)","rank_in_archive_order":1,"of":2,"metrics":{"A-OSE":"4117","MAP":"36.75","Unknown Recall":"50.89","WI":"0.0179"},"uses_additional_data":false},{"leaderboard":"/sota/open-world-object-detection-on-pascal-voc","task":"Open World Object Detection","dataset":"PASCAL VOC 2007","model":"ORE (MDef-DETR)","rank_in_archive_order":1,"of":2,"metrics":{"A-OSE":"7322","MAP":"64.03","Unknown Recall":"50.13","WI":"0.0474"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2111.11430","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}