Methods › Computer Vision › Object Detection Models › MDETR
MDETR
Introduced by Aishwarya Kamath et al. in MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
MDETR is an end-to-end modulated detector that detects objects in an image conditioned on a raw text query, like a caption or a question. It utilizes a transformer-based architecture to reason jointly over text and image by fusing the two modalities at an early stage of the model. The network is pre-trained on 1.3M text-image pairs, mined from pre-existing multi-modal datasets having explicit alignment between phrases in text and objects in the image. The network is then fine-tuned on several downstream tasks such as phrase grounding, referring expression comprehension and segmentation.
Papers archive 2025-07-28
13 shown of 13, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Disambiguating Reference in Visually Grounded Dialogues through Joint Modeling of Textual and Multimodal Semantic Structures 16 May 2025 · 1 repository · arXiv:2505.11726Syntology ran 2 of 6 samples · 4 unverified
-
Seeing More with Less: Human-like Representations in Vision Models 1 Jan 2025 · 0 repositories
-
A Lightweight Modular Framework for Low-Cost Open-Vocabulary Object Detection Training 20 Aug 2024 · 1 repository · arXiv:2408.10787
-
ELSA: Evaluating Localization of Social Activities in Urban Streets using Open-Vocabulary Detection 3 Jun 2024 · 0 repositories · arXiv:2406.01551
-
Augment the Pairs: Semantics-Preserving Image-Caption Pair Augmentation for Grounding-Based Vision and Language Models 5 Nov 2023 · 1 repository · arXiv:2311.02536
-
3D-Aware Visual Question Answering about Parts, Poses and Occlusions 27 Oct 2023 · 2 repositories · arXiv:2310.17914Syntology ran 11 of 28 samples · 17 unverified · 14 pointer-only (licence)
-
Dynamic Inference With Grounding Based Vision and Language Models 1 Jan 2023 · 0 repositories
-
Dynamic MDETR: A Dynamic Multimodal Transformer Decoder for Visual Grounding 28 Sep 2022 · 0 repositories · arXiv:2209.13959
-
Exploring Modulated Detection Transformer as a Tool for Action Recognition in Videos 21 Sep 2022 · 1 repository · arXiv:2209.10126
-
Bottom Up Top Down Detection Transformers for Language Grounding in Images and Point Clouds 16 Dec 2021 · 1 repository · arXiv:2112.08879
-
Augmented 2D-TAN: A Two-stage Approach for Human-centric Spatio-Temporal Video Grounding 20 Jun 2021 · 0 repositories · arXiv:2106.10634
-
Team RUC_AIM3 Technical Report at ActivityNet 2021: Entities Object Localization 11 Jun 2021 · 1 repository · arXiv:2106.06138
-
MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding 26 Apr 2021 · 5 repositories · arXiv:2104.12763Syntology ran 6 of 11 samples · 5 unverified
Tasks archive 2025-07-28
20 shown of 34 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections