Papers › Enhancing Novel Object Detection via Cooperative Foundational Models
Enhancing Novel Object Detection via Cooperative Foundational Models
Rohit Bharadwaj, Muzammal Naseer, Salman Khan, Fahad Shahbaz Khan
In this work, we address the challenging and emergent problem of novel object detection (NOD), focusing on the accurate detection of both known and novel object categories during inference. Traditional object detection algorithms are inherently closed-set, limiting their capability to handle NOD. We present a novel approach to transform existing closed-set detectors into open-set detectors. This transformation is achieved by leveraging the complementary strengths of pre-trained foundational models, specifically CLIP and SAM, through our cooperative mechanism. Furthermore, by integrating this mechanism with state-of-the-art open-set detectors such as GDINO, we establish new benchmarks in object detection performance. Our method achieves 17.42 mAP in novel object detection and 42.08 mAP for known objects on the challenging LVIS dataset. Adapting our approach to the COCO OVD split, we surpass the current state-of-the-art by a margin of 7.2 AP₅₀ for novel classes. Our code is available at https://rohit901.github.io/coop-foundation-models/ .
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Novel Object Detection | LVIS v1.0 val | Cooperative Foundational Models | All mAP | 19.33 | #1 of 5 | Archive leaderboard | report |
| Novel Object Detection | LVIS v1.0 val | Cooperative Foundational Models | Known mAP | 42.08 | #1 of 5 | Archive leaderboard | report |
| Novel Object Detection | LVIS v1.0 val | Cooperative Foundational Models | Novel mAP | 17.42 | #1 of 5 | Archive leaderboard | report |
| Open Vocabulary Object Detection | MSCOCO | Cooperative Foundational Models | AP 0.5 | 50.3 | #1 of 32 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections