Browse State-of-the-Art › Monocular 3D Object Detection
Monocular 3D Object Detection
87 papers with code · 16 benchmarks · 7 datasets archive 2025-07-28
Monocular 3D Object Detection is the task to draw 3D bounding box around objects in a single 2D RGB image. It is localization task but without any extra information like depth or other sensors or multiple-images.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
16 leaderboard tables shown for this task, 16 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 16 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
7 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 87 papers with code (187 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
22 Apr 2021 9 repositories listed Syntology ran 7 of 22 samples · 15 unverifiedIn this paper, we study this problem with a practice built on a fully convolutional single-stage detector and propose a general framework FCOS3D.
-
13 Jul 2019 4 repositories listed Syntology ran 1 of 21 samples · 20 unverifiedUnderstanding the world in 3D is a critical component of urban autonomous driving.
-
6 Apr 2021 3 repositories listed Syntology ran 11 of 20 samples · 9 unverifiedThe precise localization of 3D objects from a single image without depth information is a highly challenging problem.
-
24 Feb 2020 3 repositories listed Syntology ran 2 of 9 samples · 7 unverifiedEstimating 3D orientation and translation of objects is essential for infrastructure-less autonomous navigation and driving.
-
21 Jul 2022 2 repositories listed Syntology ran 9 of 15 samples · 6 unverifiedAs a result, DEVIANT is equivariant to the depth translations in the projective manifold whereas vanilla networks are not.
-
9 Dec 2021 2 repositories listed Syntology ran 11 of 15 samples · 4 unverified · 12 pointer-only (licence)It presents the MonoCon method which learns Monocular Contexts, as auxiliary tasks in training, to help monocular 3D object detection.
-
20 Sep 2021 2 repositories listed Syntology ran 4 of 15 samples · 11 unverifiedUnlike the existing methods that use sparse LiDAR mainly in a manner of time-consuming iterative post-processing, our model fuses monocular image features and sparse LiDAR features to predict initial depth maps.
-
13 Aug 2021 2 repositories listed Syntology ran 0 of 3 samples · 3 unverifiedRecent progress in 3D object detection from single images leverages monocular depth estimation as a way to produce 3D pointclouds, turning cameras into pseudo-lidar sensors.
-
2 Jun 2021 2 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedTo address this problem, we propose ImVoxelNet, a novel fully convolutional method of 3D object detection based on monocular or multi-view RGB images.
-
1 Mar 2021 2 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedWe validate our approach on the KITTI 3D object detection benchmark, where we rank 1st among published monocular methods.
-
19 Jul 2020 2 repositories listed Syntology ran 3 of 20 samples · 17 unverifiedIn this work, we propose a novel method for monocular video-based 3D object detection which carefully leverages kinematic motion to improve precision of 3D localization.
-
10 Dec 2019 2 repositories listed Syntology ran 2 of 16 samples · 14 unverified3D object detection from a single image without LiDAR is a challenging task due to the lack of accurate depth information.
-
25 Nov 2024 1 repository listedIn this work, we pioneer the study of open-vocabulary monocular 3D object detection, a novel task that aims to detect and localize objects in 3D space from a single RGB image without limiting detection to a predefined…
-
5 Nov 2024 1 repository listedSpecifically, EH-FAM employs multi-head attention with a global receptive field to extract semantic features for small-scale objects and leverages lightweight convolutional modules to efficiently aggregate visual…
-
25 Oct 2024 1 repository listedHowever, due to depth errors originating from the object's visual surface, the height of the bounding box often fails to represent the actual projected central height, which undermines the effectiveness of geometric…
-
30 Jul 2024 1 repository listedExploring the use of cost-effective, extensive synthetic datasets offers a viable solution to tackle this challenge and enhance the performance of roadside monocular 3D detection.
-
23 Jul 2024 1 repository listed Syntology ran 15 of 26 samples · 11 unverifiedIt contains two components: (1) the weather codebook to memorize the knowledge of the clear weather and generate a weather-reference feature for any input, and (2) the weather-adaptive diffusion model to enhance the…
-
14 Jul 2024 1 repository listedRecent advancements in camera-based 3D object detection have introduced cross-modal knowledge distillation to bridge the performance gap with LiDAR 3D detectors, leveraging the precise geometric information in LiDAR…
-
30 May 2024 1 repository listedHowever, applying this paradigm in Mono 3Det poses significant challenges due to OOD test data causing a remarkable decline in object detection scores.
-
7 Apr 2024 1 repository listedMonocular 3D object detection (Mono3D) holds noteworthy promise for autonomous driving applications owing to the cost-effectiveness and rich visual context of monocular camera sensors.
-
4 Apr 2024 1 repository listed Syntology ran 3 of 6 samples · 3 unverifiedMonocular 3D object detection has attracted widespread attention due to its potential to accurately obtain object 3D localization from a single image at a low cost.
-
29 Mar 2024 1 repository listed Syntology ran 0 of 1 samples · 1 unverifiedIn the auto-labeling stage, we represent the surface of each instance as a signed distance field (SDF) and render its silhouette as an instance mask through our proposed instance-aware volumetric silhouette rendering.
-
22 Dec 2023 1 repository listedTo select samples adaptively, we propose a Learnable Sample Selection (LSS) module, which is based on Gumbel-Softmax and a relative-distance sample divider.
-
13 Dec 2023 1 repository listed Syntology ran 5 of 7 samples · 2 unverified · 7 pointer-only (licence)To foster this task, we propose Mono3DVG-TR, an end-to-end transformer-based network, which takes advantage of both the appearance and geometry information in text embeddings for multi-modal learning and 3D object…
-
8 Dec 2023 1 repository listed Syntology ran 1 of 2 samples · 1 unverifiedA major challenge in monocular 3D object detection is the limited diversity and quantity of objects in real datasets.
-
28 Oct 2023 1 repository listedMonocular 3D object detection (M3OD) is a significant yet inherently challenging task in autonomous driving due to absence of explicit depth cues in a single RGB image.
-
24 Oct 2023 1 repository listedIt models the uncertainty propagation relationship of the geometry projection during training, improving the stability and efficiency of the end-to-end model learning.
-
17 Oct 2023 1 repository listedMonocular 3D object detection is an inherently ill-posed problem, as it is challenging to predict accurate 3D localization from a single image.
-
4 Oct 2023 1 repository listedRoadside camera-driven 3D object detection is a crucial task in intelligent transportation systems, which extends the perception range beyond the limitations of vision-centric vehicles and enhances road safety.
-
21 Sep 2023 1 repository listedMonocular 3D detection of vehicle and infrastructure sides are two important topics in autonomous driving.
Syntology lines on 17 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections