Browse State-of-the-Art › 3D Object Detection From Monocular Images
3D Object Detection From Monocular Images
12 papers with code · 3 benchmarks · 3 datasets archive 2025-07-28
This is the task of detecting 3D objects from monocular images (as opposed to LiDAR based counterparts). It is usually associated with autonomous driving based tasks.
( Image credit: Orthographic Feature Transform for Monocular 3D Object Detection )
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
3 leaderboard tables shown for this task, 3 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| KITTI-360 (11 rows) | SeaBird + PanopticBEV | SeaBird: Segmentation in Bird's View with Dice Loss Improves... | code | — | Compare |
| Waymo Open Dataset (3 rows) | DEVIANT | DEVIANT: Depth EquiVarIAnt NeTwork for Monocular 3D Object Detection | code | Syntology ran 9 of 15 samples · 6 unverified | Compare |
| nuScenes Cars (1 row) | MonoDIS | Disentangling Monocular 3D Object Detection | — | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
3 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
12 shown of 12 papers with code (17 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
21 Apr 2019 13 repositories listed Syntology ran 2 of 10 samples · 8 unverified · 10 pointer-only (licence)Current 3D object detection methods are heavily influenced by 2D detectors.
-
13 Jul 2019 4 repositories listed Syntology ran 1 of 21 samples · 20 unverifiedUnderstanding the world in 3D is a critical component of urban autonomous driving.
-
21 Jul 2022 2 repositories listed Syntology ran 9 of 15 samples · 6 unverifiedAs a result, DEVIANT is equivariant to the depth translations in the projective manifold whereas vanilla networks are not.
-
SeaBird: Segmentation in Bird's View with Dice Loss Improves Monocular 3D Detection of Large Objects29 Mar 2024 1 repository listedWe argue that the cause of failure is the sensitivity of depth regression losses to noise of larger objects.
-
21 Jul 2022 1 repository listedIn 3D, existing benchmarks are small in size and approaches specialize in few object categories and specific domains, e.
-
24 Mar 2022 1 repository listedIn this paper, we introduce the first DETR framework for Monocular DEtection with a depth-guided TRansformer, named MonoDETR.
-
21 Mar 2022 1 repository listed Syntology ran 3 of 8 samples · 5 unverifiedMoreover, different from conventional pixel-wise positional encodings, we introduce a novel depth positional encoding (DPE) to inject depth positional hints into transformers.
-
3 Dec 2021 1 repository listedWe present ROCA, a novel end-to-end approach that retrieves and aligns 3D CAD models from a shape database to a single input image.
-
29 Jul 2021 1 repository listed Syntology ran 8 of 21 samples · 13 unverifiedIn this paper, we propose a Geometry Uncertainty Projection Network (GUP Net) to tackle the error amplification problem at both inference and training stages.
-
31 Mar 2021 1 repository listed Syntology ran 5 of 7 samples · 2 unverifiedIn this paper, we present and integrate GrooMeD-NMS -- a novel Grouped Mathematically Differentiable NMS for monocular 3D object detection, such that the network is trained end-to-end with a loss on the boxes after NMS.
-
30 Mar 2021 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedEstimating 3D bounding boxes from monocular images is an essential component in autonomous driving, while accurate 3D object detection from this kind of data is very challenging.
-
20 Nov 2018 1 repository listed Syntology ran 2 of 14 samples · 12 unverifiedThis allows us to reason holistically about the spatial configuration of the scene in a domain where scale is consistent and distances between objects are meaningful.
Syntology lines on 8 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections