Papers › CubifAE-3D: Monocular Camera Space Cubification for Auto-Encoder based 3D Object Detection
CubifAE-3D: Monocular Camera Space Cubification for Auto-Encoder based 3D Object Detection
Shubham Shrivastava, Punarjay Chakravarty
We introduce a method for 3D object detection using a single monocular image. Starting from a synthetic dataset, we pre-train an RGB-to-Depth Auto-Encoder (AE). The embedding learnt from this AE is then used to train a 3D Object Detector (3DOD) CNN which is used to regress the parameters of 3D object poses after the encoder from the AE generates a latent embedding from the RGB image. We show that we can pre-train the AE using paired RGB and depth images from simulation data once and subsequently only train the 3DOD network using real data, comprising of RGB images and 3D object pose labels (without the requirement of dense depth). Our 3DOD network utilizes a particular `cubification' of 3D space around the camera, where each cuboid is tasked with predicting N object poses, along with their class and confidence values. The AE pre-training and this method of dividing the 3D space around the camera into cuboids give our method its name - CubifAE-3D. We demonstrate results for monocular 3D object detection in the Autonomous Vehicle (AV) use-case with the Virtual KITTI 2 and the KITTI datasets.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Monocular 3D Object Detection | KITTI Cars Hard | CubifAE-3D | AP Hard | 6.42 | #7 of 7 | Archive leaderboard | report |
| Monocular 3D Object Detection | KITTI Cars Moderate | CubifAE-3D | AP Medium | 7.94 | #26 of 29 | Archive leaderboard | report |
| Monocular 3D Object Detection | KITTI Pedestrian Hard | CubifAE-3D | AP Hard | 4.82 | #4 of 4 | Archive leaderboard | report |
| Monocular 3D Object Detection | KITTI Pedestrians Moderate val | CubifAE-3D | AP Medium | 5.43 | #1 of 1 | Archive leaderboard | report |
| Monocular 3D Object Detection | Virtual KITTI 2 | CubifAE-3D | mAP@0.3 | 86.6 | #1 of 1 | Archive leaderboard | report |
| Monocular 3D Object Detection | Virtual KITTI 2 | CubifAE-3D | mAP@0.5 | 66.7 | #1 of 1 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections