Browse State-of-the-Art › 3D Object Detection

3D Object Detection

764 papers with code · 67 benchmarks · 67 datasets archive 2025-07-28

Computer Vision

3D Object Detection is a task in computer vision where the goal is to identify and locate objects in a 3D environment based on their shape, location, and orientation. It involves detecting the presence of objects and determining their location in the 3D space in real-time. This task is crucial for applications such as autonomous vehicles, robotics, and augmented reality.

( Image credit: AVOD )

Description from the archive archive 2025-07-28.

Benchmarks archive 2025-07-28

67 leaderboard tables shown for this task, 67 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 67 until expanded.

DatasetBest model (first row in archive order)PaperCodeSyntologyCompare
nuScenes (372 rows) EA-LSS EA-LSS: Edge-aware Lift-splat-shot Framework for 3D BEV Object Detection code Syntology ran 2 of 2 samples · 0 unverified Compare
ScanNetV2 (33 rows) DEST (based on V-DETR) (TTA) State Space Model Meets Transformer: A New Paradigm for 3D Object Detection code Syntology ran 2 of 10 samples · 8 unverified Compare
SUN-RGBD val (32 rows) Point-GCC+TR3D+FF Point-GCC: Universal Self-supervised 3D Scene Pre-training via... code — Compare
KITTI Cars Easy (26 rows) TRTConv — — — Compare
KITTI Cars Hard (25 rows) TRTConv — — — Compare
nuScenes Camera Only (19 rows) Far3D Far3D: Expanding the Horizon for Surround-view 3D Object Detection code Syntology ran 8 of 8 samples · 0 unverified Compare
KITTI Cyclists Moderate (13 rows) 3D-FCT 3D-FCT: Simultaneous 3D Object Detection and Tracking Using... — — Compare
KITTI Cyclists Easy (12 rows) 3D-FCT 3D-FCT: Simultaneous 3D Object Detection and Tracking Using... — — Compare
KITTI Cyclists Hard (12 rows) SA-Det3D SA-Det3D: Self-Attention Based Context-Aware 3D Object Detection code — Compare
KITTI Pedestrians Moderate (12 rows) 3D-FCT 3D-FCT: Simultaneous 3D Object Detection and Tracking Using... — — Compare
KITTI Cars Easy val (11 rows) SA-SSD+EBM Accurate 3D Object Detection using Energy-Based Models code — Compare
KITTI Cars Moderate val (11 rows) SA-SSD+EBM Accurate 3D Object Detection using Energy-Based Models code — Compare
nuscenes Camera-Radar (11 rows) SpaRC SpaRC: Sparse Radar-Camera Fusion for 3D Object Detection — — Compare
View-of-Delft (val) (11 rows) HyDRa Unleashing HyDRa: Hybrid Fusion, Depth Consistency and Radar for... code — Compare
KITTI Cars Hard val (10 rows) M3DeTR M3DeTR: Multi-representation, Multi-scale, Mutual-relation 3D... code Syntology ran 0 of 7 samples · 7 unverified Compare
DAIR-V2X-I (9 rows) MonoUNI MonoUNI: A Unified Vehicle and Infrastructure-side Monocular 3D... code — Compare
KITTI Pedestrians Easy (9 rows) IPOD IPOD: Intensive Point-based Object Detector for Point Cloud — — Compare
KITTI Pedestrians Hard (9 rows) SVGA-Net SVGA-Net: Sparse Voxel-Graph Attention Network for 3D Object... — — Compare
Rope3D (8 rows) MonoUNI MonoUNI: A Unified Vehicle and Infrastructure-side Monocular 3D... code — Compare
SUN-RGBD (8 rows) Uni3DETR Uni3DETR: Unified 3D Detection Transformer code Syntology ran 1 of 1 samples · 0 unverified Compare
Waymo Open Dataset (8 rows) LION LION: Linear Group RNN for 3D Object Detection in Point Clouds code Syntology ran 0 of 7 samples · 7 unverified Compare
waymo vehicle (8 rows) PillarNeXt PillarNeXt: Rethinking Network Designs for 3D Object Detection in... code — Compare
nuScenes LiDAR only (7 rows) LION LION: Linear Group RNN for 3D Object Detection in Point Clouds code Syntology ran 0 of 7 samples · 7 unverified Compare
S3DIS (7 rows) UniDet3D UniDet3D: Multi-dataset Indoor 3D Object Detection code — Compare
waymo cyclist (7 rows) DSVT(val) DSVT: Dynamic Sparse Voxel Transformer with Rotated Sets code — Compare
waymo pedestrian (7 rows) DSVT(val) DSVT: Dynamic Sparse Voxel Transformer with Rotated Sets code — Compare
Argoverse (6 rows) tempVar — — — Compare
V2XSet (6 rows) V2X-ViT V2X-ViT: Vehicle-to-Everything Cooperative Perception with Vision... code Syntology ran 0 of 3 samples · 3 unverified Compare
DTTD-Mobile (5 rows) DTTDNet Robust 6DoF Pose Estimation Against Depth Noise and a... code — Compare
OPV2V (5 rows) V2VNet (PointPillar backbone) V2VNet: Vehicle-to-Vehicle Communication for Joint Perception and... code Syntology ran 2 of 2 samples · 0 unverified Compare
SimBEV (5 rows) UniTR+LSS SimBEV: A Synthetic Multi-Task Multi-Sensor Driving Data... code — Compare
V2X-SIM (5 rows) QUEST QUEST: Query Stream for Practical Cooperative Perception code — Compare
Aria Everyday Objects (4 rows) EVL EFM3D: A Benchmark for Measuring Progress Towards 3D Egocentric... code Syntology ran 3 of 3 samples · 0 unverified Compare
Aria Synthetic Environments (4 rows) EVL EFM3D: A Benchmark for Measuring Progress Towards 3D Egocentric... code Syntology ran 3 of 3 samples · 0 unverified Compare
ARKitScenes (4 rows) UniDet3D UniDet3D: Multi-dataset Indoor 3D Object Detection code — Compare
KITTI Cyclist Easy val (4 rows) M3DeTR M3DeTR: Multi-representation, Multi-scale, Mutual-relation 3D... code Syntology ran 0 of 7 samples · 7 unverified Compare
KITTI Cyclist Hard val (4 rows) M3DeTR M3DeTR: Multi-representation, Multi-scale, Mutual-relation 3D... code Syntology ran 0 of 7 samples · 7 unverified Compare
KITTI Cyclist Moderate val (4 rows) M3DeTR M3DeTR: Multi-representation, Multi-scale, Mutual-relation 3D... code Syntology ran 0 of 7 samples · 7 unverified Compare
KITTI Pedestrian Moderate val (4 rows) PVCNN Point-Voxel CNN for Efficient 3D Deep Learning code Syntology ran 5 of 8 samples · 3 unverified Compare
KITTI Pedestrian Easy val (4 rows) PVCNN Point-Voxel CNN for Efficient 3D Deep Learning code Syntology ran 5 of 8 samples · 3 unverified Compare
KITTI Pedestrian Hard val (4 rows) PVCNN Point-Voxel CNN for Efficient 3D Deep Learning code Syntology ran 5 of 8 samples · 3 unverified Compare
TruckScenes (4 rows) SpaRC SpaRC: Sparse Radar-Camera Fusion for 3D Object Detection — — Compare
3D Object Detection on Argoverse2 Camera Only (3 rows) Far3D Far3D: Expanding the Horizon for Surround-view 3D Object Detection code Syntology ran 8 of 8 samples · 0 unverified Compare
3RScan (3 rows) UniDet3D UniDet3D: Multi-dataset Indoor 3D Object Detection code — Compare
aiMotive Dataset (3 rows) Lidar-Radar-Camera aiMotive Dataset: A Multimodal Dataset for Robust Autonomous... code Syntology ran 0 of 5 samples · 5 unverified Compare
MultiScan (3 rows) UniDet3D UniDet3D: Multi-dataset Indoor 3D Object Detection code — Compare
ScanNet++ (3 rows) UniDet3D UniDet3D: Multi-dataset Indoor 3D Object Detection code — Compare
Argoverse2 (2 rows) LION LION: Linear Group RNN for 3D Object Detection in Point Clouds code Syntology ran 0 of 7 samples · 7 unverified Compare
DAIR-V2X (2 rows) CoBEVFlow — — — Compare
ONCE (2 rows) LION LION: Linear Group RNN for 3D Object Detection in Point Clouds code Syntology ran 0 of 7 samples · 7 unverified Compare
Spiideo SoccerNet SynLoc (2 rows) Baseline-960x960 Spiideo SoccerNet SynLoc: Single Frame World Coordinate Athlete... code — Compare
waymo all_ns (2 rows) CenterPoint Center-based 3D Object Detection and Tracking code Syntology ran 7 of 22 samples · 15 unverified Compare
Cityscapes 3D (1 row) TaskPrompter Joint 2D-3D Multi-Task Learning on Cityscapes-3D: 3D Detection,... code — Compare
Clear Weather (1 row) PV-RCNN LiDAR Snowfall Simulation for Robust 3D Object Detection code — Compare
IRV2V (1 row) CoBEVFlow — — — Compare
KITTI Cyclists Moderate val (1 row) Deformable PV-RCNN Deformable PV-RCNN: Improving 3D Object Detection with Learned Deformations code Syntology ran 1 of 1 samples · 0 unverified Compare
KITTI Pedestrian Moderate (1 row) PiFeNet Accurate and Real-time 3D Pedestrian Detection Using an Efficient... code — Compare
KITTI Pedestrian (1 row) PiFeNet Accurate and Real-time 3D Pedestrian Detection Using an Efficient... code — Compare
KITTI Pedestrian Easy (1 row) PiFeNet Accurate and Real-time 3D Pedestrian Detection Using an Efficient... code — Compare
KITTI Pedestrian Hard (1 row) PiFeNet Accurate and Real-time 3D Pedestrian Detection Using an Efficient... code — Compare
KITTI Pedestrians Moderate val (1 row) Deformable PV-RCNN Deformable PV-RCNN: Improving 3D Object Detection with Learned Deformations code Syntology ran 1 of 1 samples · 0 unverified Compare
Light Snowfall (1 row) PV-RCNN LiDAR Snowfall Simulation for Robust 3D Object Detection code — Compare
nuScenes-F (1 row) RRPN + R101 - F RRPN: Radar Region Proposal Network for Object Detection in... code — Compare
nuScenes-FB (1 row) RRPN + R101 - FB RRPN: Radar Region Proposal Network for Object Detection in... code — Compare
NYU Depth v2 (1 row) SGPN-CNN SGPN: Similarity Group Proposal Network for 3D Point Cloud... code — Compare
Dense Fog (1 row) PV-RCNN Fog Simulation on Real LiDAR Point Clouds for 3D Object Detection... code Syntology ran 1 of 6 samples · 5 unverified Compare
Heavy Snowfall (1 row) PV-RCNN LiDAR Snowfall Simulation for Robust 3D Object Detection code — Compare

Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.

Libraries

Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.

Datasets archive 2025-07-28

67 datasets whose archive record lists this task, ordered by the archive's paper count. 30 shown of 67 until expanded.

Subtasks archive 2025-07-28

5 subtasks in the archive's task tree.

Parent tasks archive 2025-07-28

Most implemented papers archive 2025-07-28

30 shown of 764 papers with code (1,576 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.

Syntology lines on 23 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections