Papers › BEVFormer: Learning Bird's-Eye-View Representation from Multi-Camera Images via...

BEVFormer: Learning Bird's-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers

31 Mar 2022arXiv:2203.17270archive 2025-07-28

Zhiqi Li, Wenhai Wang, Hongyang Li, Enze Xie, Chonghao Sima, Tong Lu, Qiao Yu, Jifeng Dai

3D visual perception tasks, including 3D detection and map segmentation based on multi-camera images, are essential for autonomous driving systems. In this work, we present a new framework termed BEVFormer, which learns unified BEV representations with spatiotemporal transformers to support multiple autonomous driving perception tasks. In a nutshell, BEVFormer exploits both spatial and temporal information by interacting with spatial and temporal space through predefined grid-shaped BEV queries. To aggregate spatial information, we design spatial cross-attention that each BEV query extracts the spatial features from the regions of interest across camera views. For temporal information, we propose temporal self-attention to recurrently fuse the history BEV information. Our approach achieves the new state-of-the-art 56.9\% in terms of NDS metric on the nuScenes \texttt{test} set, which is 9.0 points higher than previous best arts and on par with the performance of LiDAR-based baselines. We further show that BEVFormer remarkably improves the accuracy of velocity estimation and recall of objects under low visibility conditions. The code is available at \url{https://github.com/zhiqi-li/BEVFormer}.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

zhiqi-li/BEVFormer officialmentioned in papermentioned on GitHub report
fundamentalvision/BEVFormer mentioned on GitHubpytorchApache-2.0 report
valeoai/pointbev mentioned on GitHubpytorchMIT report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

3D Object DetectionAutonomous DrivingBird's-Eye View Semantic SegmentationRobust Camera Only 3D Object Detection

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
3D Object Detection DAIR-V2X-I BEVFormer AP|R40(easy) 61.4 #8 of 9 Archive leaderboard report
3D Object Detection DAIR-V2X-I BEVFormer AP|R40(hard) 50.7 #8 of 9 Archive leaderboard report
3D Object Detection DAIR-V2X-I BEVFormer AP|R40(moderate) 50.7 #8 of 9 Archive leaderboard report
3D Object Detection nuScenes BEVFormer NDS 0.57 #222 of 372 Archive leaderboard report
3D Object Detection nuScenes BEVFormer mAAE 0.13 #222 of 372 Archive leaderboard report
3D Object Detection nuScenes BEVFormer mAOE 0.38 #222 of 372 Archive leaderboard report
3D Object Detection nuScenes BEVFormer mAP 0.48 #222 of 372 Archive leaderboard report
3D Object Detection nuScenes BEVFormer mASE 0.26 #222 of 372 Archive leaderboard report
3D Object Detection nuScenes BEVFormer mATE 0.58 #222 of 372 Archive leaderboard report
3D Object Detection nuScenes BEVFormer mAVE 0.38 #222 of 372 Archive leaderboard report
3D Object Detection nuScenes Camera Only BEVFormer Future Frame false #19 of 19 Archive leaderboard report
3D Object Detection nuScenes Camera Only BEVFormer NDS 56.9 #19 of 19 Archive leaderboard report
Bird's-Eye View Semantic Segmentation Lyft Level 5 BEVFormer (EfficientNet-b4) IoU vehicle - 224x480 - Long 44.5 #4 of 7 Archive leaderboard report
Bird's-Eye View Semantic Segmentation Lyft Level 5 BEVFormer (EfficientNet-b4) IoU vehicle - 224x480 - Short 69.9 #4 of 7 Archive leaderboard report
Bird's-Eye View Semantic Segmentation Lyft Level 5 BEVFormer(ResNet-50) IoU vehicle - 224x480 - Long 43.2 #6 of 7 Archive leaderboard report
Bird's-Eye View Semantic Segmentation Lyft Level 5 BEVFormer(ResNet-50) IoU vehicle - 224x480 - Short 68.8 #6 of 7 Archive leaderboard report
Bird's-Eye View Semantic Segmentation nuScenes BEVFormer IoU lane - 224x480 - 100x100 at 0.5 25.7 #4 of 17 Archive leaderboard report
Bird's-Eye View Semantic Segmentation nuScenes BEVFormer IoU veh - 224x480 - No vis filter - 100x100 at 0.5 35.8 #4 of 17 Archive leaderboard report
Bird's-Eye View Semantic Segmentation nuScenes BEVFormer IoU veh - 224x480 - Vis filter. - 100x100 at 0.5 42.0 #4 of 17 Archive leaderboard report
Bird's-Eye View Semantic Segmentation nuScenes BEVFormer IoU veh - 448x800 - No vis filter - 100x100 at 0.5 39.0 #4 of 17 Archive leaderboard report
Bird's-Eye View Semantic Segmentation nuScenes BEVFormer IoU veh - 448x800 - Vis filter. - 100x100 at 0.5 45.5 #4 of 17 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections