Papers › Projecting Your View Attentively: Monocular Road Scene Layout Estimation via...

Projecting Your View Attentively: Monocular Road Scene Layout Estimation via Cross-View Transformation

19 Jun 2021CVPR 2021 1archive 2025-07-28

Weixiang Yang, Qi Li, Wenxi Liu, Yuanlong Yu, Yuexin Ma, Shengfeng He, Jia Pan

HD map reconstruction is crucial for autonomous driving. LiDAR-based methods are limited due to the deployed expensive sensors and time-consuming computation. Camera-based methods usually need to separately perform road segmentation and view transformation, which often causes distortion and the absence of content. To push the limits of the technology, we present a novel framework that enables reconstructing a local map formed by road layout and vehicle occupancy in the bird's-eye view given a front-view monocular image only. In particular, we propose a cross-view transformation module, which takes the constraint of cycle consistency between views into account and makes full use of their correlation to strengthen the view transformation and scene understanding. Considering the relationship between vehicles and roads, we also design a context-aware discriminator to further refine the results. Experiments on public benchmarks show that our method achieves the state-of-the-art performance in the tasks of road layout estimation and vehicle occupancy estimation. Especially for the latter task, our model outperforms all competitors by a large margin. Furthermore, our model runs at 35 FPS on a single GPU, which is efficient and applicable for real-time panorama HD map reconstruction.

PaperPDFCode

Code

JonDoe-297/cross-view officialpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Autonomous DrivingMonocular Cross-View Road Scene Parsing(Road)Monocular Cross-View Road Scene Parsing(Vehicle)Road SegmentationScene Understanding

1 archive task tag without a task page not shown.

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Monocular Cross-View Road Scene Parsing(Road) Argoverse crossView mAP 87.30% #2 of 2 Archive leaderboard report
Monocular Cross-View Road Scene Parsing(Road) Argoverse crossView mIOU 76.56% #2 of 2 Archive leaderboard report
Monocular Cross-View Road Scene Parsing(Road) Kitti Odometry crossView mAP 86.39% #2 of 2 Archive leaderboard report
Monocular Cross-View Road Scene Parsing(Road) Kitti Odometry crossView mIOU 77.47% #2 of 2 Archive leaderboard report
Monocular Cross-View Road Scene Parsing(Road) Kitti Raw crossView mAP 79.65% #2 of 2 Archive leaderboard report
Monocular Cross-View Road Scene Parsing(Road) Kitti Raw crossView mIoU 68.26% #2 of 2 Archive leaderboard report
Monocular Cross-View Road Scene Parsing(Vehicle) Argoverse crossView mAP 62.69% #2 of 2 Archive leaderboard report
Monocular Cross-View Road Scene Parsing(Vehicle) Argoverse crossView mIoU 47.87% #2 of 2 Archive leaderboard report
Monocular Cross-View Road Scene Parsing(Vehicle) KITTI2012 crossView mAP 51.04% #2 of 2 Archive leaderboard report
Monocular Cross-View Road Scene Parsing(Vehicle) KITTI2012 crossView mIoU 38.85% #2 of 2 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections