Papers › Rethinking the Open-Loop Evaluation of End-to-End Autonomous Driving in nuScenes

Rethinking the Open-Loop Evaluation of End-to-End Autonomous Driving in nuScenes

17 May 2023arXiv:2305.10430archive 2025-07-28

Jiang-Tian Zhai, Ze Feng, Jinhao Du, Yongqiang Mao, Jiang-Jiang Liu, Zichang Tan, Yifu Zhang, Xiaoqing Ye, Jingdong Wang

Modern autonomous driving systems are typically divided into three main tasks: perception, prediction, and planning. The planning task involves predicting the trajectory of the ego vehicle based on inputs from both internal intention and the external environment, and manipulating the vehicle accordingly. Most existing works evaluate their performance on the nuScenes dataset using the L2 error and collision rate between the predicted trajectories and the ground truth. In this paper, we reevaluate these existing evaluation metrics and explore whether they accurately measure the superiority of different methods. Specifically, we design an MLP-based method that takes raw sensor data (e.g., past trajectory, velocity, etc.) as input and directly outputs the future trajectory of the ego vehicle, without using any perception or prediction information such as camera images or LiDAR. Our simple method achieves similar end-to-end planning performance on the nuScenes dataset with other perception-based methods, reducing the average L2 error by about 20%. Meanwhile, the perception-based methods have an advantage in terms of collision rate. We further conduct in-depth analysis and provide new insights into the factors that are critical for the success of the planning task on nuScenes dataset. Our observation also indicates that we need to rethink the current open-loop evaluation scheme of end-to-end autonomous driving in nuScenes. Codes are available at https://github.com/E2E-AD/AD-MLP.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

E2E-AD/AD-MLP officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Autonomous DrivingTrajectory Planning

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Trajectory Planning nuScenes AD-MLP Collision-1s 0.17 #1 of 4 Archive leaderboard report
Trajectory Planning nuScenes AD-MLP Collision-2s 0.18 #1 of 4 Archive leaderboard report
Trajectory Planning nuScenes AD-MLP Collision-3s 0.24 #1 of 4 Archive leaderboard report
Trajectory Planning nuScenes AD-MLP Collision-Avg 0.19 #1 of 4 Archive leaderboard report
Trajectory Planning nuScenes AD-MLP L2-1s 0.20 #1 of 4 Archive leaderboard report
Trajectory Planning nuScenes AD-MLP L2-2s 0.26 #1 of 4 Archive leaderboard report
Trajectory Planning nuScenes AD-MLP L2-3s 0.41 #1 of 4 Archive leaderboard report
Trajectory Planning nuScenes AD-MLP L2-Avg 0.29 #1 of 4 Archive leaderboard report
Trajectory Planning nuScenes VAD-Base [jiang2023vad] Collision-1s 0.07 #2 of 4 Archive leaderboard report
Trajectory Planning nuScenes VAD-Base [jiang2023vad] Collision-2s 0.10 #2 of 4 Archive leaderboard report
Trajectory Planning nuScenes VAD-Base [jiang2023vad] Collision-3s 0.24 #2 of 4 Archive leaderboard report
Trajectory Planning nuScenes VAD-Base [jiang2023vad] Collision-Avg 0.14 #2 of 4 Archive leaderboard report
Trajectory Planning nuScenes VAD-Base [jiang2023vad] L2-1s 0.17 #2 of 4 Archive leaderboard report
Trajectory Planning nuScenes VAD-Base [jiang2023vad] L2-2s 0.34 #2 of 4 Archive leaderboard report
Trajectory Planning nuScenes VAD-Base [jiang2023vad] L2-3s 0.60 #2 of 4 Archive leaderboard report
Trajectory Planning nuScenes VAD-Base [jiang2023vad] L2-Avg 0.37 #2 of 4 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections