{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/end-to-end-multi-modal-multi-task-vehicle","title":"End-to-end Multi-Modal Multi-Task Vehicle Control for Self-Driving Cars with Visual Perception","arxiv_id":"1801.06734","date":"2018-01-20","proceeding":null,"authors":["Zhengyuan Yang","Yixuan Zhang","Jerry Yu","Junjie Cai","Jiebo Luo"],"abstract":"Convolutional Neural Networks (CNN) have been successfully applied to\nautonomous driving tasks, many in an end-to-end manner. Previous end-to-end\nsteering control methods take an image or an image sequence as the input and\ndirectly predict the steering angle with CNN. Although single task learning on\nsteering angles has reported good performances, the steering angle alone is not\nsufficient for vehicle control. In this work, we propose a multi-task learning\nframework to predict the steering angle and speed control simultaneously in an\nend-to-end manner. Since it is nontrivial to predict accurate speed values with\nonly visual inputs, we first propose a network to predict discrete speed\ncommands and steering angles with image sequences. Moreover, we propose a\nmulti-modal multi-task network to predict speed values and steering angles by\ntaking previous feedback speeds and visual recordings as inputs. Experiments\nare conducted on the public Udacity dataset and a newly collected SAIC dataset.\nResults show that the proposed model predicts steering angles and speed values\naccurately. Furthermore, we improve the failure data synthesis methods to solve\nthe problem of error accumulation in real road tests.","url_abs":"http://arxiv.org/abs/1801.06734v2","url_pdf":"http://arxiv.org/pdf/1801.06734v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"end-to-end-multi-modal-multi-task-vehicle","repo_url":"https://github.com/rehamessameltagoury/Urban-Self-Driving-Car","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"autonomous-driving","task_name":"Autonomous Driving"},{"task_slug":"multi-task-learning","task_name":"Multi-Task Learning"},{"task_slug":"self-driving-cars","task_name":"Self-Driving Cars"},{"task_slug":"steering-control","task_name":"Steering Control"}],"methods":[{"method_slug":"speed","method_name":"SPEED"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1801.06734","atlas_url":"https://app.syntology.ai/?focus=1801.06734","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}