{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/deep-steering-learning-end-to-end-driving","title":"Deep Steering: Learning End-to-End Driving Model from Spatial and Temporal Visual Cues","arxiv_id":"1708.03798","date":"2017-08-12","proceeding":null,"authors":["Lu Chi","Yadong Mu"],"abstract":"In recent years, autonomous driving algorithms using low-cost vehicle-mounted\ncameras have attracted increasing endeavors from both academia and industry.\nThere are multiple fronts to these endeavors, including object detection on\nroads, 3-D reconstruction etc., but in this work we focus on a vision-based\nmodel that directly maps raw input images to steering angles using deep\nnetworks. This represents a nascent research topic in computer vision. The\ntechnical contributions of this work are three-fold. First, the model is\nlearned and evaluated on real human driving videos that are time-synchronized\nwith other vehicle sensors. This differs from many prior models trained from\nsynthetic data in racing games. Second, state-of-the-art models, such as\nPilotNet, mostly predict the wheel angles independently on each video frame,\nwhich contradicts common understanding of driving as a stateful process.\nInstead, our proposed model strikes a combination of spatial and temporal cues,\njointly investigating instantaneous monocular camera observations and vehicle's\nhistorical states. This is in practice accomplished by inserting\ncarefully-designed recurrent units (e.g., LSTM and Conv-LSTM) at proper network\nlayers. Third, to facilitate the interpretability of the learned model, we\nutilize a visual back-propagation scheme for discovering and visualizing image\nregions crucially influencing the final steering prediction. Our experimental\nstudy is based on about 6 hours of human driving data provided by Udacity.\nComprehensive quantitative evaluations demonstrate the effectiveness and\nrobustness of our model, even under scenarios like drastic lighting changes and\nabrupt turning. The comparison with other state-of-the-art models clearly\nreveals its superior performance in predicting the due wheel angle for a\nself-driving car.","url_abs":"http://arxiv.org/abs/1708.03798v1","url_pdf":"http://arxiv.org/pdf/1708.03798v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"deep-steering-learning-end-to-end-driving","repo_url":"https://github.com/abhileshborode/Behavorial-Clonng-Self-driving-cars","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"autonomous-driving","task_name":"Autonomous Driving"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[{"method_slug":"interpretability","method_name":"Interpretability"},{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1708.03798","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}