{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/textual-explanations-for-self-driving","title":"Textual Explanations for Self-Driving Vehicles","arxiv_id":"1807.11546","date":"2018-07-30","proceeding":"ECCV 2018 9","authors":["Jinkyu Kim","Anna Rohrbach","Trevor Darrell","John Canny","Zeynep Akata"],"abstract":"Deep neural perception and control networks have become key components of\nself-driving vehicles. User acceptance is likely to benefit from\neasy-to-interpret textual explanations which allow end-users to understand what\ntriggered a particular behavior. Explanations may be triggered by the neural\ncontroller, namely introspective explanations, or informed by the neural\ncontroller's output, namely rationalizations. We propose a new approach to\nintrospective explanations which consists of two parts. First, we use a visual\n(spatial) attention model to train a convolutional network end-to-end from\nimages to the vehicle control commands, i.e., acceleration and change of\ncourse. The controller's attention identifies image regions that potentially\ninfluence the network's output. Second, we use an attention-based video-to-text\nmodel to produce textual explanations of model actions. The attention maps of\ncontroller and explanation model are aligned so that explanations are grounded\nin the parts of the scene that mattered to the controller. We explore two\napproaches to attention alignment, strong- and weak-alignment. Finally, we\nexplore a version of our model that generates rationalizations, and compare\nwith introspective explanations on the same video segments. We evaluate these\nmodels on a novel driving dataset with ground-truth human explanations, the\nBerkeley DeepDrive eXplanation (BDD-X) dataset. Code is available at\nhttps://github.com/JinkyuKimUCB/explainable-deep-driving.","url_abs":"http://arxiv.org/abs/1807.11546v1","url_pdf":"http://arxiv.org/pdf/1807.11546v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"textual-explanations-for-self-driving","repo_url":"https://github.com/JinkyuKimUCB/explainable-deep-driving","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok","spdx":"NOASSERTION"}},{"paper_slug":"textual-explanations-for-self-driving","repo_url":"https://github.com/Manojbhat09/Commonsense_action_recognition","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[],"methods":[],"datasets_introduced":[{"slug":"bdd-x","name":"BDD-X","full_name":"Berkeley Deep Drive-X (eXplanation)"}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1807.11546","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}