{"url":"/dataset/bdd-a","name":"BDD-A","full_name":"Berkeley DeepDrive Attention","description_markdown":"Dataset Statistics: The statistics of our dataset are summarized and compared with the\r\nlargest existing dataset (DR(eye)VE) [1] in Table 1. Our dataset was collected using videos\r\nselected from a publicly available, large-scale, crowd-sourced driving video dataset, BDD100k [30,\r\n31]. BDD100K contains human-demonstrated dashboard videos and time-stamped sensor\r\nmeasurements collected during urban driving in various weather and lighting conditions. To\r\nefficiently collect attention data for critical driving situations, we specifically selected video clips\r\nthat both included braking events and took place in busy areas (see supplementary materials\r\nfor technical details). We then trimmed videos to include 6.5 seconds prior to and 3.5 seconds\r\nafter each braking event. It turned out that other driving actions, e.g., turning, lane switching\r\nand accelerating, were also included. 1,232 videos (=3.5 hours) in total were collected following\r\nthese procedures. Some example images from our dataset are shown in Fig. 6. Our selected\r\nvideos contain a large number of different road users. We detected the objects in our videos\r\nusing YOLO [22].On average, each video frame contained 4.4 cars and 0.3 pedestrians, multiple\r\ntimes more than the DR(eye)VE dataset (Table 1).\r\nData Collection Procedure: For our eye-tracking experiment, we recruited 45 participants\r\nwho each had more than one year of driving experience. The participants watched the selected\r\ndriving videos in the lab while performing a driving instructor task: participants were asked\r\nto imagine that they were driving instructors sitting in the copilot seat and needed to press\r\nthe space key whenever they felt it necessary to correct or warn the student driver of potential\r\ndangers. Their eye movements during the task were recorded at 1000 Hz with an EyeLink 1000\r\ndesktop-mounted infrared eye tracker, used in conjunction with the Eyelink Toolbox scripts [7]\r\nfor MATLAB. Each participant completed the task for 200 driving videos. Each driving video\r\nwas viewed by at least 4 participants. The gaze patterns made by these independent participants\r\nwere aggregated and smoothed to make an attention map for each frame of the stimulus video\r\n(see Fig. 6 and supplementary materials for technical details).\r\nPsychological studies [19, 11] have shown that when humans look through multiple visual\r\ncues that simultaneously demand attention, the order in which humans look at those cues is\r\nhighly subjective. Therefore, by aggregating gazes of independent observers, we could record\r\nmultiple important visual cues in one frame. In addition, it has been shown that human drivers\r\nlook at buildings, trees, flowerbeds, and other unimportant objects non-negligibly frequently\r\n[1]. Presumably, these eye movements should be regarded as noise for driving-related machine\r\nlearning purposes. By averaging the eye movements of independent observers, we were able to\r\neffectively wash out those sources of noise (see Fig. 2B).\r\nComparison with In-Car Attention Data: We collected in-lab driver attention data using\r\nvideos from the DR(eye)VE dataset. This allowed us to compare in-lab and in-car attention\r\nmaps of each video. The DR(eye)VE videos we used were 200 randomly selected 10-second\r\nvideo clips, half of them containing braking events and half without braking events.\r\nWe tested how well in-car and in-lab attention maps highlighted driving-relevant objects.\r\nWe used YOLO [22] to detect the objects in the videos of our dataset. We identified three\r\nobject categories that are important for driving and that had sufficient instances in the videos\r\n(car, pedestrian and cyclist). We calculated the proportion of attended objects out of total\r\ndetected instances for each category for both in-lab and in-car attention maps (see supplementary\r\nmaterials for technical details). The results showed that in-car attention maps highlighted\r\nsignificantly less driving-relevant objects than in-lab attention maps (see Fig. 2A).\r\nThe difference in the number of attended objects between the in-car and in-lab attention maps\r\ncan be due to the fact that eye movements collected from a single driver do not completely indicate\r\nall the objects that demand attention in the particular driving situation. One individual’s eye\r\nmovements are only an approximation of their attention [23], and humans can also track objects\r\nwith covert attention without looking at them [6]. The difference in the number of attended\r\nobjects may also reflect the difference between first-person driver attention and third-person\r\ndriver attention. It may be that the human observers in our in-lab eye-tracking experiment also\r\nlooked at objects that were not relevant for driving. We ran a human evaluation experiment to\r\naddress this concern.\r\nHuman Evaluation: To verify that our in-lab driver attention maps highlight regions that\r\nshould indeed demand drivers’ attention, we conducted an online study to let humans compare\r\nin-lab and in-car driver attention maps. In each trial of the online study, participants watched\r\none driving video clip three times: the first time with no edit, and then two more times in\r\nrandom order with overlaid in-lab and in-car attention maps, respectively. The participant was\r\nthen asked to choose which heatmap-coded video was more similar to where a good driver would\r\nlook. In total, we collected 736 trials from 32 online participants. We found that our in-lab\r\nattention maps were more often preferred by the participants than the in-car attention maps\r\n(71% versus 29% of all trials, statistically significant as p = 1×10−29, see Table 2). Although\r\nthis result cannot suggest that in-lab driver attention maps are superior to in-car attention maps\r\nin general, it does show that the driver attention maps collected with our protocol represent\r\nwhere a good driver should look from a third-person perspective.\r\nIn addition, we will show in the Experiments section that in-lab attention data collected\r\nusing our protocol can be used to train a model to effectively predict actual, in-car driver\r\nattention. This result proves that our dataset can also serve as a substitute for in-car driver\r\nattention data, especially in crucial situations where in-car data collection is not practical.\r\nTo summarize, compared with driver attention data collected in-car, our dataset has three\r\nclear advantages: multi-focus, little driving-irrelevant noise, and efficiently tailored to crucial\r\ndriving situations.","description_withheld":null,"homepage":"https://bdd-data.berkeley.edu/","introduced_date":"2018-12-05","introduced_date_note":null,"introduced_by":{"paper":"/paper/predicting-driver-attention-in-critical","title":"Predicting Driver Attention in Critical Situations","first_author":"Ye Xia","url":null},"license":{"name":"Public","url":"https://bdd-data.berkeley.edu/"},"modalities":[{"name":"Videos","url":"/datasets/modality/videos"}],"tasks":[{"name":"Driver Attention Monitoring","url":"/task/driver-attention-monitoring","datasets_with_task":"/datasets/task/driver-attention-monitoring"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["BDD-A"],"data_loaders":[],"num_papers_in_archive":25,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}