{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/vision-meets-drones-a-challenge","title":"Vision Meets Drones: A Challenge","arxiv_id":"1804.07437","date":"2018-04-20","proceeding":null,"authors":["Pengfei Zhu","Longyin Wen","Xiao Bian","Haibin Ling","QinGhua Hu"],"abstract":"In this paper we present a large-scale visual object detection and tracking\nbenchmark, named VisDrone2018, aiming at advancing visual understanding tasks\non the drone platform. The images and video sequences in the benchmark were\ncaptured over various urban/suburban areas of 14 different cities across China\nfrom north to south. Specifically, VisDrone2018 consists of 263 video clips and\n10,209 images (no overlap with video clips) with rich annotations, including\nobject bounding boxes, object categories, occlusion, truncation ratios, etc.\nWith intensive amount of effort, our benchmark has more than 2.5 million\nannotated instances in 179,264 images/video frames. Being the largest such\ndataset ever published, the benchmark enables extensive evaluation and\ninvestigation of visual analysis algorithms on the drone platform. In\nparticular, we design four popular tasks with the benchmark, including object\ndetection in images, object detection in videos, single object tracking, and\nmulti-object tracking. All these tasks are extremely challenging in the\nproposed dataset due to factors such as occlusion, large scale and pose\nvariation, and fast motion. We hope the benchmark largely boost the research\nand development in visual analysis on drone platforms.","url_abs":"http://arxiv.org/abs/1804.07437v2","url_pdf":"http://arxiv.org/pdf/1804.07437v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"multi-object-tracking","task_name":"Multi-Object Tracking"},{"task_slug":"object","task_name":"Object"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"object-tracking","task_name":"Object Tracking"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[],"datasets_introduced":[{"slug":"visdrone","name":"VisDrone","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1804.07437","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}