{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/playing-for-benchmarks","title":"Playing for Benchmarks","arxiv_id":"1709.07322","date":"2017-09-21","proceeding":"ICCV 2017 10","authors":["Stephan R. Richter","Zeeshan Hayder","Vladlen Koltun"],"abstract":"We present a benchmark suite for visual perception. The benchmark is based on\nmore than 250K high-resolution video frames, all annotated with ground-truth\ndata for both low-level and high-level vision tasks, including optical flow,\nsemantic instance segmentation, object detection and tracking, object-level 3D\nscene layout, and visual odometry. Ground-truth data for all tasks is available\nfor every frame. The data was collected while driving, riding, and walking a\ntotal of 184 kilometers in diverse ambient conditions in a realistic virtual\nworld. To create the benchmark, we have developed a new approach to collecting\nground-truth data from simulated worlds without access to their source code or\ncontent. We conduct statistical analyses that show that the composition of the\nscenes in the benchmark closely matches the composition of corresponding\nphysical environments. The realism of the collected data is further validated\nvia perceptual experiments. We analyze the performance of state-of-the-art\nmethods for multiple tasks, providing reference baselines and highlighting\nchallenges for future research. The supplementary video can be viewed at\nhttps://youtu.be/T9OybWv923Y","url_abs":"http://arxiv.org/abs/1709.07322v1","url_pdf":"http://arxiv.org/pdf/1709.07322v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"instance-segmentation","task_name":"Instance Segmentation"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"optical-flow-estimation","task_name":"Optical Flow Estimation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"visual-odometry","task_name":"Visual Odometry"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[],"datasets_introduced":[{"slug":"visual-perception-viper","name":"VIsual PERception (VIPER)","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1709.07322","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}