{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/scenenet-rgb-d-5m-photorealistic-images-of","title":"SceneNet RGB-D: 5M Photorealistic Images of Synthetic Indoor Trajectories with Ground Truth","arxiv_id":"1612.05079","date":"2016-12-15","proceeding":null,"authors":["John McCormac","Ankur Handa","Stefan Leutenegger","Andrew J. Davison"],"abstract":"We introduce SceneNet RGB-D, expanding the previous work of SceneNet to\nenable large scale photorealistic rendering of indoor scene trajectories. It\nprovides pixel-perfect ground truth for scene understanding problems such as\nsemantic segmentation, instance segmentation, and object detection, and also\nfor geometric computer vision problems such as optical flow, depth estimation,\ncamera pose estimation, and 3D reconstruction. Random sampling permits\nvirtually unlimited scene configurations, and here we provide a set of 5M\nrendered RGB-D images from over 15K trajectories in synthetic layouts with\nrandom but physically simulated object poses. Each layout also has random\nlighting, camera trajectories, and textures. The scale of this dataset is well\nsuited for pre-training data-driven computer vision techniques from scratch\nwith RGB-D inputs, which previously has been limited by relatively small\nlabelled datasets in NYUv2 and SUN RGB-D. It also provides a basis for\ninvestigating 3D scene labelling tasks by providing perfect camera poses and\ndepth data as proxy for a SLAM system. We host the dataset at\nhttp://robotvault.bitbucket.io/scenenet-rgbd.html","url_abs":"http://arxiv.org/abs/1612.05079v3","url_pdf":"http://arxiv.org/pdf/1612.05079v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"scenenet-rgb-d-5m-photorealistic-images-of","repo_url":"https://github.com/fzi-forschungszentrum-informatik/mrf","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"3d-reconstruction","task_name":"3D Reconstruction"},{"task_slug":"camera-pose-estimation","task_name":"Camera Pose Estimation"},{"task_slug":"depth-estimation","task_name":"Depth Estimation"},{"task_slug":"instance-segmentation","task_name":"Instance Segmentation"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"optical-flow-estimation","task_name":"Optical Flow Estimation"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"},{"task_slug":"scene-understanding","task_name":"Scene Understanding"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1612.05079","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}