{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/the-rgb-d-triathlon-towards-agile-visual","title":"The RGB-D Triathlon: Towards Agile Visual Toolboxes for Robots","arxiv_id":"1904.00912","date":"2019-04-01","proceeding":null,"authors":["Fabio Cermelli","Massimiliano Mancini","Elisa Ricci","Barbara Caputo"],"abstract":"Deep networks have brought significant advances in robot perception, enabling\nto improve the capabilities of robots in several visual tasks, ranging from\nobject detection and recognition to pose estimation, semantic scene\nsegmentation and many others. Still, most approaches typically address visual\ntasks in isolation, resulting in overspecialized models which achieve strong\nperformances in specific applications but work poorly in other (often related)\ntasks. This is clearly sub-optimal for a robot which is often required to\nperform simultaneously multiple visual recognition tasks in order to properly\nact and interact with the environment. This problem is exacerbated by the\nlimited computational and memory resources typically available onboard to a\nrobotic platform. The problem of learning flexible models which can handle\nmultiple tasks in a lightweight manner has recently gained attention in the\ncomputer vision community and benchmarks supporting this research have been\nproposed. In this work we study this problem in the robot vision context,\nproposing a new benchmark, the RGB-D Triathlon, and evaluating state of the art\nalgorithms in this novel challenging scenario. We also define a new evaluation\nprotocol, better suited to the robot vision setting. Results shed light on the\nstrengths and weaknesses of existing approaches and on open issues, suggesting\ndirections for future research.","url_abs":"http://arxiv.org/abs/1904.00912v2","url_pdf":"http://arxiv.org/pdf/1904.00912v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"the-rgb-d-triathlon-towards-agile-visual","repo_url":"https://github.com/fcdl94/RobotChallenge","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"},{"task_slug":"scene-segmentation","task_name":"Scene Segmentation"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}