{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/video-object-segmentation-using-teacher","title":"Video Object Segmentation using Teacher-Student Adaptation in a Human Robot Interaction (HRI) Setting","arxiv_id":"1810.07733","date":"2018-10-17","proceeding":null,"authors":["Mennatullah Siam","Chen Jiang","Steven Lu","Laura Petrich","Mahmoud Gamal","Mohamed Elhoseiny","Martin Jagersand"],"abstract":"Video object segmentation is an essential task in robot manipulation to\nfacilitate grasping and learning affordances. Incremental learning is important\nfor robotics in unstructured environments, since the total number of objects\nand their variations can be intractable. Inspired by the children learning\nprocess, human robot interaction (HRI) can be utilized to teach robots about\nthe world guided by humans similar to how children learn from a parent or a\nteacher. A human teacher can show potential objects of interest to the robot,\nwhich is able to self adapt to the teaching signal without providing manual\nsegmentation labels. We propose a novel teacher-student learning paradigm to\nteach robots about their surrounding environment. A two-stream motion and\nappearance \"teacher\" network provides pseudo-labels to adapt an appearance\n\"student\" network. The student network is able to segment the newly learned\nobjects in other scenes, whether they are static or in motion. We also\nintroduce a carefully designed dataset that serves the proposed HRI setup,\ndenoted as (I)nteractive (V)ideo (O)bject (S)egmentation. Our IVOS dataset\ncontains teaching videos of different objects, and manipulation tasks. Unlike\nprevious datasets, IVOS provides manipulation tasks sequences with segmentation\nannotation along with the waypoints for the robot trajectories. It also\nprovides segmentation annotation for the different transformations such as\ntranslation, scale, planar rotation, and out-of-plane rotation. Our proposed\nadaptation method outperforms the state-of-the-art on DAVIS and FBMS with 6.8%\nand 1.2% in F-measure respectively. It improves over the baseline on IVOS\ndataset with 46.1% and 25.9% in mIoU.","url_abs":"http://arxiv.org/abs/1810.07733v4","url_pdf":"http://arxiv.org/pdf/1810.07733v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"video-object-segmentation-using-teacher","repo_url":"https://github.com/MSiam/motion_adaptation","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"incremental-learning","task_name":"Incremental Learning"},{"task_slug":"robot-manipulation","task_name":"Robot Manipulation"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"unsupervised-video-object-segmentation","task_name":"Unsupervised Video Object Segmentation"},{"task_slug":"video-object-segmentation","task_name":"Video Object Segmentation"},{"task_slug":"video-semantic-segmentation","task_name":"Video Semantic Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1810.07733","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}