{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/one-shot-segmentation-in-clutter","title":"One-Shot Segmentation in Clutter","arxiv_id":"1803.09597","date":"2018-03-26","proceeding":"ICML 2018 7","authors":["Claudio Michaelis","Matthias Bethge","Alexander S. Ecker"],"abstract":"We tackle the problem of one-shot segmentation: finding and segmenting a\npreviously unseen object in a cluttered scene based on a single instruction\nexample. We propose a novel dataset, which we call $\\textit{cluttered\nOmniglot}$. Using a baseline architecture combining a Siamese embedding for\ndetection with a U-net for segmentation we show that increasing levels of\nclutter make the task progressively harder. Using oracle models with access to\nvarious amounts of ground-truth information, we evaluate different aspects of\nthe problem and show that in this kind of visual search task, detection and\nsegmentation are two intertwined problems, the solution to each of which helps\nsolving the other. We therefore introduce $\\textit{MaskNet}$, an improved model\nthat attends to multiple candidate locations, generates segmentation proposals\nto mask out background clutter and selects among the segmented objects. Our\nfindings suggest that such image recognition models based on an iterative\nrefinement of object detection and foreground segmentation may provide a way to\ndeal with highly cluttered scenes.","url_abs":"http://arxiv.org/abs/1803.09597v2","url_pdf":"http://arxiv.org/pdf/1803.09597v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"one-shot-segmentation-in-clutter","repo_url":"https://github.com/michaelisc/cluttered-omniglot","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"foreground-segmentation","task_name":"Foreground Segmentation"},{"task_slug":"one-shot-segmentation","task_name":"One-Shot Segmentation"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[{"method_slug":"concatenated-skip-connection","method_name":"Concatenated Skip Connection"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"max-pooling","method_name":"Max Pooling"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"u-net","method_name":"U-Net"}],"datasets_introduced":[{"slug":"cluttered-omniglot","name":"Cluttered Omniglot","full_name":""}],"methods_introduced":[],"results":[{"leaderboard":"/sota/one-shot-segmentation-on-cluttered-omniglot","task":"One-Shot Segmentation","dataset":"Cluttered Omniglot","model":"MaskNet","rank_in_archive_order":1,"of":2,"metrics":{"IoU [256 distractors]":"43.7","IoU [32 distractors]":"65.6","IoU [4 distractors]":"95.8"},"uses_additional_data":false},{"leaderboard":"/sota/one-shot-segmentation-on-cluttered-omniglot","task":"One-Shot Segmentation","dataset":"Cluttered Omniglot","model":"Siamese-U-Net","rank_in_archive_order":2,"of":2,"metrics":{"IoU [256 distractors]":"38.4","IoU [32 distractors]":"62.4","IoU [4 distractors]":"97.1"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1803.09597","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}