{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/tokencut-segmenting-objects-in-images-and","title":"TokenCut: Segmenting Objects in Images and Videos with Self-supervised Transformer and Normalized Cut","arxiv_id":"2209.00383","date":"2022-09-01","proceeding":null,"authors":["Yangtao Wang","Xi Shen","Yuan Yuan","Yuming Du","Maomao Li","Shell Xu Hu","James L Crowley","Dominique Vaufreydaz"],"abstract":"In this paper, we describe a graph-based algorithm that uses the features obtained by a self-supervised transformer to detect and segment salient objects in images and videos. With this approach, the image patches that compose an image or video are organised into a fully connected graph, where the edge between each pair of patches is labeled with a similarity score between patches using features learned by the transformer. Detection and segmentation of salient objects is then formulated as a graph-cut problem and solved using the classical Normalized Cut algorithm. Despite the simplicity of this approach, it achieves state-of-the-art results on several common image and video detection and segmentation tasks. For unsupervised object discovery, this approach outperforms the competing approaches by a margin of 6.1%, 5.7%, and 2.6%, respectively, when tested with the VOC07, VOC12, and COCO20K datasets. For the unsupervised saliency detection task in images, this method improves the score for Intersection over Union (IoU) by 4.4%, 5.6% and 5.2%. When tested with the ECSSD, DUTS, and DUT-OMRON datasets, respectively, compared to current state-of-the-art techniques. This method also achieves competitive results for unsupervised video object segmentation tasks with the DAVIS, SegTV2, and FBMS datasets.","url_abs":"https://arxiv.org/abs/2209.00383v3","url_pdf":"https://arxiv.org/pdf/2209.00383v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"object-discovery","task_name":"Object Discovery"},{"task_slug":"saliency-detection","task_name":"Saliency Detection"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"unsupervised-instance-segmentation","task_name":"Unsupervised Instance Segmentation"},{"task_slug":"unsupervised-object-segmentation","task_name":"Unsupervised Object Segmentation"},{"task_slug":"unsupervised-saliency-detection","task_name":"Unsupervised Saliency Detection"},{"task_slug":"unsupervised-video-object-segmentation","task_name":"Unsupervised Video Object Segmentation"},{"task_slug":"video-object-segmentation","task_name":"Video Object Segmentation"},{"task_slug":"video-semantic-segmentation","task_name":"Video Semantic Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/unsupervised-instance-segmentation-on-coco","task":"Unsupervised Instance Segmentation","dataset":"COCO val2017","model":"TokenCut","rank_in_archive_order":5,"of":5,"metrics":{"AP":"2.4","AP50":"4.8","AP75":"1.9"},"uses_additional_data":false},{"leaderboard":"/sota/unsupervised-object-segmentation-on-fbms-59","task":"Unsupervised Object Segmentation","dataset":"FBMS-59","model":"TokenCut","rank_in_archive_order":6,"of":7,"metrics":{"mIoU":"60.2"},"uses_additional_data":false},{"leaderboard":"/sota/unsupervised-object-segmentation-on-segtrack","task":"Unsupervised Object Segmentation","dataset":"SegTrack-v2","model":"TokenCut","rank_in_archive_order":7,"of":8,"metrics":{"mIoU":"59.6"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2209.00383","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}