Papers › TokenCut: Segmenting Objects in Images and Videos with Self-supervised Transformer and...

TokenCut: Segmenting Objects in Images and Videos with Self-supervised Transformer and Normalized Cut

1 Sep 2022arXiv:2209.00383archive 2025-07-28

Yangtao Wang, Xi Shen, Yuan Yuan, Yuming Du, Maomao Li, Shell Xu Hu, James L Crowley, Dominique Vaufreydaz

In this paper, we describe a graph-based algorithm that uses the features obtained by a self-supervised transformer to detect and segment salient objects in images and videos. With this approach, the image patches that compose an image or video are organised into a fully connected graph, where the edge between each pair of patches is labeled with a similarity score between patches using features learned by the transformer. Detection and segmentation of salient objects is then formulated as a graph-cut problem and solved using the classical Normalized Cut algorithm. Despite the simplicity of this approach, it achieves state-of-the-art results on several common image and video detection and segmentation tasks. For unsupervised object discovery, this approach outperforms the competing approaches by a margin of 6.1%, 5.7%, and 2.6%, respectively, when tested with the VOC07, VOC12, and COCO20K datasets. For the unsupervised saliency detection task in images, this method improves the score for Intersection over Union (IoU) by 4.4%, 5.6% and 5.2%. When tested with the ECSSD, DUTS, and DUT-OMRON datasets, respectively, compared to current state-of-the-art techniques. This method also achieves competitive results for unsupervised video object segmentation tasks with the DAVIS, SegTV2, and FBMS datasets.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Object DiscoverySaliency DetectionSegmentationSemantic SegmentationUnsupervised Instance SegmentationUnsupervised Object SegmentationUnsupervised Saliency DetectionUnsupervised Video Object SegmentationVideo Object SegmentationVideo Semantic Segmentation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Unsupervised Instance Segmentation COCO val2017 TokenCut AP 2.4 #5 of 5 Archive leaderboard report
Unsupervised Instance Segmentation COCO val2017 TokenCut AP50 4.8 #5 of 5 Archive leaderboard report
Unsupervised Instance Segmentation COCO val2017 TokenCut AP75 1.9 #5 of 5 Archive leaderboard report
Unsupervised Object Segmentation FBMS-59 TokenCut mIoU 60.2 #6 of 7 Archive leaderboard report
Unsupervised Object Segmentation SegTrack-v2 TokenCut mIoU 59.6 #7 of 8 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections