Papers › MAST: A Memory-Augmented Self-supervised Tracker

MAST: A Memory-Augmented Self-supervised Tracker

18 Feb 2020CVPR 2020 6arXiv:2002.07793archive 2025-07-28

Zihang Lai, Erika Lu, Weidi Xie

Recent interest in self-supervised dense tracking has yielded rapid progress, but performance still remains far from supervised methods. We propose a dense tracking model trained on videos without any annotations that surpasses previous self-supervised methods on existing benchmarks by a significant margin (+15%), and achieves performance comparable to supervised methods. In this paper, we first reassess the traditional choices used for self-supervised training and reconstruction loss by conducting thorough experiments that finally elucidate the optimal choices. Second, we further improve on existing methods by augmenting our architecture with a crucial memory component. Third, we benchmark on large-scale semi-supervised video object segmentation(aka. dense tracking), and propose a new metric: generalizability. Our first two contributions yield a self-supervised network that for the first time is competitive with supervised methods on standard evaluation metrics of dense tracking. When measuring generalizability, we show self-supervised approaches are actually superior to the majority of supervised methods. We believe this new generalizability metric can better capture the real-world use-cases for dense tracking, and will spur new interest in this research direction.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

zlai0/MAST officialmentioned in papermentioned on GitHubpytorch report
bo-miao/MAMP mentioned on GitHubpytorchBSD-3-Clause report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Semantic SegmentationSemi-Supervised Video Object SegmentationUnsupervised Video Object SegmentationVideo Object SegmentationVideo Semantic Segmentation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Semi-Supervised Video Object Segmentation DAVIS 2017 (val) MAST F-measure (Mean) 67.6 #67 of 81 Archive leaderboard report
Semi-Supervised Video Object Segmentation DAVIS 2017 (val) MAST F-measure (Recall) 77.7 #67 of 81 Archive leaderboard report
Semi-Supervised Video Object Segmentation DAVIS 2017 (val) MAST J&F 65.5 #67 of 81 Archive leaderboard report
Semi-Supervised Video Object Segmentation DAVIS 2017 (val) MAST Jaccard (Mean) 63.3 #67 of 81 Archive leaderboard report
Semi-Supervised Video Object Segmentation DAVIS 2017 (val) MAST Jaccard (Recall) 73.2 #67 of 81 Archive leaderboard report
Unsupervised Video Object Segmentation DAVIS 2017 (val) MAST F-measure (Mean) 67.6 #4 of 10 Archive leaderboard report
Unsupervised Video Object Segmentation DAVIS 2017 (val) MAST F-measure (Recall) 77.7 #4 of 10 Archive leaderboard report
Unsupervised Video Object Segmentation DAVIS 2017 (val) MAST J&F 65.5 #4 of 10 Archive leaderboard report
Unsupervised Video Object Segmentation DAVIS 2017 (val) MAST Jaccard (Mean) 63.3 #4 of 10 Archive leaderboard report
Unsupervised Video Object Segmentation DAVIS 2017 (val) MAST Jaccard (Recall) 73.2 #4 of 10 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections