Papers › Boosting the Performance of Semi-Supervised Learning with Unsupervised Clustering

Boosting the Performance of Semi-Supervised Learning with Unsupervised Clustering

1 Dec 2020arXiv:2012.00504archive 2025-07-28

Boaz Lerner, Guy Shiran, Daphna Weinshall

Recently, Semi-Supervised Learning (SSL) has shown much promise in leveraging unlabeled data while being provided with very few labels. In this paper, we show that ignoring the labels altogether for whole epochs intermittently during training can significantly improve performance in the small sample regime. More specifically, we propose to train a network on two tasks jointly. The primary classification task is exposed to both the unlabeled and the scarcely annotated data, whereas the secondary task seeks to cluster the data without any labels. As opposed to hand-crafted pretext tasks frequently used in self-supervision, our clustering phase utilizes the same classification network and head in an attempt to relax the primary task and propagate the information from the labels without overfitting them. On top of that, the self-supervised technique of classifying image rotations is incorporated during the unsupervised learning phase to stabilize training. We demonstrate our method's efficacy in boosting several state-of-the-art SSL algorithms, significantly improving their results and reducing running time in various standard semi-supervised benchmarks, including 92.6% accuracy on CIFAR-10 and 96.9% on SVHN, using only 4 labels per class in each task. We also notably improve the results in the extreme cases of 1,2 and 3 labels per class, and show that features learned by our model are more meaningful for separating the data.

PaperPDFCode

In Syntology View this paper on Syntology: its repositories, every harvested function with whether it ran, its licence and the call to fetch it.

Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

boazlern/SSClustering officialmentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

ClusteringSemi-Supervised Image Classification

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Semi-Supervised Image Classification CIFAR-10, 20 Labels Semi-MMDC Percentage error 28.1±5.5 #3 of 3 Archive leaderboard report
Semi-Supervised Image Classification CIFAR-10, 250 Labels Semi-MMDC Percentage error 5.51±0.25 #17 of 27 Archive leaderboard report
Semi-Supervised Image Classification CIFAR-10, 40 Labels Semi-MMDC Percentage error 7.39±0.61 #16 of 21 Archive leaderboard report
Semi-Supervised Image Classification STL-10, 1000 Labels Semi-MMDC Accuracy 95.22±0.29 #5 of 13 Archive leaderboard report
Semi-Supervised Image Classification SVHN, 250 Labels Semi-MMDC Accuracy 97.7±0.03 #2 of 15 Archive leaderboard report
Semi-Supervised Image Classification SVHN, 40 Labels Semi-MMDC Percentage error 3.09±0.54 #2 of 5 Archive leaderboard report
Semi-Supervised Image Classification cifar-10, 10 Labels Semi-MMDC Accuracy (Test) 70.84±8.1 #3 of 3 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections