Papers › DP-SSL: Towards Robust Semi-supervised Learning with A Few Labeled Samples

DP-SSL: Towards Robust Semi-supervised Learning with A Few Labeled Samples

26 Oct 2021NeurIPS 2021 12arXiv:2110.13740archive 2025-07-28

Yi Xu, Jiandong Ding, Lu Zhang, Shuigeng Zhou

The scarcity of labeled data is a critical obstacle to deep learning. Semi-supervised learning (SSL) provides a promising way to leverage unlabeled data by pseudo labels. However, when the size of labeled data is very small (say a few labeled samples per class), SSL performs poorly and unstably, possibly due to the low quality of learned pseudo labels. In this paper, we propose a new SSL method called DP-SSL that adopts an innovative data programming (DP) scheme to generate probabilistic labels for unlabeled data. Different from existing DP methods that rely on human experts to provide initial labeling functions (LFs), we develop a multiple-choice learning~(MCL) based approach to automatically generate LFs from scratch in SSL style. With the noisy labels produced by the LFs, we design a label model to resolve the conflict and overlap among the noisy labels, and finally infer probabilistic labels for unlabeled samples. Extensive experiments on four standard SSL benchmarks show that DP-SSL can provide reliable labels for unlabeled data and achieve better classification performance on test sets than existing SSL methods, especially when only a small number of labeled samples are available. Concretely, for CIFAR-10 with only 40 labeled samples, DP-SSL achieves 93.82% annotation accuracy on unlabeled data and 93.46% classification accuracy on test data, which are higher than the SOTA results.

PaperPDFConference PDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Multiple-choiceSemi-Supervised Image Classification

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Semi-Supervised Image Classification CIFAR-10, 250 Labels DP-SSL Percentage error 4.78±0.26 #9 of 27 Archive leaderboard report
Semi-Supervised Image Classification CIFAR-10, 40 Labels DP-SSL Percentage error 6.54±0.98 #13 of 21 Archive leaderboard report
Semi-Supervised Image Classification CIFAR-10, 4000 Labels DP-SSL Percentage error 4.23±0.20 #16 of 49 Archive leaderboard report
Semi-Supervised Image Classification CIFAR-100, 400 Labels DP-SSL Percentage error 43.17±1.29 #16 of 21 Archive leaderboard report
Semi-Supervised Image Classification cifar-100, 10000 Labels DP-SSL Percentage error 22.24±0.31 #15 of 29 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Test

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections