Papers › The Re-Label Method For Data-Centric Machine Learning

The Re-Label Method For Data-Centric Machine Learning

9 Feb 2023arXiv:2302.04391archive 2025-07-28

Tong Guo

In industry deep learning application, our manually labeled data has a certain number of noisy data. To solve this problem and achieve more than 90 score in dev dataset, we present a simple method to find the noisy data and re-label the noisy data by human, given the model predictions as references in human labeling. In this paper, we illustrate our idea for a broad set of deep learning tasks, includes classification, sequence tagging, object detection, sequence generation, click-through rate prediction. The dev dataset evaluation results and human evaluation results verify our idea.

PaperPDFCode

Code

guotong1988/Automatic-Label-Error-Correction officialmentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Click-Through Rate PredictionDeep LearningLabel Error DetectionObject DetectionText Classificationobject-detection

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Label Error Detection TREC-6 github.com/guotong1988/Automatic-Label-Error-Correction Accuracy 99.0 #1 of 1 Archive leaderboard report
Text Classification TREC-6 Automatic Label Error Correction Error 0.40 #1 of 19 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections