{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-with-instance-dependent-label-noise-1","title":"Learning with Instance-Dependent Label Noise: A Sample Sieve Approach","arxiv_id":"2010.02347","date":"2020-10-05","proceeding":"ICLR 2021 1","authors":["Hao Cheng","Zhaowei Zhu","Xingyu Li","Yifei Gong","Xing Sun","Yang Liu"],"abstract":"Human-annotated labels are often prone to noise, and the presence of such noise will degrade the performance of the resulting deep neural network (DNN) models. Much of the literature (with several recent exceptions) of learning with noisy labels focuses on the case when the label noise is independent of features. Practically, annotations errors tend to be instance-dependent and often depend on the difficulty levels of recognizing a certain task. Applying existing results from instance-independent settings would require a significant amount of estimation of noise rates. Therefore, providing theoretically rigorous solutions for learning with instance-dependent label noise remains a challenge. In this paper, we propose CORES$^{2}$ (COnfidence REgularized Sample Sieve), which progressively sieves out corrupted examples. The implementation of CORES$^{2}$ does not require specifying noise rates and yet we are able to provide theoretical guarantees of CORES$^{2}$ in filtering out the corrupted examples. This high-quality sample sieve allows us to treat clean examples and the corrupted ones separately in training a DNN solution, and such a separation is shown to be advantageous in the instance-dependent noise setting. We demonstrate the performance of CORES$^{2}$ on CIFAR10 and CIFAR100 datasets with synthetic instance-dependent label noise and Clothing1M with real-world human noise. As of independent interests, our sample sieve provides a generic machinery for anatomizing noisy datasets and provides a flexible interface for various robust training techniques to further improve the performance. Code is available at https://github.com/UCSC-REAL/cores.","url_abs":"https://arxiv.org/abs/2010.02347v2","url_pdf":"https://arxiv.org/pdf/2010.02347v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-with-instance-dependent-label-noise-1","repo_url":"https://github.com/UCSC-REAL/cores","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"image-classification-with-label-noise","task_name":"Image Classification with Label Noise"},{"task_slug":"learning-with-noisy-labels","task_name":"Learning with noisy labels"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/image-classification-on-clothing1m","task":"Image Classification","dataset":"Clothing1M","model":"CORES2","rank_in_archive_order":33,"of":51,"metrics":{"Accuracy":"73.24%"},"uses_additional_data":false},{"leaderboard":"/sota/learning-with-noisy-labels-on-cifar-100n","task":"Learning with noisy labels","dataset":"CIFAR-100N","model":"CORES","rank_in_archive_order":9,"of":24,"metrics":{"Accuracy (mean)":"61.15"},"uses_additional_data":false},{"leaderboard":"/sota/learning-with-noisy-labels-on-cifar-100n","task":"Learning with noisy labels","dataset":"CIFAR-100N","model":"CORES*","rank_in_archive_order":22,"of":24,"metrics":{"Accuracy (mean)":"55.72"},"uses_additional_data":false},{"leaderboard":"/sota/learning-with-noisy-labels-on-cifar-10n","task":"Learning with noisy labels","dataset":"CIFAR-10N-Aggregate","model":"CORES*","rank_in_archive_order":6,"of":26,"metrics":{"Accuracy (mean)":"95.25"},"uses_additional_data":false},{"leaderboard":"/sota/learning-with-noisy-labels-on-cifar-10n","task":"Learning with noisy labels","dataset":"CIFAR-10N-Aggregate","model":"CORES","rank_in_archive_order":17,"of":26,"metrics":{"Accuracy (mean)":"91.23"},"uses_additional_data":false},{"leaderboard":"/sota/learning-with-noisy-labels-on-cifar-10n-1","task":"Learning with noisy labels","dataset":"CIFAR-10N-Random1","model":"CORES*","rank_in_archive_order":6,"of":24,"metrics":{"Accuracy (mean)":"94.45"},"uses_additional_data":false},{"leaderboard":"/sota/learning-with-noisy-labels-on-cifar-10n-1","task":"Learning with noisy labels","dataset":"CIFAR-10N-Random1","model":"CORES","rank_in_archive_order":18,"of":24,"metrics":{"Accuracy (mean)":"89.66"},"uses_additional_data":false},{"leaderboard":"/sota/learning-with-noisy-labels-on-cifar-10n-2","task":"Learning with noisy labels","dataset":"CIFAR-10N-Random2","model":"CORES*","rank_in_archive_order":4,"of":23,"metrics":{"Accuracy (mean)":"94.88"},"uses_additional_data":false},{"leaderboard":"/sota/learning-with-noisy-labels-on-cifar-10n-2","task":"Learning with noisy labels","dataset":"CIFAR-10N-Random2","model":"CORES","rank_in_archive_order":13,"of":23,"metrics":{"Accuracy (mean)":"89.91"},"uses_additional_data":false},{"leaderboard":"/sota/learning-with-noisy-labels-on-cifar-10n-3","task":"Learning with noisy labels","dataset":"CIFAR-10N-Random3","model":"CORES*","rank_in_archive_order":4,"of":23,"metrics":{"Accuracy (mean)":"94.74"},"uses_additional_data":false},{"leaderboard":"/sota/learning-with-noisy-labels-on-cifar-10n-3","task":"Learning with noisy labels","dataset":"CIFAR-10N-Random3","model":"CORES","rank_in_archive_order":14,"of":23,"metrics":{"Accuracy (mean)":"89.79"},"uses_additional_data":false},{"leaderboard":"/sota/learning-with-noisy-labels-on-cifar-10n-worst","task":"Learning with noisy labels","dataset":"CIFAR-10N-Worst","model":"CORES*","rank_in_archive_order":7,"of":25,"metrics":{"Accuracy (mean)":"91.66"},"uses_additional_data":false},{"leaderboard":"/sota/learning-with-noisy-labels-on-cifar-10n-worst","task":"Learning with noisy labels","dataset":"CIFAR-10N-Worst","model":"CORES","rank_in_archive_order":12,"of":25,"metrics":{"Accuracy (mean)":"83.60"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2010.02347","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2010.02347"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/UCSC-REAL/cores","reach":null}],"summary":{"ran_honours":1,"ran_violates":1,"unverified":2},"by_repo_kind":{"official":{"samples":4,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":4,"samples":[{"code_sha256_prefix":"66e2df8b49609716","entry":"f_beta","repo":"UCSC-REAL/cores","repo_kind":"official","path":"loss.py","file_url":"https://github.com/UCSC-REAL/cores/blob/HEAD/loss.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"66e2df8b49609716"}},{"code_sha256_prefix":"08c40c52150e5ee8","entry":"get_noise_pred","repo":"ucsc-real/cores","repo_kind":"official","path":"phase1.py","file_url":"https://github.com/ucsc-real/cores/blob/HEAD/phase1.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"08c40c52150e5ee8"}},{"code_sha256_prefix":"12130d8056932840","entry":"accuracy","repo":"ucsc-real/cores","repo_kind":"official","path":"phase1.py","file_url":"https://github.com/ucsc-real/cores/blob/HEAD/phase1.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"12130d8056932840"}},{"code_sha256_prefix":"b703cc7763eaf4e4","entry":"loss_cores","repo":"UCSC-REAL/cores","repo_kind":"official","path":"loss.py","file_url":"https://github.com/UCSC-REAL/cores/blob/HEAD/loss.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"b703cc7763eaf4e4"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}