{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/how-does-disagreement-help-generalization","title":"How does Disagreement Help Generalization against Label Corruption?","arxiv_id":"1901.04215","date":"2019-01-14","proceeding":null,"authors":["Xingrui Yu","Bo Han","Jiangchao Yao","Gang Niu","Ivor W. Tsang","Masashi Sugiyama"],"abstract":"Learning with noisy labels is one of the hottest problems in weakly-supervised learning. Based on memorization effects of deep neural networks, training on small-loss instances becomes very promising for handling noisy labels. This fosters the state-of-the-art approach \"Co-teaching\" that cross-trains two deep neural networks using the small-loss trick. However, with the increase of epochs, two networks converge to a consensus and Co-teaching reduces to the self-training MentorNet. To tackle this issue, we propose a robust learning paradigm called Co-teaching+, which bridges the \"Update by Disagreement\" strategy with the original Co-teaching. First, two networks feed forward and predict all data, but keep prediction disagreement data only. Then, among such disagreement data, each network selects its small-loss data, but back propagates the small-loss data from its peer network and updates its own parameters. Empirical results on benchmark datasets demonstrate that Co-teaching+ is much superior to many state-of-the-art methods in the robustness of trained models.","url_abs":"https://arxiv.org/abs/1901.04215v3","url_pdf":"https://arxiv.org/pdf/1901.04215v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"how-does-disagreement-help-generalization","repo_url":"https://github.com/bhanML/coteaching_plus","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"how-does-disagreement-help-generalization","repo_url":"https://github.com/xingruiyu/coteaching_plus","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"how-does-disagreement-help-generalization","repo_url":"https://github.com/ziegler-ingo/cleavage_prediction","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"learning-with-noisy-labels","task_name":"Learning with noisy labels"},{"task_slug":"memorization","task_name":"Memorization"},{"task_slug":"weakly-supervised-learning","task_name":"Weakly-supervised Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/learning-with-noisy-labels-on-cifar-100n","task":"Learning with noisy labels","dataset":"CIFAR-100N","model":"Co-Teaching+","rank_in_archive_order":14,"of":24,"metrics":{"Accuracy (mean)":"57.88"},"uses_additional_data":false},{"leaderboard":"/sota/learning-with-noisy-labels-on-cifar-10n","task":"Learning with noisy labels","dataset":"CIFAR-10N-Aggregate","model":"Co-Teaching+","rank_in_archive_order":20,"of":26,"metrics":{"Accuracy (mean)":"90.61"},"uses_additional_data":false},{"leaderboard":"/sota/learning-with-noisy-labels-on-cifar-10n-1","task":"Learning with noisy labels","dataset":"CIFAR-10N-Random1","model":"Co-Teaching+","rank_in_archive_order":16,"of":24,"metrics":{"Accuracy (mean)":"89.70"},"uses_additional_data":false},{"leaderboard":"/sota/learning-with-noisy-labels-on-cifar-10n-2","task":"Learning with noisy labels","dataset":"CIFAR-10N-Random2","model":"Co-Teaching+","rank_in_archive_order":15,"of":23,"metrics":{"Accuracy (mean)":"89.47"},"uses_additional_data":false},{"leaderboard":"/sota/learning-with-noisy-labels-on-cifar-10n-3","task":"Learning with noisy labels","dataset":"CIFAR-10N-Random3","model":"Co-Teaching+","rank_in_archive_order":16,"of":23,"metrics":{"Accuracy (mean)":"89.54"},"uses_additional_data":false},{"leaderboard":"/sota/learning-with-noisy-labels-on-cifar-10n-worst","task":"Learning with noisy labels","dataset":"CIFAR-10N-Worst","model":"Co-Teaching+","rank_in_archive_order":15,"of":25,"metrics":{"Accuracy (mean)":"83.26"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1901.04215","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}