{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/when-optimizing-f-divergence-is-robust-with-1","title":"When Optimizing $f$-divergence is Robust with Label Noise","arxiv_id":"2011.03687","date":"2020-11-07","proceeding":"ICLR 2021 1","authors":["Jiaheng Wei","Yang Liu"],"abstract":"We show when maximizing a properly defined $f$-divergence measure with respect to a classifier's predictions and the supervised labels is robust with label noise. Leveraging its variational form, we derive a nice decoupling property for a family of $f$-divergence measures when label noise presents, where the divergence is shown to be a linear combination of the variational difference defined on the clean distribution and a bias term introduced due to the noise. The above derivation helps us analyze the robustness of different $f$-divergence functions. With established robustness, this family of $f$-divergence functions arises as useful metrics for the problem of learning with noisy labels, which do not require the specification of the labels' noise rate. When they are possibly not robust, we propose fixes to make them so. In addition to the analytical results, we present thorough experimental evidence. Our code is available at https://github.com/UCSC-REAL/Robust-f-divergence-measures.","url_abs":"https://arxiv.org/abs/2011.03687v3","url_pdf":"https://arxiv.org/pdf/2011.03687v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"when-optimizing-f-divergence-is-robust-with-1","repo_url":"https://github.com/weijiaheng/Robust-f-divergence-measures","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"when-optimizing-f-divergence-is-robust-with-1","repo_url":"https://github.com/weijiaheng/Multi-class-Peer-Loss-functions","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"learning-with-noisy-labels","task_name":"Learning with noisy labels"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/image-classification-on-clothing1m","task":"Image Classification","dataset":"Clothing1M","model":"Robust f-divergence","rank_in_archive_order":35,"of":51,"metrics":{"Accuracy":"73.09%"},"uses_additional_data":false},{"leaderboard":"/sota/learning-with-noisy-labels-on-cifar-100n","task":"Learning with noisy labels","dataset":"CIFAR-100N","model":"F-div","rank_in_archive_order":18,"of":24,"metrics":{"Accuracy (mean)":"57.10"},"uses_additional_data":false},{"leaderboard":"/sota/learning-with-noisy-labels-on-cifar-10n","task":"Learning with noisy labels","dataset":"CIFAR-10N-Aggregate","model":"F-div","rank_in_archive_order":14,"of":26,"metrics":{"Accuracy (mean)":"91.64"},"uses_additional_data":false},{"leaderboard":"/sota/learning-with-noisy-labels-on-cifar-10n-1","task":"Learning with noisy labels","dataset":"CIFAR-10N-Random1","model":"F-div","rank_in_archive_order":17,"of":24,"metrics":{"Accuracy (mean)":"89.70"},"uses_additional_data":false},{"leaderboard":"/sota/learning-with-noisy-labels-on-cifar-10n-2","task":"Learning with noisy labels","dataset":"CIFAR-10N-Random2","model":"F-div","rank_in_archive_order":14,"of":23,"metrics":{"Accuracy (mean)":"89.79"},"uses_additional_data":false},{"leaderboard":"/sota/learning-with-noisy-labels-on-cifar-10n-3","task":"Learning with noisy labels","dataset":"CIFAR-10N-Random3","model":"F-div","rank_in_archive_order":15,"of":23,"metrics":{"Accuracy (mean)":"89.55"},"uses_additional_data":false},{"leaderboard":"/sota/learning-with-noisy-labels-on-cifar-10n-worst","task":"Learning with noisy labels","dataset":"CIFAR-10N-Worst","model":"F-div","rank_in_archive_order":18,"of":25,"metrics":{"Accuracy (mean)":"82.53"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2011.03687","atlas_url":"https://app.syntology.ai/?focus=2011.03687","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}