{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/generalized-cross-entropy-loss-for-training","title":"Generalized Cross Entropy Loss for Training Deep Neural Networks with Noisy Labels","arxiv_id":"1805.07836","date":"2018-05-20","proceeding":"NeurIPS 2018 12","authors":["Zhilu Zhang","Mert R. Sabuncu"],"abstract":"Deep neural networks (DNNs) have achieved tremendous success in a variety of\napplications across many disciplines. Yet, their superior performance comes\nwith the expensive cost of requiring correctly annotated large-scale datasets.\nMoreover, due to DNNs' rich capacity, errors in training labels can hamper\nperformance. To combat this problem, mean absolute error (MAE) has recently\nbeen proposed as a noise-robust alternative to the commonly-used categorical\ncross entropy (CCE) loss. However, as we show in this paper, MAE can perform\npoorly with DNNs and challenging datasets. Here, we present a theoretically\ngrounded set of noise-robust loss functions that can be seen as a\ngeneralization of MAE and CCE. Proposed loss functions can be readily applied\nwith any existing DNN architecture and algorithm, while yielding good\nperformance in a wide range of noisy label scenarios. We report results from\nexperiments conducted with CIFAR-10, CIFAR-100 and FASHION-MNIST datasets and\nsynthetically generated noisy labels.","url_abs":"http://arxiv.org/abs/1805.07836v4","url_pdf":"http://arxiv.org/pdf/1805.07836v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"generalized-cross-entropy-loss-for-training","repo_url":"https://github.com/AlanChou/Truncated-Loss","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"generalized-cross-entropy-loss-for-training","repo_url":"https://github.com/arghosh/noisy_label_pretrain","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"generalized-cross-entropy-loss-for-training","repo_url":"https://github.com/awasthiabhijeet/Learning-From-Rules","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"generalized-cross-entropy-loss-for-training","repo_url":"https://github.com/dmizr/phuber","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"learning-with-noisy-labels","task_name":"Learning with noisy labels"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/image-classification-on-clothing1m","task":"Image Classification","dataset":"Clothing1M","model":"GCE","rank_in_archive_order":49,"of":51,"metrics":{"Accuracy":"69.75%"},"uses_additional_data":false},{"leaderboard":"/sota/learning-with-noisy-labels-on-cifar-100n","task":"Learning with noisy labels","dataset":"CIFAR-100N","model":"GCE","rank_in_archive_order":20,"of":24,"metrics":{"Accuracy (mean)":"56.73"},"uses_additional_data":false},{"leaderboard":"/sota/learning-with-noisy-labels-on-cifar-10n","task":"Learning with noisy labels","dataset":"CIFAR-10N-Aggregate","model":"GCE","rank_in_archive_order":25,"of":26,"metrics":{"Accuracy (mean)":"87.85"},"uses_additional_data":false},{"leaderboard":"/sota/learning-with-noisy-labels-on-cifar-10n-1","task":"Learning with noisy labels","dataset":"CIFAR-10N-Random1","model":"GCE","rank_in_archive_order":22,"of":24,"metrics":{"Accuracy (mean)":"87.61"},"uses_additional_data":false},{"leaderboard":"/sota/learning-with-noisy-labels-on-cifar-10n-2","task":"Learning with noisy labels","dataset":"CIFAR-10N-Random2","model":"GCE","rank_in_archive_order":20,"of":23,"metrics":{"Accuracy (mean)":"87.70"},"uses_additional_data":false},{"leaderboard":"/sota/learning-with-noisy-labels-on-cifar-10n-3","task":"Learning with noisy labels","dataset":"CIFAR-10N-Random3","model":"GCE","rank_in_archive_order":20,"of":23,"metrics":{"Accuracy (mean)":"87.58"},"uses_additional_data":false},{"leaderboard":"/sota/learning-with-noisy-labels-on-cifar-10n-worst","task":"Learning with noisy labels","dataset":"CIFAR-10N-Worst","model":"GCE","rank_in_archive_order":20,"of":25,"metrics":{"Accuracy (mean)":"80.66"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1805.07836","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}