{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/herolt-benchmarking-heterogeneous-long-tailed","title":"Towards Heterogeneous Long-tailed Learning: Benchmarking, Metrics, and Toolbox","arxiv_id":"2307.08235","date":"2023-07-17","proceeding":null,"authors":["Haohui Wang","Weijie Guan","Jianpeng Chen","Zi Wang","Dawei Zhou"],"abstract":"Long-tailed data distributions pose challenges for a variety of domains like e-commerce, finance, biomedical science, and cyber security, where the performance of machine learning models is often dominated by head categories while tail categories are inadequately learned. This work aims to provide a systematic view of long-tailed learning with regard to three pivotal angles: (A1) the characterization of data long-tailedness, (A2) the data complexity of various domains, and (A3) the heterogeneity of emerging tasks. We develop HeroLT, a comprehensive long-tailed learning benchmark integrating 18 state-of-the-art algorithms, 10 evaluation metrics, and 17 real-world datasets across 6 tasks and 4 data modalities. HeroLT with novel angles and extensive experiments (315 in total) enables effective and fair evaluation of newly proposed methods compared with existing baselines on varying dataset types. Finally, we conclude by highlighting the significant applications of long-tailed learning and identifying several promising future directions. For accessibility and reproducibility, we open-source our benchmark HeroLT and corresponding results at https://github.com/SSSKJ/HeroLT.","url_abs":"https://arxiv.org/abs/2307.08235v2","url_pdf":"https://arxiv.org/pdf/2307.08235v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"herolt-benchmarking-heterogeneous-long-tailed","repo_url":"https://github.com/ssskj/herolt","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"benchmarking","task_name":"Benchmarking"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2307.08235","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2307.08235"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/SSSKJ/HeroLT","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ssskj/herolt","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran":3,"unverified":1},"by_repo_kind":{"official":{"samples":4,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"9dc354a7407a75b8","entry":"balanced_softmax_loss","repo":"SSSKJ/HeroLT","repo_kind":"official","path":"HeroLT/nn/Loss/BalancedSoftmaxLoss.py","file_url":"https://github.com/SSSKJ/HeroLT/blob/HEAD/HeroLT/nn/Loss/BalancedSoftmaxLoss.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"9dc354a7407a75b8"}},{"code_sha256_prefix":"2c5959a4a4cf5260","entry":"create_loss","repo":"SSSKJ/HeroLT","repo_kind":"official","path":"HeroLT/nn/Loss/BalancedSoftmaxLoss.py","file_url":"https://github.com/SSSKJ/HeroLT/blob/HEAD/HeroLT/nn/Loss/BalancedSoftmaxLoss.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"2c5959a4a4cf5260"}},{"code_sha256_prefix":"756b76bed831be5a","entry":"create_loss","repo":"SSSKJ/HeroLT","repo_kind":"official","path":"HeroLT/nn/Loss/DiscCentroidsLoss.py","file_url":"https://github.com/SSSKJ/HeroLT/blob/HEAD/HeroLT/nn/Loss/DiscCentroidsLoss.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"756b76bed831be5a"}},{"code_sha256_prefix":"247a09b49a2abf27","entry":"create_loss","repo":"SSSKJ/HeroLT","repo_kind":"official","path":"HeroLT/nn/Loss/SoftmaxLoss.py","file_url":"https://github.com/SSSKJ/HeroLT/blob/HEAD/HeroLT/nn/Loss/SoftmaxLoss.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"247a09b49a2abf27"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}