{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2606-05927","title":"Addressing Imbalance in Multi-Label Data via Label-Specific Distance-based Oversampling","arxiv_id":"2606.05927","date":"2026-06-04","proceeding":null,"authors":["Bin Liu","Jun Wu","Haoyu Peng","Ao Zhou","Jin Wang","QiaoSong Chen","Grigorios Tsoumakas"],"abstract":"The complex imbalanced label distribution poses a crucial challenge to multi-label classification, as most classifiers are biased towards the majority class and high-frequent labels. Oversampling is an efficient and flexible solution that augments instances to provide a more balanced training dataset for multi-label classifiers. Most existing oversampling methods create synthetic instances in a heuristic way that essentially relies on neighborhood information retrieved using Euclidean distance within the entire feature space. However, they fail to consider the varying semantic relevance of features to different labels, leading to label inconsistency among proximate neighbors and further introducing label confusion and overfitting to synthetic instances. To overcome the above issue, we propose a novel sampling approach called Label-Specific Distance-based Multi-Label Oversampling (LSDMLO) that creates more useful and well-labeled synthetic instances to address the imbalance in multi-label datasets. LSDMLO derives the label-specific distance to identify label-consistent neighbors based on the weighted pertinent feature space, which facilitates selecting seed instances that express more label correlations in boundary areas and generating synthetic instances aligned with the label distribution of original data. The comprehensive experiments verify that the proposed LSDMLO outperforms the state-of-the-art multi-label sampling approaches under various base classifiers.","url_abs":"https://arxiv.org/abs/2606.05927","url_pdf":"https://arxiv.org/pdf/2606.05927","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2606.05927","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2606.05927"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/CquptZA/LSDMLO-PR-","reach":null}],"summary":{"unverified":6},"by_repo_kind":{"found_in_text":{"samples":6,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":6,"samples":[{"code_sha256_prefix":"4ed417e4fee4b9f7","entry":"Distance","repo":"CquptZA/LSDMLO-PR-","repo_kind":"found_in_text","path":"Sampling.py","file_url":"https://github.com/CquptZA/LSDMLO-PR-/blob/HEAD/Sampling.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"4ed417e4fee4b9f7"}},{"code_sha256_prefix":"ede6c10b51045339","entry":"LSDMLOsampling","repo":"CquptZA/LSDMLO-PR-","repo_kind":"found_in_text","path":"Sampling.py","file_url":"https://github.com/CquptZA/LSDMLO-PR-/blob/HEAD/Sampling.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"ede6c10b51045339"}},{"code_sha256_prefix":"a74bb0817bd5a798","entry":"Labeltype","repo":"CquptZA/LSDMLO-PR-","repo_kind":"found_in_text","path":"Sampling.py","file_url":"https://github.com/CquptZA/LSDMLO-PR-/blob/HEAD/Sampling.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"a74bb0817bd5a798"}},{"code_sha256_prefix":"b47c75341a3eada9","entry":"assign_weights","repo":"CquptZA/LSDMLO-PR-","repo_kind":"found_in_text","path":"Sampling.py","file_url":"https://github.com/CquptZA/LSDMLO-PR-/blob/HEAD/Sampling.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"b47c75341a3eada9"}},{"code_sha256_prefix":"bf61b2fd759c376f","entry":"label_assign_distance","repo":"CquptZA/LSDMLO-PR-","repo_kind":"found_in_text","path":"Sampling.py","file_url":"https://github.com/CquptZA/LSDMLO-PR-/blob/HEAD/Sampling.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"bf61b2fd759c376f"}},{"code_sha256_prefix":"9932b9c77285c184","entry":"label_similarity","repo":"CquptZA/LSDMLO-PR-","repo_kind":"found_in_text","path":"Sampling.py","file_url":"https://github.com/CquptZA/LSDMLO-PR-/blob/HEAD/Sampling.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"9932b9c77285c184"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.LG","source":"arxiv_2026.jsonl"},"syntology_extracted_results":null}