{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unsupervised-learning-for-class-distribution","title":"Unsupervised Learning for Class Distribution Mismatch","arxiv_id":"2505.06948","date":"2025-05-11","proceeding":null,"authors":["Pan Du","Wangbo Zhao","Xinai Lu","Nian Liu","Zhikai Li","Chaoyu Gong","Suyun Zhao","Hong Chen","Cuiping Li","Kai Wang","Yang You"],"abstract":"Class distribution mismatch (CDM) refers to the discrepancy between class distributions in training data and target tasks. Previous methods address this by designing classifiers to categorize classes known during training, while grouping unknown or new classes into an \"other\" category. However, they focus on semi-supervised scenarios and heavily rely on labeled data, limiting their applicability and performance. To address this, we propose Unsupervised Learning for Class Distribution Mismatch (UCDM), which constructs positive-negative pairs from unlabeled data for classifier training. Our approach randomly samples images and uses a diffusion model to add or erase semantic classes, synthesizing diverse training pairs. Additionally, we introduce a confidence-based labeling mechanism that iteratively assigns pseudo-labels to valuable real-world data and incorporates them into the training process. Extensive experiments on three datasets demonstrate UCDM's superiority over previous semi-supervised methods. Specifically, with a 60% mismatch proportion on Tiny-ImageNet dataset, our approach, without relying on labeled data, surpasses OpenMatch (with 40 labels per class) by 35.1%, 63.7%, and 72.5% in classifying known, unknown, and new classes.","url_abs":"https://arxiv.org/abs/2505.06948v1","url_pdf":"https://arxiv.org/pdf/2505.06948v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"unsupervised-learning-for-class-distribution","repo_url":"https://github.com/ruc-dwbi-ml/research","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[],"methods":[{"method_slug":"diffusion","method_name":"Diffusion"},{"method_slug":"focus","method_name":"Focus"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2505.06948","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2505.06948"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ruc-dwbi-ml/research","reach":null}],"summary":{"ran_honours":1,"ran_draft_wrong":1,"unverified":1},"by_repo_kind":{"official":{"samples":3,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"2dd56a19a0a350a0","entry":"compute_u_loss","repo":"ruc-dwbi-ml/research","repo_kind":"official","path":"UCDM-master/classifier_training/loss.py","file_url":"https://github.com/ruc-dwbi-ml/research/blob/HEAD/UCDM-master/classifier_training/loss.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"2dd56a19a0a350a0"}},{"code_sha256_prefix":"d06825dadf50c9a9","entry":"detect_loss","repo":"ruc-dwbi-ml/research","repo_kind":"official","path":"UCDM-master/classifier_training/loss.py","file_url":"https://github.com/ruc-dwbi-ml/research/blob/HEAD/UCDM-master/classifier_training/loss.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"d06825dadf50c9a9"}},{"code_sha256_prefix":"244c282155d38083","entry":"initialize_distributed","repo":"ruc-dwbi-ml/research","repo_kind":"official","path":"UCDM-master/synthetic_data_pipeline/Dmain.py","file_url":"https://github.com/ruc-dwbi-ml/research/blob/HEAD/UCDM-master/synthetic_data_pipeline/Dmain.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"244c282155d38083"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}