{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-clip-powered-framework-for-robust-and","title":"A CLIP-Powered Framework for Robust and Generalizable Data Selection","arxiv_id":"2410.11215","date":"2024-10-15","proceeding":null,"authors":["Suorong Yang","Peng Ye","Wanli Ouyang","Dongzhan Zhou","Furao Shen"],"abstract":"Large-scale datasets have been pivotal to the advancements of deep learning models in recent years, but training on such large datasets invariably incurs substantial storage and computational overhead. Meanwhile, real-world datasets often contain redundant and noisy data, imposing a negative impact on training efficiency and model performance. Data selection has shown promise in identifying the most representative samples from the entire dataset, which aims to minimize the performance gap with reduced training costs. Existing works typically rely on single-modality information to assign importance scores for individual samples, which may lead to inaccurate assessments, especially when dealing with noisy or corrupted samples. To address this limitation, we propose a novel CLIP-powered data selection framework that leverages multimodal information for more robust and generalizable sample selection. Specifically, our framework consists of three key modules-dataset adaptation, sample scoring, and selection optimization-that together harness extensive pre-trained multimodal knowledge to comprehensively assess sample influence and optimize the selection results through multi-objective optimization. Extensive experiments demonstrate that our approach consistently outperforms existing state-of-the-art baselines on various benchmark datasets. Notably, our method effectively removes noisy or damaged samples from the dataset, enabling it to achieve even higher performance with less data. This indicates that it is not only a way to accelerate training but can also improve overall data quality.","url_abs":"https://arxiv.org/abs/2410.11215v1","url_pdf":"https://arxiv.org/pdf/2410.11215v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2410.11215","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2410.11215"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/Jackbrocp/clip-powered-data-selection","reach":{"status":"ok"}}],"summary":{"ran":5},"by_repo_kind":{"found_in_text":{"samples":5,"ran":5,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":5,"samples":[{"code_sha256_prefix":"8dd8c30ee9fbcd8d","entry":"has_file_allowed_extension","repo":"Jackbrocp/clip-powered-data-selection","repo_kind":"found_in_text","path":"dataset.py","file_url":"https://github.com/Jackbrocp/clip-powered-data-selection/blob/HEAD/dataset.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"8dd8c30ee9fbcd8d"}},{"code_sha256_prefix":"13976a63c7d68a7f","entry":"is_image_file","repo":"Jackbrocp/clip-powered-data-selection","repo_kind":"found_in_text","path":"dataset.py","file_url":"https://github.com/Jackbrocp/clip-powered-data-selection/blob/HEAD/dataset.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"13976a63c7d68a7f"}},{"code_sha256_prefix":"6b96ab7b8bc2a0e4","entry":"make_dataset","repo":"Jackbrocp/clip-powered-data-selection","repo_kind":"found_in_text","path":"dataset.py","file_url":"https://github.com/Jackbrocp/clip-powered-data-selection/blob/HEAD/dataset.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"6b96ab7b8bc2a0e4"}},{"code_sha256_prefix":"b4018a8b149ef777","entry":"obtain_classnames","repo":"Jackbrocp/clip-powered-data-selection","repo_kind":"found_in_text","path":"utils.py","file_url":"https://github.com/Jackbrocp/clip-powered-data-selection/blob/HEAD/utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"b4018a8b149ef777"}},{"code_sha256_prefix":"3fe433f932c428b5","entry":"pdist_torch","repo":"Jackbrocp/clip-powered-data-selection","repo_kind":"found_in_text","path":"optimize_selection.py","file_url":"https://github.com/Jackbrocp/clip-powered-data-selection/blob/HEAD/optimize_selection.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"3fe433f932c428b5"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}