{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/to-share-or-not-to-share-predicting-sets-of","title":"To Share or not to Share: Predicting Sets of Sources for Model Transfer Learning","arxiv_id":"2104.08078","date":"2021-04-16","proceeding":"EMNLP 2021 11","authors":["Lukas Lange","Jannik Strötgen","Heike Adel","Dietrich Klakow"],"abstract":"In low-resource settings, model transfer can help to overcome a lack of labeled data for many tasks and domains. However, predicting useful transfer sources is a challenging problem, as even the most similar sources might lead to unexpected negative transfer results. Thus, ranking methods based on task and text similarity -- as suggested in prior work -- may not be sufficient to identify promising sources. To tackle this problem, we propose a new approach to automatically determine which and how many sources should be exploited. For this, we study the effects of model transfer on sequence labeling across various domains and tasks and show that our methods based on model similarity and support vector machines are able to predict promising sources, resulting in performance increases of up to 24 F1 points.","url_abs":"https://arxiv.org/abs/2104.08078v2","url_pdf":"https://arxiv.org/pdf/2104.08078v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"to-share-or-not-to-share-predicting-sets-of","repo_url":"https://github.com/boschresearch/predicting_sets_of_sources","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"transfer-learning","task_name":"Transfer Learning"},{"task_slug":"text-similarity","task_name":"text similarity"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2104.08078","atlas_url":"https://app.syntology.ai/?focus=2104.08078","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2104.08078"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/boschresearch/predicting_sets_of_sources","reach":null}],"summary":{"unverified":4},"by_repo_kind":{"official":{"samples":4,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":4,"samples":[{"code_sha256_prefix":"5058b431a8e082dd","entry":"AbstractRanking","repo":"boschresearch/predicting_sets_of_sources","repo_kind":"official","path":"src/predictors.py","file_url":"https://github.com/boschresearch/predicting_sets_of_sources/blob/HEAD/src/predictors.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"AGPL-3.0","inline_ok":false,"mcp_get_code":{"code_sha256":"5058b431a8e082dd"}},{"code_sha256_prefix":"52437fee2644b775","entry":"Predictor","repo":"boschresearch/predicting_sets_of_sources","repo_kind":"official","path":"src/predictors.py","file_url":"https://github.com/boschresearch/predicting_sets_of_sources/blob/HEAD/src/predictors.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"AGPL-3.0","inline_ok":false,"mcp_get_code":{"code_sha256":"52437fee2644b775"}},{"code_sha256_prefix":"cd1c2fdd54600ca6","entry":"Ranking","repo":"boschresearch/predicting_sets_of_sources","repo_kind":"official","path":"src/predictors.py","file_url":"https://github.com/boschresearch/predicting_sets_of_sources/blob/HEAD/src/predictors.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"AGPL-3.0","inline_ok":false,"mcp_get_code":{"code_sha256":"cd1c2fdd54600ca6"}},{"code_sha256_prefix":"723ed505f0bada5e","entry":"TopDynamic","repo":"boschresearch/predicting_sets_of_sources","repo_kind":"official","path":"src/predictors.py","file_url":"https://github.com/boschresearch/predicting_sets_of_sources/blob/HEAD/src/predictors.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"AGPL-3.0","inline_ok":false,"mcp_get_code":{"code_sha256":"723ed505f0bada5e"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}