{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/data-selection-strategies-for-multi-domain","title":"Data Selection Strategies for Multi-Domain Sentiment Analysis","arxiv_id":"1702.02426","date":"2017-02-08","proceeding":null,"authors":["Sebastian Ruder","Parsa Ghaffari","John G. Breslin"],"abstract":"Domain adaptation is important in sentiment analysis as sentiment-indicating\nwords vary between domains. Recently, multi-domain adaptation has become more\npervasive, but existing approaches train on all available source domains\nincluding dissimilar ones. However, the selection of appropriate training data\nis as important as the choice of algorithm. We undertake -- to our knowledge\nfor the first time -- an extensive study of domain similarity metrics in the\ncontext of sentiment analysis and propose novel representations, metrics, and a\nnew scope for data selection. We evaluate the proposed methods on two\nlarge-scale multi-domain adaptation settings on tweets and reviews and\ndemonstrate that they consistently outperform strong random and balanced\nbaselines, while our proposed selection strategy outperforms instance-level\nselection and yields the best score on a large reviews corpus.","url_abs":"http://arxiv.org/abs/1702.02426v1","url_pdf":"http://arxiv.org/pdf/1702.02426v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"data-selection-strategies-for-multi-domain","repo_url":"https://github.com/andy-yangz/writing_style_transfer","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"domain-adaptation","task_name":"Domain Adaptation"},{"task_slug":"sentiment-analysis","task_name":"Sentiment Analysis"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1702.02426","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}