{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/cynical-selection-of-language-model-training","title":"Cynical Selection of Language Model Training Data","arxiv_id":"1709.02279","date":"2017-09-07","proceeding":null,"authors":["Amittai Axelrod"],"abstract":"The Moore-Lewis method of \"intelligent selection of language model training\ndata\" is very effective, cheap, efficient... and also has structural problems.\n(1) The method defines relevance by playing language models trained on the\nin-domain and the out-of-domain (or data pool) corpora against each other. This\npowerful idea-- which we set out to preserve-- treats the two corpora as the\nopposing ends of a single spectrum. This lack of nuance does not allow for the\ntwo corpora to be very similar. In the extreme case where the come from the\nsame distribution, all of the sentences have a Moore-Lewis score of zero, so\nthere is no resulting ranking. (2) The selected sentences are not guaranteed to\nbe able to model the in-domain data, nor to even cover the in-domain data. They\nare simply well-liked by the in-domain model; this is necessary, but not\nsufficient. (3) There is no way to tell what is the optimal number of sentences\nto select, short of picking various thresholds and building the systems.\n  We present a greedy, lazy, approximate, and generally efficient\ninformation-theoretic method of accomplishing the same goal using only\nvocabulary counts. The method has the following properties: (1) Is responsive\nto the extent to which two corpora differ. (2) Quickly reaches near-optimal\nvocabulary coverage. (3) Takes into account what has already been selected. (4)\nDoes not involve defining any kind of domain, nor any kind of classifier. (6)\nKnows approximately when to stop. This method can be used as an\ninherently-meaningful measure of similarity, as it measures the bits of\ninformation to be gained by adding one text to another.","url_abs":"http://arxiv.org/abs/1709.02279v1","url_pdf":"http://arxiv.org/pdf/1709.02279v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"cynical-selection-of-language-model-training","repo_url":"https://github.com/amittai/cynical","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"model","task_name":"model"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1709.02279","atlas_url":"https://app.syntology.ai/?focus=1709.02279","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}