{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/on-the-effect-of-low-frequency-terms-on","title":"On the Effect of Low-Frequency Terms on Neural-IR Models","arxiv_id":"1904.12683","date":"2019-04-29","proceeding":null,"authors":["Sebastian Hofstätter","Navid Rekabsaz","Carsten Eickhoff","Allan Hanbury"],"abstract":"Low-frequency terms are a recurring challenge for information retrieval\nmodels, especially neural IR frameworks struggle with adequately capturing\ninfrequently observed words. While these terms are often removed from neural\nmodels - mainly as a concession to efficiency demands - they traditionally play\nan important role in the performance of IR models. In this paper, we analyze\nthe effects of low-frequency terms on the performance and robustness of neural\nIR models. We conduct controlled experiments on three recent neural IR models,\ntrained on a large-scale passage retrieval collection. We evaluate the neural\nIR models with various vocabulary sizes for their respective word embeddings,\nconsidering different levels of constraints on the available GPU memory. We\nobserve that despite the significant benefits of using larger vocabularies, the\nperformance gap between the vocabularies can be, to a great extent, mitigated\nby extensive tuning of a related parameter: the number of documents to re-rank.\nWe further investigate the use of subword-token embedding models, and in\nparticular FastText, for neural IR models. Our experiments show that using\nFastText brings slight improvements to the overall performance of the neural IR\nmodels in comparison to models trained on the full vocabulary, while the\nimprovement becomes much more pronounced for queries containing low-frequency\nterms.","url_abs":"http://arxiv.org/abs/1904.12683v2","url_pdf":"http://arxiv.org/pdf/1904.12683v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"on-the-effect-of-low-frequency-terms-on","repo_url":"https://github.com/sebastian-hofstaetter/sigir19-neural-ir","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":null,"task_name":"GPU"},{"task_slug":"information-retrieval","task_name":"Information Retrieval"},{"task_slug":"passage-retrieval","task_name":"Passage Retrieval"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"word-embeddings","task_name":"Word Embeddings"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1904.12683","atlas_url":"https://app.syntology.ai/?focus=1904.12683","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}