{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/an-evaluation-dataset-for-intent","title":"An Evaluation Dataset for Intent Classification and Out-of-Scope Prediction","arxiv_id":"1909.02027","date":"2019-09-04","proceeding":"IJCNLP 2019 11","authors":["Stefan Larson","Anish Mahendran","Joseph J. Peper","Christopher Clarke","Andrew Lee","Parker Hill","Jonathan K. Kummerfeld","Kevin Leach","Michael A. Laurenzano","Lingjia Tang","Jason Mars"],"abstract":"Task-oriented dialog systems need to know when a query falls outside their range of supported intents, but current text classification corpora only define label sets that cover every example. We introduce a new dataset that includes queries that are out-of-scope---i.e., queries that do not fall into any of the system's supported intents. This poses a new challenge because models cannot assume that every query at inference time belongs to a system-supported intent class. Our dataset also covers 150 intent classes over 10 domains, capturing the breadth that a production task-oriented agent must handle. We evaluate a range of benchmark classifiers on our dataset along with several different out-of-scope identification schemes. We find that while the classifiers perform well on in-scope intent classification, they struggle to identify out-of-scope queries. Our dataset and evaluation fill an important gap in the field, offering a way of more rigorously and realistically benchmarking text classification in task-driven dialog systems.","url_abs":"https://arxiv.org/abs/1909.02027v1","url_pdf":"https://arxiv.org/pdf/1909.02027v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"an-evaluation-dataset-for-intent","repo_url":"https://github.com/clinc/oos-eval","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"NOASSERTION"}},{"paper_slug":"an-evaluation-dataset-for-intent","repo_url":"https://github.com/Abdulk084/Intent_classification","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"an-evaluation-dataset-for-intent","repo_url":"https://github.com/farbodtaymouri/BERT-GAN","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"GPL-3.0"}},{"paper_slug":"an-evaluation-dataset-for-intent","repo_url":"https://github.com/thuiar/textoir","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"an-evaluation-dataset-for-intent","repo_url":"https://github.com/thuiar/textoir-demo","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"GPL-3.0"}}],"tasks":[{"task_slug":"benchmarking","task_name":"Benchmarking"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"intent-classification","task_name":"Intent Classification"},{"task_slug":"text-classification","task_name":"Text Classification"},{"task_slug":"intent-classification-1","task_name":"intent-classification"}],"methods":[],"datasets_introduced":[{"slug":"clinc150","name":"CLINC150","full_name":"CLINC150"}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1909.02027","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}