{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/benchmarking-natural-language-understanding","title":"Benchmarking Natural Language Understanding Services for building Conversational Agents","arxiv_id":"1903.05566","date":"2019-03-13","proceeding":null,"authors":["Xingkun Liu","Arash Eshghi","Pawel Swietojanski","Verena Rieser"],"abstract":"We have recently seen the emergence of several publicly available Natural\nLanguage Understanding (NLU) toolkits, which map user utterances to structured,\nbut more abstract, Dialogue Act (DA) or Intent specifications, while making\nthis process accessible to the lay developer. In this paper, we present the\nfirst wide coverage evaluation and comparison of some of the most popular NLU\nservices, on a large, multi-domain (21 domains) dataset of 25K user utterances\nthat we have collected and annotated with Intent and Entity Type specifications\nand which will be released as part of this submission. The results show that on\nIntent classification Watson significantly outperforms the other platforms,\nnamely, Dialogflow, LUIS and Rasa; though these also perform well.\nInterestingly, on Entity Type recognition, Watson performs significantly worse\ndue to its low Precision. Again, Dialogflow, LUIS and Rasa perform well on this\ntask.","url_abs":"http://arxiv.org/abs/1903.05566v3","url_pdf":"http://arxiv.org/pdf/1903.05566v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"benchmarking-natural-language-understanding","repo_url":"https://github.com/xliuhw/NLU-Evaluation-Data","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"CC-BY-4.0"}},{"paper_slug":"benchmarking-natural-language-understanding","repo_url":"https://github.com/Lackel/DNA","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"benchmarking-natural-language-understanding","repo_url":"https://github.com/PolyAI-LDN/polyai-models","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"benchmarking-natural-language-understanding","repo_url":"https://github.com/PolyAI-LDN/task-specific-datasets","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"CC-BY-4.0"}},{"paper_slug":"benchmarking-natural-language-understanding","repo_url":"https://github.com/amazon-science/intent-aware-encoder","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"benchmarking-natural-language-understanding","repo_url":"https://github.com/haodeqi/BenchmarkingIntentDetection","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}},{"paper_slug":"benchmarking-natural-language-understanding","repo_url":"https://github.com/lackel/hierarchical_weighted_scl","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"benchmarking-natural-language-understanding","repo_url":"https://github.com/lackel/sdc","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"benchmarking-natural-language-understanding","repo_url":"https://github.com/RasaHQ/rasa","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"benchmarking","task_name":"Benchmarking"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"intent-classification","task_name":"Intent Classification"},{"task_slug":"natural-language-understanding","task_name":"Natural Language Understanding"},{"task_slug":"intent-classification-1","task_name":"intent-classification"}],"methods":[],"datasets_introduced":[{"slug":"hwu64","name":"HWU64","full_name":"HWU64"}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1903.05566","atlas_url":"https://app.syntology.ai/?focus=1903.05566","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}