Papers › Benchmarking Natural Language Understanding Services for building Conversational Agents

Benchmarking Natural Language Understanding Services for building Conversational Agents

13 Mar 2019arXiv:1903.05566archive 2025-07-28

Xingkun Liu, Arash Eshghi, Pawel Swietojanski, Verena Rieser

We have recently seen the emergence of several publicly available Natural Language Understanding (NLU) toolkits, which map user utterances to structured, but more abstract, Dialogue Act (DA) or Intent specifications, while making this process accessible to the lay developer. In this paper, we present the first wide coverage evaluation and comparison of some of the most popular NLU services, on a large, multi-domain (21 domains) dataset of 25K user utterances that we have collected and annotated with Intent and Entity Type specifications and which will be released as part of this submission. The results show that on Intent classification Watson significantly outperforms the other platforms, namely, Dialogflow, LUIS and Rasa; though these also perform well. Interestingly, on Entity Type recognition, Watson performs significantly worse due to its low Precision. Again, Dialogflow, LUIS and Rasa perform well on this task.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

xliuhw/NLU-Evaluation-Data officialmentioned in papermentioned on GitHubCC-BY-4.0 report
Lackel/DNA mentioned on GitHubpytorch report
PolyAI-LDN/polyai-models mentioned on GitHubtf report
PolyAI-LDN/task-specific-datasets mentioned on GitHubCC-BY-4.0 report
amazon-science/intent-aware-encoder mentioned on GitHubpytorchApache-2.0 report
lackel/hierarchical_weighted_scl mentioned on GitHubpytorch report
lackel/sdc mentioned on GitHubpytorch report
RasaHQ/rasa Apache-2.0 report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

BenchmarkingGeneral ClassificationIntent ClassificationNatural Language Understandingintent-classification

Datasets

Introduced by this paper, per the archive.

HWU64

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections