{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/milkqa-a-dataset-of-consumer-questions-for","title":"MilkQA: a Dataset of Consumer Questions for the Task of Answer Selection","arxiv_id":"1801.03460","date":"2018-01-10","proceeding":null,"authors":["Marcelo Criscuolo","Erick Rocha Fonseca","Sandra Maria Aluísio","Ana Carolina Sperança-Criscuolo"],"abstract":"We introduce MilkQA, a question answering dataset from the dairy domain\ndedicated to the study of consumer questions. The dataset contains 2,657 pairs\nof questions and answers, written in the Portuguese language and originally\ncollected by the Brazilian Agricultural Research Corporation (Embrapa). All\nquestions were motivated by real situations and written by thousands of authors\nwith very different backgrounds and levels of literacy, while answers were\nelaborated by specialists from Embrapa's customer service. Our dataset was\nfiltered and anonymized by three human annotators. Consumer questions are a\nchallenging kind of question that is usually employed as a form of seeking\ninformation. Although several question answering datasets are available, most\nof such resources are not suitable for research on answer selection models for\nconsumer questions. We aim to fill this gap by making MilkQA publicly\navailable. We study the behavior of four answer selection models on MilkQA: two\nbaseline models and two convolutional neural network archictetures. Our results\nshow that MilkQA poses real challenges to computational models, particularly\ndue to linguistic characteristics of its questions and to their unusually\nlonger lengths. Only one of the experimented models gives reasonable results,\nat the cost of high computational requirements.","url_abs":"http://arxiv.org/abs/1801.03460v1","url_pdf":"http://arxiv.org/pdf/1801.03460v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"answer-selection","task_name":"Answer Selection"},{"task_slug":"question-answering","task_name":"Question Answering"}],"methods":[],"datasets_introduced":[{"slug":"milkqa","name":"MilkQA","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}