{"url":"/dataset/feedbackqa","name":"FeedbackQA","full_name":null,"description_markdown":"[📄 Read](https://arxiv.org/abs/2204.03025)<br>\r\n[💾 Code](https://github.com/McGill-NLP/feedbackqa)<br>\r\n[🔗 Webpage](https://mcgill-nlp.github.io/feedbackqa/)<br>\r\n[💻 Demo](http://206.12.100.48:8080/)<br>\r\n[🤗 Huggingface Dataset](https://huggingface.co/datasets/McGill-NLP/feedbackQA)<br>\r\n[💬 Discussions](https://github.com/McGill-NLP/feedbackqa/discussions)\r\n\r\n# Overview\r\n\r\nUsers interact with QA systems and leave feedback. In this project, we investigate methods of improving QA systems further post-deployment based on user interactions.\r\n\r\n# Dataset\r\n\r\nWe collect a retrieval-based QA dataset, FeedbackQA, which contains interactive feedback from users. We collect this dataset by deploying a base QA system to crowdworkers who then engage with the system and provide feedback on the quality of its answers. The feedback contains both structured ratings and unstructured natural language explanations. Check the \r\ndataset explorer at the bottom for some real examples.\r\n\r\n# Methods\r\n\r\nWe propose a method to improve the RQA model with the feedback data, training a reranker to select an answer candidate as well as generate the explanation. We find that this approach not only increases the accuracy of the deployed model but also other stronger models for which feedback data is not collected. Moreover, our human evaluation results show that both human-written and model-generated explanations help users to make informed and accurate decisions about whether to accept an answer. Read our paper for more details, and play with our demo for an intuitive understanding of what we have done.","description_withheld":null,"homepage":"https://mcgill-nlp.github.io/feedbackqa/","introduced_date":"2021-11-16","introduced_date_note":null,"introduced_by":{"paper":"/paper/using-interactive-feedback-to-improve-the","title":"Using Interactive Feedback to Improve the Accuracy and Explainability of Question Answering Systems Post-Deployment","first_author":"Anonymous","url":null},"license":{"name":"Apache 2.0","url":null},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Overall - Test","url":"/task/overall-test","datasets_with_task":"/datasets/task/overall-test"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["FeedbackQA"],"data_loaders":[{"repo":"https://github.com/McGill-NLP/feedbackqa","url":"https://github.com/McGill-NLP/feedbackqa","frameworks":[]}],"num_papers_in_archive":2,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}