Browse State-of-the-Art › Answer Generation
Answer Generation
111 papers with code · 2 benchmarks · 4 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
2 leaderboard tables shown for this task, 2 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| WeiboPolls (3 rows) | UniPoll | UniPoll: A Unified Social Media Poll Generation Framework via... | code | — | Compare |
| CICERO (2 rows) | T5-large pre-trained on GLUCOSE | CICERO: A Dataset for Contextualized Commonsense Inference in Dialogues | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
4 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 111 papers with code (280 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
23 Oct 2019 57 repositories listed Syntology ran 2 of 31 samples · 29 unverifiedTransfer learning, where a model is first pre-trained on a data-rich task before being fine-tuned on a downstream task, has emerged as a powerful technique in natural language processing (NLP).
-
9 Sep 2024 4 repositories listedRegulatory documents, issued by governmental regulatory bodies, establish rules, guidelines, and standards that organizations must adhere to for legal compliance.
-
30 May 2022 3 repositories listedThis paper introduces our proposed system for the MIA Shared Task on Cross-lingual Open-retrieval Question Answering (COQA).
-
24 Jun 2021 3 repositories listedThe VOGUE framework attempts to generate a verbalized answer using a hybrid approach through a multi-task learning paradigm.
-
21 Nov 2020 3 repositories listedWe show that LRTA makes a step towards truly understanding the question while the state-of-the-art model tends to learn superficial correlations from the training data.
-
21 Jul 2016 3 repositories listedWhile question answering (QA) with neural network, i.
-
28 Jun 2024 2 repositories listedThis paper introduces MM-Instruct, a large-scale dataset of diverse and high-quality visual instruction data designed to enhance the instruction-following capabilities of large multimodal models (LMMs).
-
PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models12 Feb 2024 2 repositories listed Syntology ran 6 of 17 samples · 11 unverifiedBased on this attack surface, we propose PoisonedRAG, the first knowledge corruption attack to RAG, where an attacker could inject a few malicious texts into the knowledge database of a RAG system to induce an LLM to…
-
19 May 2023 2 repositories listed Syntology ran 7 of 7 samples · 0 unverifiedIn this paper, we propose Visual Question Localized-Answering in Robotic Surgery (Surgical-VQLA) to localize the specific surgical area during the answer prediction.
-
25 May 2022 2 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedOur results highlight the need for developing ODQA models that handle a broad range of question types, including single and multi-answer questions.
-
20 Dec 2021 2 repositories listedSpecifically, the task involves multi-hop questions that require reasoning over image-caption pairs to identify the grounded visual object being referred to and then predicting a span from the news body text to answer…
-
It is AI's Turn to Ask Humans a Question: Question-Answer Pair Generation for Children's Story Books8 Sep 2021 2 repositories listedExisting question answering (QA) techniques are created mainly to answer questions asked by humans.
-
9 Jun 2021 2 repositories listed Syntology ran 4 of 4 samples · 0 unverified · 4 pointer-only (licence)We model retrieval decisions as latent variables over sets of relevant documents.
-
18 Feb 2021 2 repositories listedAs a first step towards measuring news informedness at a scale, we study the problem of quiz-style multiple-choice question generation, which may be used to survey users about their knowledge of recent news.
-
27 Jan 2020 2 repositories listed Syntology ran 1 of 8 samples · 7 unverified · 2 pointer-only (licence)In this paper, we propose Answer-Clue-Style-aware Question Generation (ACS-QG), which aims at automatically generating high-quality and diverse question-answer pairs from unlabeled text corpus at scale by imitating the…
-
17 Jun 2025 1 repository listedThis paper presents the RMIT--ADM+S participation in the SIGIR 2025 LiveRAG Challenge.
-
12 Jun 2025 1 repository listedRetrieval-Augmented Generation (RAG) has demonstrated considerable effectiveness in open-domain question answering.
-
5 Jun 2025 1 repository listedECoRAG improves LLM performance by compressing retrieved documents based on evidentiality, ensuring whether answer generation is supported by the correct evidence.
-
22 May 2025 1 repository listed Syntology ran 0 of 2 samples · 2 unverifiedLarge Language Models (LLMs), despite their advancements, are fundamentally limited by their static parametric knowledge, hindering performance on tasks requiring open-domain up-to-date information.
-
21 May 2025 1 repository listed Syntology ran 1 of 10 samples · 9 unverifiedIn this paper, we first conduct extensive analysis experiments of three key components of that training pipeline: input design, output evaluation, and policy update-each revealing distinct challenges arising from…
-
21 May 2025 1 repository listed Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)Large language models are able to exploit in-context learning to access external knowledge beyond their training data through retrieval-augmentation.
-
20 May 2025 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)To address these limitations, agentic RAG systems (e.
-
18 May 2025 1 repository listedThe increasing context length of modern language models has created a need for evaluating their ability to retrieve and process information across extensive documents.
-
11 Apr 2025 1 repository listedRecent advances in large language models (LLMs) provide new opportunities for context understanding in virtual reality (VR).
-
3 Apr 2025 1 repository listedThe stronger clean logit difference for definitional questions further supports this localized representation.
-
19 Mar 2025 1 repository listedInspired by this cognitive process, we propose \textbf{MetaLadder}, a novel framework that explicitly prompts LLMs to recall and reflect on meta-problems, those structurally or semantically analogous problems, alongside…
-
12 Mar 2025 1 repository listedBuilt on this resource, we provide a framework for long-form answer generation evaluation, involving nuggets extraction and nuggets matching, linked to retrieval.
-
4 Mar 2025 1 repository listedWith the rapid development in Transformer-based language models, the reading comprehension tasks on short documents and simple questions have been largely addressed.
-
27 Feb 2025 1 repository listed Syntology ran 0 of 3 samples · 3 unverifiedTo improve and evaluate the normative reasoning capability of vision-language models (VLMs), we present \dataset{} ϵ, consisting of 1, 853 challenging, multi-stage MCQ questions based on ego-centric videos of human…
-
24 Feb 2025 1 repository listedRegulatory texts are inherently long and complex, presenting significant challenges for information retrieval systems in supporting regulatory officers with compliance tasks.
Syntology lines on 11 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections