Browse State-of-the-Art › Long Form Question Answering
Long Form Question Answering
32 papers with code · 0 benchmarks · 5 datasets archive 2025-07-28
Long-form question answering is a task requiring elaborate and in-depth answers to open-ended questions.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
5 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 32 papers with code (61 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
22 Jul 2019 3 repositories listed Syntology ran 4 of 11 samples · 7 unverified · 4 pointer-only (licence)We introduce the first large-scale corpus for long-form question answering, a task requiring elaborate and in-depth answers to open-ended questions.
-
8 Nov 2023 2 repositories listedExperimenting with several LLMs in various settings, we find this task to be surprisingly challenging, demonstrating the importance of QuoteSum for developing and studying such consolidation capabilities.
-
5 Oct 2023 2 repositories listedA single language model, even when aligned with labelers through reinforcement learning from human feedback (RLHF), may not suit all human preferences.
-
17 Apr 2023 2 repositories listedWe generate instructions via LLMs for human-written corpus examples using reverse instructions.
-
10 Mar 2021 2 repositories listedThe task of long-form question answering (LFQA) involves retrieving documents relevant to a given question and using them to generate a paragraph-length answer.
-
17 Jun 2025 1 repository listedRecent large language models (LLMs) achieve impressive performance in source-conditioned text generation but often fail to correctly provide fine-grained attributions for their outputs, undermining verifiability and…
-
14 May 2025 1 repository listedLarge Language Models (LLMs) frequently produce factoid hallucinations - plausible yet incorrect answers.
-
25 Apr 2025 1 repository listedWe address this gap by conducting an in-depth study of long-form answer evaluation with the following research questions: (i) To what extent do existing automatic evaluation metrics serve as a substitute for human…
-
20 Feb 2025 1 repository listedFinally, we show that different general-purpose LLMs excel in the biomedical domain than the encyclopedic one, and that open-domain evidence retrieval in large corpora is challenging.
-
18 Feb 2025 1 repository listedIn the age of misinformation, hallucination -- the tendency of Large Language Models (LLMs) to generate non-factual or unfaithful responses -- represents the main risk for their global utility.
-
13 Feb 2025 1 repository listedWe introduce SelfCite, a novel self-supervised approach that aligns LLMs to generate high-quality, fine-grained, sentence-level citations for the statements in their generated responses.
-
16 Oct 2024 1 repository listedTo bridge this gap, we introduce a new claim decomposition benchmark, which requires building system that can identify atomic and checkworthy claims for LLM responses.
-
21 Aug 2024 1 repository listedWe develop and benchmark a RAG model against a standard, non-RAG LLM, focusing on transcription, retrieval, and generation performance.
-
20 Aug 2024 1 repository listedBy enhancing the intelligibility of human questions for black-box LLMs, our question rewriter improves the quality of generated answers.
-
16 Jul 2024 1 repository listedThis work introduces HaluQuestQA, the first hallucination dataset with localized error annotations for human-written and model-generated LFQA answers.
-
5 Jul 2024 1 repository listed Syntology ran 1 of 4 samples · 3 unverifiedSuch an annotator can not only evaluate the hallucination levels of various LLMs on the large-scale dataset but also help to mitigate the hallucination of LLMs generations, with the Natural Language Inference (NLI)…
-
25 Jun 2024 1 repository listedTo bridge this gap, we introduce CaLMQA, a collection of 1.
-
21 May 2024 1 repository listedWe also propose OLAPH, a simple and novel framework that utilizes cost-effective and multifaceted automatic evaluation to construct a synthetic preference set and answers questions in our preferred manner.
-
2 Apr 2024 1 repository listedWe present ClapNQ, a benchmark Long-form Question Answering dataset for the full RAG pipeline.
-
25 Mar 2024 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedRecent efforts to address hallucinations in Large Language Models (LLMs) have focused on attributed text generation, which supplements generated texts with citations of supporting sources for post-generation…
-
11 Mar 2024 1 repository listed Syntology ran 7 of 14 samples · 7 unverifiedWe introduce ALaRM, the first framework modeling hierarchical rewards in reinforcement learning from human feedback (RLHF), which is designed to enhance the alignment of large language models (LLMs) with human…
-
KG-Rank: Enhancing Large Language Models for Medical QA with Knowledge Graphs and Ranking Techniques9 Mar 2024 1 repository listed Syntology ran 3 of 5 samples · 2 unverified · 5 pointer-only (licence)In this work, we develop an augmented LLM framework, KG-Rank, which leverages a medical knowledge graph (KG) along with ranking and re-ranking techniques, to improve the factuality of long-form question answering (QA)…
-
29 Nov 2023 1 repository listedQuery-focused Summarization (QfS) deals with systems that generate summaries from document(s) based on a query.
-
2 Jun 2023 1 repository listed Syntology ran 1 of 2 samples · 1 unverifiedWe introduce Fine-Grained RLHF, a framework that enables training and learning from reward functions that are fine-grained in two respects: (1) density, providing a reward after every segment (e.
-
30 May 2023 1 repository listed Syntology ran 2 of 5 samples · 3 unverified · 5 pointer-only (licence)Long-form question answering systems provide rich information by presenting paragraph-level answers, often containing optional background or auxiliary information.
-
29 May 2023 1 repository listedWe present a careful analysis of experts' evaluation, which focuses on new aspects such as the comprehensiveness of the answer.
-
24 May 2023 1 repository listedIn this paper, we suggest revisiting the sentence union generation task as an effective well-defined testbed for assessing text consolidation capabilities, decoupling the consolidation challenge from subjective content…
-
11 May 2023 1 repository listed Syntology ran 1 of 4 samples · 3 unverifiedWe recruit annotators to search for relevant information using our interface and then answer questions.
-
28 Apr 2023 1 repository listedThis paper proposes a novel framework named \textbf{Search-in-the-Chain} (SearChain) for the interaction between LLM and IR to solve the challenges.
-
8 May 2021 1 repository listedPresentations are critical for communication in all areas of our lives, yet the creation of slide decks is often tedious and time-consuming.
Syntology lines on 8 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections