Browse State-of-the-Art › Answer Selection
Answer Selection
52 papers with code · 6 benchmarks · 10 datasets archive 2025-07-28
Answer Selection is the task of identifying the correct answer to a question from a pool of candidate answers. This task can be formulated as a classification or a ranking problem.
Source: Learning Analogy-Preserving Sentence Embeddings for Answer Selection
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
6 leaderboard tables shown for this task, 6 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| ASNQ (3 rows) | DeBERTa-V3-Large + SSP | Pre-training Transformer Models with Sentence-Level Objectives for... | — | — | Compare |
| CICERO (2 rows) | T5-large | CICERO: A Dataset for Contextualized Commonsense Inference in Dialogues | code | — | Compare |
| Ubuntu Dialogue (v2, Ranking) (2 rows) | BERT + Keep Learning | Keep Learning: Self-supervised Meta-learning for Learning from Inference | — | — | Compare |
| TrecQA (1 row) | RLAS-BIABC | RLAS-BIABC: A Reinforcement Learning-Based Answer Selection Using... | — | — | Compare |
| Ubuntu Dialogue (v1, Ranking) (1 row) | HRDE-LTC | Learning to Rank Question-Answer Pairs using Hierarchical... | code | — | Compare |
| WikiQA (1 row) | RLAS-BIABC | RLAS-BIABC: A Reinforcement Learning-Based Answer Selection Using... | — | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
10 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 52 papers with code (171 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
The Ubuntu Dialogue Corpus: A Large Dataset for Research in Unstructured Multi-Turn Dialogue Systems30 Jun 2015 21 repositories listed Syntology ran 1 of 26 samples · 25 unverified · 1 pointer-only (licence)This paper introduces the Ubuntu Dialogue Corpus, a dataset containing almost 1 million multi-turn dialogues, with a total of over 7 million utterances and 100 million words.
-
16 Dec 2015 8 repositories listed(ii) We propose three attention schemes that integrate mutual influence between sentences into CNN; thus, the representation of each sentence takes into consideration its counterpart.
-
19 Nov 2015 6 repositories listed Syntology ran 1 of 4 samples · 3 unverified · 1 pointer-only (licence)We validate this framework on two very different text modelling applications, generative document modelling and supervised question answering.
-
5 Jun 2016 4 repositories listed Syntology ran 1 of 3 samples · 2 unverifiedIn this paper we study the problem of answering cloze-style questions over documents.
-
1 Aug 2019 3 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedIn this paper, we present a fast and strong neural approach for general purpose text matching applications.
-
10 Oct 2017 3 repositories listedIn this paper, we propose a novel end-to-end neural architecture for ranking candidate answers, that adapts a hierarchical recurrent neural network and a latent topic clustering module.
-
11 Feb 2016 3 repositories listedIn this work, we propose Attentive Pooling (AP), a two-way attention mechanism for discriminative model training.
-
22 Mar 2021 2 repositories listedIn this paper we aim to solve the latter one by proposing a deep latent variable model, in which multiple Gaussian processes are employed as priors of latent variables to separately learn underlying abstract concepts…
-
1 Jul 2020 2 repositories listedIn a separate line of research, KG embedding methods have been proposed to reduce KG sparsity by performing missing link prediction.
-
6 Dec 2018 2 repositories listedSecond, these two tasks can benefit each other: answer selection can incorporate the external knowledge from knowledge base (KB), while KBQA can be improved by learning contextual information from answer selection.
-
6 Nov 2016 2 repositories listedWe particularly focus on the different comparison functions we can use to match two vectors.
-
4 Nov 2016 2 repositories listedIn this paper, we focus on this answer extraction task, presenting a novel model architecture that efficiently builds fixed length representations of all spans in the evidence document with a recurrent network.
-
12 Nov 2015 2 repositories listedOne direction is to define a more composite representation for questions and answers by combining convolutional neural network with the basic framework.
-
7 Aug 2015 2 repositories listedWe apply a general deep learning framework to address the non-factoid question answering task.
-
24 Apr 2025 1 repository listedOur system focuses on financial non-factoid answer selection, which retrieves a set of passage-level texts and selects the most relevant as the answer.
-
16 Apr 2025 1 repository listedIn this paper, we explore the upper bound of harnessing multilingualism in reasoning tasks, suggesting that multilingual reasoning promises significantly (by nearly 10 Acc@k points) and robustly (tolerance for…
-
25 Sep 2024 1 repository listedText-to-SQL parsing and end-to-end question answering (E2E TQA) are two main approaches for Table-based Question Answering task.
-
21 May 2024 1 repository listedRecent advancements in Chain-of-Thought prompting have facilitated significant breakthroughs for Large Language Models (LLMs) in complex reasoning tasks.
-
9 Apr 2024 1 repository listed Syntology ran 5 of 5 samples · 0 unverified · 5 pointer-only (licence)Recently, the large language model (LLM) community has shown increasing interest in enhancing LLMs' capability to handle extremely long documents.
-
14 Feb 2024 1 repository listedWith the widespread adoption of large language models (LLMs) in numerous applications, the challenge of factuality and the propensity for hallucinations has emerged as a significant concern.
-
1 Feb 2024 1 repository listed Syntology ran 2 of 2 samples · 0 unverifiedLarge Language Model (LLM) leaderboards based on benchmark rankings are regularly used to guide practitioners in model selection.
-
26 Aug 2023 1 repository listedFirstly, We propose a problem type classifier that combines the strengths of the tree-based solver and the LLM solver.
-
15 Jul 2023 1 repository listedFinally, we conduct experiments to illustrate the interpretability of CRAB in concept learning, answer selection, and global rule abstraction.
-
10 Feb 2023 1 repository listedConversational Question Answering (ConvQA) models aim at answering a question with its relevant paragraph and previous question-answer pairs that occurred during conversation multiple times.
-
22 Oct 2022 1 repository listed Syntology ran 6 of 7 samples · 1 unverifiedA more natural prompting approach is to present the question and answer options to the LLM jointly and have it output the symbol (e.
-
11 Oct 2022 1 repository listedTransformer-based models have achieved great success on sentence pair modeling tasks, such as answer selection and natural language inference (NLI).
-
1 Oct 2022 1 repository listedIn this paper, we propose a novel Spurious Correlation reduction method to improve the robustness of the neural ANswer selection models (SCAN) from the sample and feature perspectives by removing the feature…
-
1 Oct 2022 1 repository listedThe results of our extrinsic evaluation show that while there is a significant difference between the performance of the rule-based system vs.
-
2 May 2022 1 repository listedOur evaluation on three AS2 and one fact verification datasets demonstrates the superiority of our pre-training technique over the traditional ones for transformers used as joint models for multi-candidate inference…
-
30 Apr 2022 1 repository listedWe report the performance of DeBERTaV3 on CommonsenseQA in this report.
Syntology lines on 7 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections