Browse State-of-the-Art › Question Answering

Question Answering

4,171 papers with code · 142 benchmarks · 413 datasets archive 2025-07-28

MiscellaneousNatural Language ProcessingReasoning

Question answering can be segmented into domain-specific tasks like community question answering and knowledge-base question answering. Popular benchmark datasets for evaluation question answering systems include SQuAD, HotPotQA, bAbI, TriviaQA, WikiQA, and many others. Models for question answering are typically evaluated on metrics like EM and F1. Some recent top performing models are T5 and XLNet.

( Image credit: SQuAD )

Description from the archive archive 2025-07-28; Papers-with-Code links inside it are rewritten to this site.

Benchmarks archive 2025-07-28

152 leaderboard tables shown for this task, 142 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 152 until expanded.

DatasetBest model (first row in archive order)PaperCodeSyntologyCompare
SQuAD2.0 (286 rows) IE-Net (ensemble) — — — Compare
SQuAD1.1 (213 rows) {ANNA} (single model) — — — Compare
HotpotQA (72 rows) Beam Retrieval End-to-End Beam Retrieval for Multi-Hop Question Answering code Syntology ran 10 of 13 samples · 3 unverified Compare
PIQA (67 rows) Unicorn 11B (fine-tuned) UNICORN on RAINBOW: A Universal Commonsense Reasoning Model on a... code Syntology ran 0 of 5 samples · 5 unverified Compare
BoolQ (65 rows) Mistral-Nemo 12B (HPT) Hierarchical Prompting Taxonomy: A Universal Evaluation Framework... code — Compare
COPA (60 rows) PaLM 540B (finetuned) PaLM: Scaling Language Modeling with Pathways code Syntology ran 30 of 37 samples · 7 unverified Compare
TriviaQA (56 rows) Claude 2 (few-shot, k=5) Model Card and Evaluations for Claude Models — — Compare
SQuAD1.1 dev (55 rows) T5-11B Exploring the Limits of Transfer Learning with a Unified... code Syntology ran 2 of 31 samples · 29 unverified Compare
Natural Questions (47 rows) Atlas (full, Wiki-dec-2018 index) Atlas: Few-shot Learning with Retrieval Augmented Language Models code Syntology ran 2 of 2 samples · 0 unverified Compare
OpenBookQA (45 rows) GPT-4 + knowledge base — — — Compare
WebQuestions (37 rows) CoA Chain-of-Action: Faithful and Multimodal Question Answering... code Syntology ran 8 of 10 samples · 2 unverified Compare
TruthfulQA (33 rows) GPT-4 (RLHF) GPT-4 Technical Report code Syntology ran 2 of 5 samples · 3 unverified Compare
MultiRC (30 rows) PaLM 540B (finetuned) PaLM: Scaling Language Modeling with Pathways code Syntology ran 30 of 37 samples · 7 unverified Compare
PubMedQA (30 rows) Meditron-70B (CoT + SC) MEDITRON-70B: Scaling Medical Pretraining for Large Language Models code Syntology ran 9 of 14 samples · 5 unverified Compare
CronQuestions (29 rows) GenTKGQA Two-stage Generative Question Answering on Temporal Knowledge... — — Compare
MedQA (27 rows) Med-Gemini Capabilities of Gemini Models in Medicine — — Compare
WikiQA (25 rows) TANDA-DeBERTa-V3-Large + ALL Structural Self-Supervised Objectives for Transformers code — Compare
SIQA (24 rows) Unicorn 11B (fine-tuned) UNICORN on RAINBOW: A Universal Commonsense Reasoning Model on a... code Syntology ran 0 of 5 samples · 5 unverified Compare
StoryCloze (23 rows) BLOOMZ Crosslingual Generalization through Multitask Finetuning code Syntology ran 1 of 4 samples · 3 unverified Compare
DaNetQA (22 rows) Golden Transformer — — — Compare
TimeQuestions (21 rows) TimeR4 — — — Compare
Quora Question Pairs (19 rows) XLNet (single model) XLNet: Generalized Autoregressive Pretraining for Language Understanding code Syntology ran 10 of 24 samples · 14 unverified Compare
NewsQA (18 rows) OpenAI/o3-2025-01-31-high o3-mini vs DeepSeek-R1: Which One is Safer? code — Compare
CNN / Daily Mail (16 rows) GA+MAGE (32) Linguistic Knowledge as Memory for Recurrent Neural Networks — — Compare
DROP Test (16 rows) QDGAT (ensemble) Question Directed Graph Attention Network for Numerical Reasoning over Text — — Compare
bAbi (14 rows) STM Self-Attentive Associative Memory code Syntology ran 1 of 4 samples · 3 unverified Compare
Natural Questions (long) (13 rows) DensePhrases Learning Dense Representations of Phrases at Scale code — Compare
SQuAD2.0 dev (13 rows) XLNet (single model) XLNet: Generalized Autoregressive Pretraining for Language Understanding code Syntology ran 10 of 24 samples · 14 unverified Compare
TrecQA (13 rows) TANDA DeBERTa-V3-Large + ALL Structural Self-Supervised Objectives for Transformers code — Compare
StrategyQA (12 rows) PaLM 2 (few-shot, CoT, SC) PaLM 2 Technical Report code — Compare
MultiTQ (11 rows) Prog-TQA Self-Improvement Programming for Temporal Knowledge Graph Question... — — Compare
NarrativeQA (10 rows) Masque (NarrativeQA + MS MARCO) Multi-style Generative Reading Comprehension — — Compare
Bamboogle (9 rows) ReST meets ReAct (PaLM 2-L + Google Search) ReST meets ReAct: Self-Improvement for Multi-Step Reasoning LLM Agent — — Compare
CoQA (9 rows) BERT Large Augmented (single model) BERT: Pre-training of Deep Bidirectional Transformers for Language... code Syntology ran 204 of 659 samples · 455 unverified Compare
OBQA (9 rows) FLAN 137B (zero-shot) Finetuned Language Models Are Zero-Shot Learners code Syntology ran 0 of 1 samples · 1 unverified Compare
TIQ (9 rows) FAITH Faithful Temporal Question Answering over Heterogeneous Sources — — Compare
WikiHop (9 rows) BigBird-etc Big Bird: Transformers for Longer Sequences code Syntology ran 10 of 15 samples · 5 unverified Compare
Children's Book Test (8 rows) NSE Gated-Attention Readers for Text Comprehension code Syntology ran 1 of 3 samples · 2 unverified Compare
FEVER (8 rows) CoA Chain-of-Action: Faithful and Multimodal Question Answering... code Syntology ran 8 of 10 samples · 2 unverified Compare
TempQuestions (8 rows) SF-TQA Semantic Framework based Query Generation for Temporal Question... — — Compare
BioASQ (7 rows) BioLinkBERT (large) LinkBERT: Pretraining Language Models with Document Links code Syntology ran 0 of 14 samples · 14 unverified Compare
FQuAD (7 rows) CamemBERT-Large FQuAD: French Question Answering Dataset — — Compare
KILT: ELI5 (7 rows) RBG Read before Generate! Faithful Long Form Question Answering with... — — Compare
QASent (7 rows) Attentive LSTM Neural Variational Inference for Text Processing code Syntology ran 1 of 4 samples · 3 unverified Compare
Quasart-T (7 rows) Cluster-Former (#C=512) Cluster-Former: Clustering-based Sparse Transformer for Long-Range... — — Compare
RACE (7 rows) XLNet XLNet: Generalized Autoregressive Pretraining for Language Understanding code Syntology ran 10 of 24 samples · 14 unverified Compare
SQA3D (7 rows) CREMA CREMA: Generalizable and Efficient Video-Language Reasoning via... code Syntology ran 8 of 10 samples · 2 unverified Compare
Story Cloze (7 rows) Neo-6B (QA + WS) Ask Me Anything: A simple strategy for prompting language models code Syntology ran 2 of 2 samples · 0 unverified Compare
YahooCQA (7 rows) sMIM (1024) + SentenceMIM: A Latent Variable Language Model code Syntology ran 1 of 9 samples · 8 unverified Compare
DROP (6 rows) PaLM 540B (Self Improvement, Self Consistency) Large Language Models Can Self-Improve — — Compare
FinQA (6 rows) APOLLO APOLLO: An Optimized Training Approach for Long-form Numerical Reasoning code — Compare
FriendsQA (6 rows) Ma et al. - ELECTRA Enhanced Speaker-aware Multi-party Multi-turn Dialogue Comprehension — — Compare
NExT-QA (Open-ended VideoQA) (6 rows) Flash-VStream Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams code — Compare
NQ (BEIR) (6 rows) Blended RAG Blended RAG: Improving RAG (Retriever-Augmented Generation)... code Syntology ran 1 of 2 samples · 1 unverified Compare
PeerQA (6 rows) GPT-4o-2024-08-06-128k GPT-4 Technical Report code Syntology ran 2 of 5 samples · 3 unverified Compare
SemEvalCQA (5 rows) HyperQA Hyperbolic Representation Learning for Fast and Efficient Neural... code — Compare
AI2 Kaggle Dataset (4 rows) IR Baseline Tell Me Why: Using Question Answering as Distant Supervision for... — — Compare
BLURB (4 rows) BioLinkBERT (large) LinkBERT: Pretraining Language Models with Document Links code Syntology ran 0 of 14 samples · 14 unverified Compare
catbAbI LM-mode (4 rows) Fast Weight Memory Learning Associative Inference Using Fast Weight Memory code Syntology ran 2 of 3 samples · 1 unverified Compare
catbAbI QA-mode (4 rows) Fast Weight Memory Learning Associative Inference Using Fast Weight Memory code Syntology ran 2 of 3 samples · 1 unverified Compare
CheGeKa (4 rows) Human benchmark TAPE: Assessing Few-shot Russian Language Understanding code — Compare
Complex-CronQuestions (4 rows) SubGTR Temporal knowledge graph question answering via subgraph reasoning code — Compare
EgoTaskQA (4 rows) EgoVLPv2 EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in... code — Compare
FairytaleQA (4 rows) BART fine-tuned on FairytaleQA Fantastic Questions and Where to Find Them: FairytaleQA -- An... code — Compare
FiQA-2018 (BEIR) (4 rows) monoT5-3B No Parameter Left Behind: How Distillation and Model Size Affect... code — Compare
HotpotQA (BEIR) (4 rows) monoT5-3B No Parameter Left Behind: How Distillation and Model Size Affect... code — Compare
HybridQA (4 rows) MAFiD MAFiD: Moving Average Equipped Fusion-in-Decoder for Question... code — Compare
Molweni (4 rows) Ma et al. - ELECTRA Enhanced Speaker-aware Multi-party Multi-turn Dialogue Comprehension — — Compare
MS MARCO (4 rows) Masque Q&A Style Multi-style Generative Reading Comprehension — — Compare
MultiQ (4 rows) Human benchmark TAPE: Assessing Few-shot Russian Language Understanding code — Compare
NaturalQA (4 rows) DPR Dense Passage Retrieval for Open-Domain Question Answering code Syntology ran 10 of 14 samples · 4 unverified Compare
QuALITY (4 rows) Claude 1.3 (5-shot) Model Card and Evaluations for Claude Models — — Compare
RuOpenBookQA (4 rows) Human benchmark TAPE: Assessing Few-shot Russian Language Understanding code — Compare
CaseHOLD (3 rows) Custom Legal-BERT When Does Pretraining Help? Assessing Self-Supervised Learning for... code Syntology ran 0 of 9 samples · 9 unverified Compare
ConditionalQA (3 rows) FiD Leveraging Passage Retrieval with Generative Models for Open... code — Compare
ConvFinQA (3 rows) GPT-4 (8k) Are ChatGPT and GPT-4 General-Purpose Solvers for Financial Text... — — Compare
DuoRC (3 rows) Vector Database (ChromaDB) RecallM: An Adaptable Memory Mechanism with Temporal Understanding... code — Compare
Mathematics Dataset (3 rows) TP-Transformer Enhancing the Transformer with Explicit Relational Encoding for... code Syntology ran 1 of 11 samples · 10 unverified Compare
OTT-QA (3 rows) Fusion Retriever+ETC Open Question Answering over Tables and Text code Syntology ran 4 of 5 samples · 1 unverified Compare
ReClor (3 rows) XLNet-large ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning code Syntology ran 0 of 2 samples · 2 unverified Compare
SberQuAD (3 rows) DeepPavlov RuBERT SberQuAD -- Russian Reading Comprehension Dataset: Description and Analysis — — Compare
SCDE (3 rows) albert-xxlarge + APN(baseline) — — — Compare
TweetQA (3 rows) ByT5 (small) ByT5: Towards a token-free future with pre-trained byte-to-byte models code Syntology ran 0 of 6 samples · 6 unverified Compare
VNHSGE-English (3 rows) Bing Chat VNHSGE: VietNamese High School Graduation Examination Dataset for... code — Compare
AGI Eval (2 rows) Orca 2-13B Orca 2: Teaching Small Language Models How to Reason — — Compare
Aristo Kaggle Allen AI 8th grade questions (2 rows) Cardal Moving Beyond the Turing Test with the Allen AI Science Challenge code — Compare
CliCR (2 rows) Gated-Attention Reader CliCR: A Dataset of Clinical Case Reports for Machine Reading Comprehension code Syntology ran 3 of 4 samples · 1 unverified Compare
CODAH (2 rows) G-DAUG-Combo + RoBERTa-Large Generative Data Augmentation for Commonsense Reasoning code Syntology ran 2 of 4 samples · 2 unverified Compare
COMPLEXQUESTIONS (2 rows) WebQA Evaluating Semantic Parsing against a Simple Web-based Question... code — Compare
GeoQuestions1089 (2 rows) GeoQA2 Benchmarking Geospatial Question Answering Engines using the... code — Compare
MapEval-API (2 rows) Claude-3.5-Sonnet (ReAct) MapEval: A Map-Based Evaluation of Geo-Spatial Reasoning in... code Syntology ran 3 of 3 samples · 0 unverified Compare
MCTest-500 (2 rows) Parallel-Hierarchical A Parallel-Hierarchical Model for Machine Comprehension on Sparse Data code — Compare
MedTurkQuAD: Medical Turkish Question-Answering Dataset (2 rows) BERTurk (cased, 32k) Developing Question-Answering Models in Low-Resource Languages: A... — — Compare
MRQA (2 rows) LinkBERT (large) LinkBERT: Pretraining Language Models with Document Links code Syntology ran 0 of 14 samples · 14 unverified Compare
MuLD (HotpotQA) (2 rows) Longformer MuLD: The Multitask Long Document Benchmark code — Compare
MuLD (NarrativeQA) (2 rows) Longformer MuLD: The Multitask Long Document Benchmark code — Compare
PopQA (2 rows) SelfRAG-7b Self-RAG: Learning to Retrieve, Generate, and Critique through... code Syntology ran 8 of 14 samples · 6 unverified Compare
PubChemQA (2 rows) BioMedGPT-10B BioMedGPT: Open Multimodal Generative Pre-trained Transformer for... code — Compare
QuAC (2 rows) FlowQA (single model) FlowQA: Grasping Flow in History for Conversational Machine Comprehension code — Compare
Reverb (2 rows) Weakly Supervised Embeddings Open Question Answering with Weakly Supervised Embedding Models — — Compare
SQuAD (2 rows) Blended RAG Blended RAG: Improving RAG (Retriever-Augmented Generation)... code Syntology ran 1 of 2 samples · 1 unverified Compare
TempQA-WD (2 rows) BestOfBoth — — — Compare
Torque (2 rows) ECONET ECONET: Effective Continual Pretraining of Language Models for... code — Compare
UniProtQA (2 rows) BioMedGPT-10B BioMedGPT: Open Multimodal Generative Pre-trained Transformer for... code — Compare
VNHSGE-Biology (2 rows) Bing Chat VNHSGE: VietNamese High School Graduation Examination Dataset for... code — Compare
VNHSGE-Chemistry (2 rows) Bing Chat VNHSGE: VietNamese High School Graduation Examination Dataset for... code — Compare
VNHSGE-Civic (2 rows) Bing Chat VNHSGE: VietNamese High School Graduation Examination Dataset for... code — Compare
VNHSGE-Geography (2 rows) Bing Chat VNHSGE: VietNamese High School Graduation Examination Dataset for... code — Compare
VNHSGE-History (2 rows) Bing Chat VNHSGE: VietNamese High School Graduation Examination Dataset for... code — Compare
VNHSGE-Literature (2 rows) ChatGPT VNHSGE: VietNamese High School Graduation Examination Dataset for... code — Compare
VNHSGE Mathematics (2 rows) Bing Chat VNHSGE: VietNamese High School Graduation Examination Dataset for... code — Compare
VNHSGE-Physics (2 rows) Bing Chat VNHSGE: VietNamese High School Graduation Examination Dataset for... code — Compare
WikiSQL (2 rows) PieTa Piece of Table: A Divide-and-Conquer Approach for Selecting... — — Compare
WikiTableQuestions (2 rows) ChatGPT 3.5 SpatialFormat LAPDoc: Layout-Aware Prompting for Documents — — Compare
JD Product Question Answer (1 row) PAAG Product-Aware Answer Generation in E-Commerce Question-Answering code — Compare
AviationQA (1 row) KGT5 There is No Big Brother or Small Brother: Knowledge Infusion in... code — Compare
BBH (1 row) Shakti-LLM (2.5B) SHAKTI: A 2.5 Billion Parameter Small Language Model Optimized for... — — Compare
ChAII - Hindi and Tamil Question Answering (1 row) MuCoT MuCoT: Multilingual Contrastive Training for Question-Answering in... code — Compare
COCO Visual Question Answering (VQA) real images 1.0 open ended (1 row) MaMMUT (2B) MaMMUT: A Simple Architecture for Joint Learning for MultiModal Tasks code Syntology ran 3 of 3 samples · 0 unverified Compare
ComplexWebQuestions (1 row) TOME-2 Mention Memory: incorporating textual knowledge into Transformers... code — Compare
EfficientQA dev (1 row) UnitedQA UnitedQA: A Hybrid Approach for Open Domain Question Answering — — Compare
EfficientQA test (1 row) UnitedQA UnitedQA: A Hybrid Approach for Open Domain Question Answering — — Compare
GraphQuestions (1 row) ChatGPT Can ChatGPT Replace Traditional KBQA Models? An In-depth Analysis... code — Compare
HellaSwag (1 row) Shakti-LLM (2.5B) SHAKTI: A 2.5 Billion Parameter Small Language Model Optimized for... — — Compare
JaQuAD (1 row) BERT-Japanese JaQuAD: Japanese Question Answering Dataset for Machine Reading... code — Compare
KQA Pro (1 row) ChatGPT Can ChatGPT Replace Traditional KBQA Models? An In-depth Analysis... code — Compare
MapEval-Textual (1 row) Claude-3.5-Sonnet MapEval: A Map-Based Evaluation of Geo-Spatial Reasoning in... code Syntology ran 3 of 3 samples · 0 unverified Compare
MCTest-160 (1 row) syntax, frame, coreference, and word embedding features A Parallel-Hierarchical Model for Machine Comprehension on Sparse Data code — Compare
MedMCQA Dev (1 row) MedMobile (3.8B) MedMobile: A mobile-sized language model with expert-level... code — Compare
MetaQA (1 row) T5-small+prolog Domain Specific Question Answering Over Knowledge Graphs Using... code — Compare
MML (1 row) qwen-LLM 7B SHAKTI: A 2.5 Billion Parameter Small Language Model Optimized for... — — Compare
MRQA out-of-domain (1 row) RGX Cooperative Self-training of Machine Reading Comprehension code — Compare
MultiSpanQA (1 row) RoBERTa-large Tagger + LIQUID (Ensemble) LIQUID: A Framework for List Question Answering Dataset Generation code Syntology ran 2 of 2 samples · 0 unverified Compare
QASPER (1 row) Longformer Encoder Decoder (base) A Dataset of Information-Seeking Questions and Answers Anchored in... code Syntology ran 1 of 1 samples · 0 unverified Compare
RecipeQA (1 row) multimodal+LXMERT+ConstrainedMaxPooling Latent Alignment of Procedural Concepts in Multimodal Recipes code — Compare
SchizzoSQUAD (1 row) SchizzoBioBERT Question-Answering Model for Schizophrenia Symptoms and Their... — — Compare
SimpleQuestions (1 row) Memory Networks (ensemble) Large-scale Simple Question Answering with Memory Networks code — Compare
StepGame (1 row) TP-MANN StepGame: A New Benchmark for Robust Multi-Hop Spatial Reasoning in Texts code Syntology ran 5 of 7 samples · 2 unverified Compare
SWAG (1 row) DeBERTaV3large DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with... code Syntology ran 0 of 7 samples · 7 unverified Compare
TAT-QA (1 row) TagOp TAT-QA: A Question Answering Benchmark on a Hybrid of Tabular and... code Syntology ran 2 of 12 samples · 10 unverified Compare
WebQuestionsSP (1 row) ChatGPT Can ChatGPT Replace Traditional KBQA Models? An In-depth Analysis... code — Compare
WebSRC (1 row) ChatGPT 3.5 SpatialFormat LAPDoc: Layout-Aware Prompting for Documents — — Compare
adversarial_qa (0 rows) no rows in the archive — —
AgriQA (0 rows) no rows in the archive — —
squad_adversarial (0 rows) no rows in the archive — —
squad_tr (0 rows) no rows in the archive — —
squad_v2 (0 rows) no rows in the archive — —
squadshifts amazon (0 rows) no rows in the archive — —
squadshifts new_wiki (0 rows) no rows in the archive — —
squadshifts nyt (0 rows) no rows in the archive — —
squadshifts reddit (0 rows) no rows in the archive — —
Tigrinya Q&A (0 rows) no rows in the archive — —

Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.

Libraries

Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.

Datasets archive 2025-07-28

413 datasets whose archive record lists this task, ordered by the archive's paper count. 30 shown of 413 until expanded.

COCO (Common Objects in Context)MNISTMMLNatural QuestionsMS MARCOHellaSwagTriviaQAHotpotQAConceptNetPIQAWinoGrandeBoolQRedditOpenBookQATruthfulQACNN/Daily MailFEVERTextVQARACEDROPOK-VQABBHConceptual CaptionsScienceQACOPAMedQABEIRStrategyQADocVQACoQAPubMedQANewsQAWikiSQLMathVistaWebQuestionsLAMAAI2DNarrativeQAWikiQABioASQXQuADQuACNExT-QAATOMICMLQACORD-19SWAGMultiRCMathQAVisDialELI5SQuADTyDiQAActivityNet-QAe-SNLISHAPESSimpleQuestionsSIQAKILTT-RExMRQAMCTestQASCFinQADVS128 Gesture2WikiMultiHopQACosmosQAQASPERQuALITYRAVENCBTTGIF-QAPopQAST-VQAPathVQAReClorMetaQAWikiTableQuestionsTAT-QASemantic ScholarTrecQAHybridQABeaverTailsBelebeleDREAMCFQWikiHopCOCO-QAComplexWebQuestionsFigureQAWebQuestionsSPTVQA+CANARDSQA3DPrOntoQAQUASAR-TQuora Question PairsReVerb ChallengeBamboogleDAQUARDRCDMUSIC-AVQAQReCCStoryClozeTGIFTOFUChildren's Book TestemrQAPlotQADVQAMKQATQAQUASARCoS-EGeoQADARTGeometry3KShARCDuoRCSCROLLSBLURBCSQATheoremQABREAKWikiMoviesConvFinQAInsuranceQATyDiQA-GoldPDoc2DialMSLR-WEB10KOTT-QAVisualMRCMeQSumSciREXWorldtreeImage Paragraph CaptioningQuaRTzCODAHMolwenidecaNLPKaggleDBQACaseHOLDVQA-HATAdversarialQADUDEFairytaleQAGenericsKBMultiDoc2DialWikiReadingASNQDIOR-RSVGOASST1ARCDRecipeQASpoken-SQuADTechQATempQuestionsHeadQAKQA ProSituatedQASocial-IQTimeQuestionsBookTestEXAMSTopiOCQAWebSRCCliCRFQuADWIQACLOTHMINTAKAQuaRelWebChildCronQuestionsMathematics DatasetQAMPARIQulacStaQCSUTD-TrafficQATorqueTweetQADoQAORCASViolinCLEVR-Ref+ConvQuestionsEgoTaskQAFreebaseQAORConvQAToolQATVCCovidQAQEDStepGameUIT-ViQuADAmazonQAGraphQuestionsMedHopANTIQUEGooAQJEC-QAKLEJMedConceptsQAMEDIQA-AnSMovieGraphsSARASuper-CLEVRVisual MadlibsWho-did-WhatConditionalQADramaQAHabitat PlatformPeerQAPolicyQAPROSTHouse3D EnvironmentMultiTQProtoQATIQCLEVR-DialogFM-IQALAReQAPerception TestReQASberQuADWikiConvAQUAComQACQASUMME-KARHow2RKnowIT VQALEAF-QARadQASelQASPARTQASPARTQA -SQuADShiftsWebCPMCLEVR-MathGermanQuADOPIECPEYMAQUASAR-SQuizbowlSubjQACOGDaNetQADBLP-QuADForecastQAGHOSTSIQUADJGLUEk-qaLiveQAManyModalQANExT-QA (Open-ended VideoQA)RuBQVNHSGEBIPIAFrenchMedMCQAKAMELMovieFIBMuSeRCRELXV2CXQAAI2D-RSTFlenQAKG20CMATINFMoAMOCHAOneStopQAPQuADQAConvRuCoSSpaRTUNTutorialVQAWikiHowQACCPE-MCompMixConvRefCS1QADiSCQIndoNLGJaQuADSCDEUniProtQAVideoNavQAXOR-TYDI QAAfriQAAllegro ReviewsCHQ-SummClarQComplex-CronQuestionsConcurrentQA BenchmarkCOVID-QDICE: a Dataset of Italian Crime Event newsDisfl-QADrugEHRQAExplainCPEFanOutQALLeQAMIMIC-SPARQLMMEDMuLDMultiQODSQAPDFVQAPubChemQAQALD-9-PlusResearchy QuestionsResQShmoop CorpusTextBox 2.0TriviaHGUIT-ViCoV19QAAlmawave-SLUAWS DocumentationCheGeKaCII-BenchExpMRCG-VUEM2QAMilkQAMultiReQANQuADPoseScriptQUITEReviewQARoMQARuOpenBookQAShortcut QASimpleDBpediaQAStanford Schema2QA DatasetTempQA-WDVisual BeliefsVizWiz-PrivVizWiz-QualityIssuesVQA 360°X-WikiREXorQA-IN:XQuAD-INACCESS DENIED INCACCORD CSQA 0-5AviationQABDD-QAcatbAbI LM-modecatbAbI QA-modeCC-RiddleChAII - Hindi and Tamil Question AnsweringChiQACompMix-IRCOPA-HRCoreSearchCRSBCUHK-QAdiaforge-utc-r-0725Dialog-based Language Learning datasetDynToMEpiK-EvalEvent-QAFinancial Language Understanding EvaluationGeoQuestions1089ILSP Greek Evaluation SuiteMapEvalMapEval-APIMapEval-TextualMatToolsMedTurkQuAD: Medical Turkish Question-Answering DatasetMF3QAMF3QA_uncleanedMILUMLQuestionsMMInstruct-GPT4VNitiBenchNTextORKG-QAPentachromatic Cultural Palette DatasetPersian NLP HubPersianQAPhrase-in-ContextPhysiCoPiráPQ-decaNLPPQArefQASiNaQASportsRGRSSchizzoSQUADScienceExamCERSCIMATsimply-CLEVRSOTU_QA_2023TimelineKGQATinySocialToM QATQBA++TUMTraffic-VideoQATupleInf Open IE DatasetVCG+112KVDQGVisual Choice of Plausible AlternativesVisual Haystacks (VHs)VlogQAvqa-nle-llavaVTQAWangchanX-Legal-ThaiCCL-RAGWikiQAarWikiSuggestXamarin Q&AXLingEval

Subtasks archive 2025-07-28

19 subtasks in the archive's task tree.

Most implemented papers archive 2025-07-28

30 shown of 4,171 papers with code (10,817 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.

  • 12 Jun 2017 595 repositories listed Syntology ran 600 of 946 samples · 346 unverified · 451 pointer-only (licence)
    The dominant sequence transduction models are based on complex recurrent or convolutional neural networks in an encoder-decoder configuration.
  • 11 Oct 2018 534 repositories listed Syntology ran 204 of 659 samples · 455 unverified · 149 pointer-only (licence)
    We introduce a new language representation model called BERT, which stands for Bidirectional Encoder Representations from Transformers.
  • 30 Oct 2017 93 repositories listed Syntology ran 50 of 106 samples · 56 unverified · 43 pointer-only (licence)
    We present graph attention networks (GATs), novel neural network architectures that operate on graph-structured data, leveraging masked self-attentional layers to address the shortcomings of prior methods based on graph…
  • 28 May 2020 67 repositories listed Syntology ran 15 of 65 samples · 50 unverified · 4 pointer-only (licence)
    By contrast, humans can generally perform a new language task from only a few examples or from simple instructions - something which current NLP systems still largely struggle to do.
  • 26 Jul 2019 67 repositories listed Syntology ran 22 of 48 samples · 26 unverified · 23 pointer-only (licence)
    Language model pretraining has led to significant performance gains but careful comparison between different approaches is challenging.
  • 27 Feb 2023 57 repositories listed Syntology ran 26 of 58 samples · 32 unverified · 4 pointer-only (licence)
    We introduce LLaMA, a collection of foundation language models ranging from 7B to 65B parameters.
  • 23 Oct 2019 57 repositories listed Syntology ran 2 of 31 samples · 29 unverified
    Transfer learning, where a model is first pre-trained on a data-rich task before being fine-tuned on a downstream task, has emerged as a powerful technique in natural language processing (NLP).
  • 26 Sep 2019 48 repositories listed Syntology ran 46 of 126 samples · 80 unverified · 22 pointer-only (licence)
    Increasing model size when pretraining natural language representations often results in improved performance on downstream tasks.
  • 29 Oct 2019 47 repositories listed Syntology ran 22 of 53 samples · 31 unverified · 7 pointer-only (licence)
    We evaluate a number of noising approaches, finding the best performance by both randomly shuffling the order of the original sentences and using a novel in-filling scheme, where spans of text are replaced with a single…
  • 15 Feb 2018 46 repositories listed Syntology ran 23 of 58 samples · 35 unverified · 25 pointer-only (licence)
    We introduce a new type of deep contextualized word representation that models both (1) complex characteristics of word use (e.
  • 31 Mar 2015 44 repositories listed Syntology ran 2 of 15 samples · 13 unverified · 5 pointer-only (licence)
    For the former our approach is competitive with Memory Networks, but with less supervision.
  • 2 Oct 2019 37 repositories listed Syntology ran 19 of 27 samples · 8 unverified
    As Transfer Learning from large-scale pre-trained models becomes more prevalent in Natural Language Processing (NLP), operating these large models in on-the-edge and/or under constrained computational training or…
  • 1 Apr 2019 32 repositories listed Syntology ran 1 of 11 samples · 10 unverified · 9 pointer-only (licence)
    In this paper, we first study a principled layerwise adaptation strategy to accelerate training of deep neural networks using large mini-batches.
  • 19 Jun 2019 27 repositories listed Syntology ran 10 of 24 samples · 14 unverified · 3 pointer-only (licence)
    With the capability of modeling bidirectional contexts, denoising autoencoding based pretraining like BERT achieves better performance than pretraining approaches based on autoregressive language modeling.
  • 5 Nov 2016 27 repositories listed Syntology ran 8 of 11 samples · 3 unverified · 7 pointer-only (licence)
    Machine comprehension (MC), answering a query about a given context paragraph, requires modeling complex interactions between the context and the query.
  • 16 May 2014 27 repositories listed Syntology ran 0 of 4 samples · 4 unverified
    Its construction gives our algorithm the potential to overcome the weaknesses of bag-of-words models.
  • 10 Apr 2020 22 repositories listed Syntology ran 15 of 35 samples · 20 unverified · 5 pointer-only (licence)
    To address this limitation, we introduce the Longformer with an attention mechanism that scales linearly with sequence length, making it easy to process documents of thousands of tokens or longer.
  • 14 Feb 2019 21 repositories listed
    Natural language processing tasks, such as question answering, machine translation, reading comprehension, and summarization, are typically approached with supervised learning on taskspecific datasets.
  • 16 Jun 2016 21 repositories listed Syntology ran 4 of 6 samples · 2 unverified · 6 pointer-only (licence)
    We present the Stanford Question Answering Dataset (SQuAD), a new reading comprehension dataset consisting of 100, 000+ questions posed by crowdworkers on a set of Wikipedia articles, where the answer to each question…
  • 17 May 2021 20 repositories listed Syntology ran 34 of 44 samples · 10 unverified · 11 pointer-only (licence)
    Transformers have become one of the most important architectural innovations in deep learning and have enabled many breakthroughs over the past few years.
  • 5 Jun 2017 20 repositories listed Syntology ran 3 of 7 samples · 4 unverified · 1 pointer-only (licence)
    Relational reasoning is a central component of generally intelligent behavior, but has proven difficult for neural networks to learn.
  • 19 Feb 2015 20 repositories listed Syntology ran 1 of 3 samples · 2 unverified · 2 pointer-only (licence)
    One long-term goal of machine learning research is to produce methods that are applicable to reasoning and natural language, in particular building an intelligent dialogue agent.
  • 18 Jul 2023 19 repositories listed Syntology ran 31 of 52 samples · 21 unverified · 16 pointer-only (licence)
    In this work, we develop and release Llama 2, a collection of pretrained and fine-tuned large language models (LLMs) ranging in scale from 7 billion to 70 billion parameters.
  • 28 Jan 2022 19 repositories listed Syntology ran 2 of 7 samples · 5 unverified
    We explore how generating a chain of thought -- a series of intermediate reasoning steps -- significantly improves the ability of large language models to perform complex reasoning.
  • 10 Apr 2020 19 repositories listed Syntology ran 10 of 14 samples · 4 unverified · 9 pointer-only (licence)
    Open-domain question answering relies on efficient passage retrieval to select candidate contexts, where traditional sparse vector space models, such as TF-IDF or BM25, are the de facto method.
  • 23 Mar 2020 19 repositories listed Syntology ran 26 of 40 samples · 14 unverified · 10 pointer-only (licence)
    Then, instead of training a model that predicts the original identities of the corrupted tokens, we train a discriminative model that predicts whether each token in the corrupted input was replaced by a generator sample…
  • 19 Apr 2019 19 repositories listed Syntology ran 0 of 7 samples · 7 unverified
    We present a novel language representation model enhanced by knowledge called ERNIE (Enhanced Representation through kNowledge IntEgration).
  • 25 Jan 2019 19 repositories listed Syntology ran 4 of 25 samples · 21 unverified · 1 pointer-only (licence)
    Biomedical text mining is becoming increasingly important as the number of biomedical documents rapidly grows.
  • 22 May 2020 18 repositories listed Syntology ran 4 of 6 samples · 2 unverified
    Large pre-trained language models have been shown to store factual knowledge in their parameters, and achieve state-of-the-art results when fine-tuned on downstream NLP tasks.
  • 23 Apr 2018 15 repositories listed Syntology ran 7 of 19 samples · 12 unverified · 2 pointer-only (licence)
    On the SQuAD dataset, our model is 3x to 13x faster in training and 4x to 9x faster in inference, while achieving equivalent accuracy to recurrent models.

Syntology lines on 29 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections