Browse State-of-the-Art › Question Answering › Papers, page 44
Question Answering
Papers archive 2025-07-28
archive papers tagged: 10,817 · with a code link: 4,171 · where Syntology ran a sample: 1,274 (1,073 with a run with no instrument failure, 201 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (1,274 of 10,817 tagged: 1,073 with a run with no instrument failure, 201 where every run was a failure of Syntology's instrument)
Page 44 of 109: papers 4,301 to 4,400 of 10,817, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
DynamicVL: Benchmarking Multimodal Large Language Models for Dynamic City Understanding27 May 2025 0 repositories listed
-
Music's Multimodal Complexity in AVQA: Why We Need More than General Multimodal LLMs27 May 2025 0 repositories listed
-
Rethinking Information Synthesis in Multimodal Question Answering A Multi-Agent Perspective27 May 2025 0 repositories listed
-
Silence is Not Consensus: Disrupting Agreement Bias in Multi-Agent LLMs via Catfish Agent for Clinical Decision Making27 May 2025 0 repositories listed
-
SOSBENCH: Benchmarking Safety Alignment on Scientific Knowledge27 May 2025 0 repositories listed
-
ALAS: Measuring Latent Speech-Text Alignment For Spoken Language Understanding In Multimodal LLMs26 May 2025 0 repositories listed
-
Benchmarking Large Multimodal Models for Ophthalmic Visual Question Answering with OphthalWeChat26 May 2025 0 repositories listed
-
CP-Router: An Uncertainty-Aware Router Between LLM and LRM26 May 2025 0 repositories listed
-
DGRAG: Distributed Graph-based Retrieval-Augmented Generation in Edge-Cloud Systems26 May 2025 0 repositories listed
-
Interleaved Reasoning for Large Language Models via Reinforcement Learning26 May 2025 0 repositories listed
-
It's High Time: A Survey of Temporal Information Retrieval and Question Answering26 May 2025 0 repositories listed
-
S2LPP: Small-to-Large Prompt Prediction across LLMs26 May 2025 0 repositories listed
-
Self-Reflective Planning with Knowledge Graphs: Enhancing LLM Reasoning Reliability for Question Answering26 May 2025 0 repositories listed
-
SIPDO: Closed-Loop Prompt Optimization via Synthetic Data Feedback26 May 2025 0 repositories listed
-
Uncertainty-Aware Attention Heads: Efficient Unsupervised Uncertainty Quantification for LLMs26 May 2025 0 repositories listed
-
ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment25 May 2025 0 repositories listed
-
GC-KBVQA: A New Four-Stage Framework for Enhancing Knowledge Based Visual Question Answering Performance25 May 2025 0 repositories listed
-
SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs25 May 2025 0 repositories listed
-
UNCERTAINTY-LINE: Length-Invariant Estimation of Uncertainty for Large Language Models25 May 2025 0 repositories listed
-
Weaver: Interweaving SQL and LLM for Table Reasoning25 May 2025 0 repositories listed
-
When Two LLMs Debate, Both Think They'll Win25 May 2025 0 repositories listed
-
Benchmarking Poisoning Attacks against Retrieval-Augmented Generation24 May 2025 0 repositories listed
-
BRIT: Bidirectional Retrieval over Unified Image-Text Graph24 May 2025 0 repositories listed
-
Climate-Eval: A Comprehensive Benchmark for NLP Tasks Related to Climate Change24 May 2025 0 repositories listed
-
Multilingual Question Answering in Low-Resource Settings: A Dzongkha-English Benchmark for Foundation Models24 May 2025 0 repositories listed
-
The Silent Saboteur: Imperceptible Adversarial Attacks against Black-Box Retrieval-Augmented Generation Systems24 May 2025 0 repositories listed
-
Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models23 May 2025 0 repositories listed
-
FinRAGBench-V: A Benchmark for Multimodal RAG with Visual Citation in the Financial Domain23 May 2025 0 repositories listed
-
PerMedCQA: Benchmarking Large Language Models on Medical Consumer Question Answering in Persian Language23 May 2025 0 repositories listed
-
PPT: A Process-based Preference Learning Framework for Self Improving Table Question Answering Models23 May 2025 0 repositories listed
-
Task Specific Pruning with LLM-Sieve: How Many Parameters Does Your Task Really Need?23 May 2025 0 repositories listed
-
A Causal Approach to Mitigate Modality Preference Bias in Medical Visual Question Answering22 May 2025 0 repositories listed
-
Are the Hidden States Hiding Something? Testing the Limits of Factuality-Encoding Capabilities in LLMs22 May 2025 0 repositories listed
-
Augmenting LLM Reasoning with Dynamic Notes Writing for Complex QA22 May 2025 0 repositories listed
-
Collaboration among Multiple Large Language Models for Medical Question Answering22 May 2025 0 repositories listed
-
Continually Self-Improving Language Models for Bariatric Surgery Question--Answering22 May 2025 0 repositories listed
-
CT-Agent: A Multimodal-LLM Agent for 3D CT Radiology Question Answering22 May 2025 0 repositories listed
-
CUB: Benchmarking Context Utilisation Techniques for Language Models22 May 2025 0 repositories listed
-
EarthSE: A Benchmark Evaluating Earth Scientific Exploration Capability for Large Language Models22 May 2025 0 repositories listed
-
Evaluating Large Language Model with Knowledge Oriented Language Specific Simple Question Answering22 May 2025 0 repositories listed
-
Grounding Chest X-Ray Visual Question Answering with Generated Radiology Reports22 May 2025 0 repositories listed
-
Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding22 May 2025 0 repositories listed
-
Steering LVLMs via Sparse Autoencoder for Hallucination Mitigation22 May 2025 0 repositories listed
-
Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools22 May 2025 0 repositories listed
-
UNCLE: Uncertainty Expressions in Long-Form Generation22 May 2025 0 repositories listed
-
VoxRAG: A Step Toward Transcription-Free RAG Systems in Spoken Question Answering22 May 2025 0 repositories listed
-
Zero-Shot Anomaly Detection in Battery Thermal Images Using Visual Question Answering with Prior Knowledge22 May 2025 0 repositories listed
-
BR-TaxQA-R: A Dataset for Question Answering with References for Brazilian Personal Income Tax Law, including case law21 May 2025 0 repositories listed
-
CRAFT: Training-Free Cascaded Retrieval for Tabular QA21 May 2025 0 repositories listed
-
Discovering Pathology Rationale and Token Allocation for Efficient Multimodal Pathology Reasoning21 May 2025 0 repositories listed
-
Human-centered Interactive Learning via MLLMs for Text-to-Image Person Re-identification21 May 2025 0 repositories listed
-
Improving LLM First-Token Predictions in Multiple-Choice Question Answering via Prefilling Attack21 May 2025 0 repositories listed
-
KaFT: Knowledge-aware Fine-tuning for Boosting LLMs' Domain-specific Question-Answering Performance21 May 2025 0 repositories listed
-
LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval21 May 2025 0 repositories listed
-
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets21 May 2025 0 repositories listed
-
Set-LLM: A Permutation-Invariant LLM21 May 2025 0 repositories listed
-
Single LLM, Multiple Roles: A Unified Retrieval-Augmented Generation Framework Using Role-Specific Token Optimization21 May 2025 0 repositories listed
-
Social Bias in Popular Question-Answering Benchmarks21 May 2025 0 repositories listed
-
StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization21 May 2025 0 repositories listed
-
TinyDrive: Multiscale Visual Question Answering with Selective Token Routing for Autonomous Driving21 May 2025 0 repositories listed
-
Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems21 May 2025 0 repositories listed
-
Visual Question Answering on Multiple Remote Sensing Image Modalities21 May 2025 0 repositories listed
-
Abacus: A Cost-Based Optimizer for Semantic Operator Systems20 May 2025 0 repositories listed
-
Automatic Dataset Generation for Knowledge Intensive Question Answering Tasks20 May 2025 0 repositories listed
-
AutoRev: Automatic Peer Review System for Academic Research Papers20 May 2025 0 repositories listed
-
Beyond Chains: Bridging Large Language Models and Knowledge Bases in Complex Question Answering20 May 2025 0 repositories listed
-
Debating for Better Reasoning: An Unsupervised Multimodal Approach20 May 2025 0 repositories listed
-
Domain Adaptation of VLM for Soccer Video Understanding20 May 2025 0 repositories listed
-
Exploring Jailbreak Attacks on LLMs through Intent Concealment and Diversion20 May 2025 0 repositories listed
-
HausaNLP: Current Status, Challenges and Future Directions for Hausa Natural Language Processing20 May 2025 0 repositories listed
-
Interpretable Traces, Unexpected Outcomes: Investigating the Disconnect in Trace-Based Knowledge Distillation20 May 2025 0 repositories listed
-
Memory-Centric Embodied Question Answer20 May 2025 0 repositories listed
-
Reinforcing Question Answering Agents with Minimalist Policy Gradient Optimization20 May 2025 0 repositories listed
-
Studying the Role of Input-Neighbor Overlap in Retrieval-Augmented Language Models Training Efficiency20 May 2025 0 repositories listed
-
The Hallucination Tax of Reinforcement Finetuning20 May 2025 0 repositories listed
-
Toward Effective Reinforcement Learning Fine-Tuning for Medical VQA in Vision-Language Models20 May 2025 0 repositories listed
-
Towards Omnidirectional Reasoning with 360-R1: A Dataset, Benchmark, and GRPO-based Method20 May 2025 0 repositories listed
-
Visual Instruction Bottleneck Tuning20 May 2025 0 repositories listed
-
YESciEval: Robust LLM-as-a-Judge for Scientific Question Answering20 May 2025 0 repositories listed
-
Alignment-Augmented Speculative Decoding with Alignment Sampling and Conditional Verification19 May 2025 0 repositories listed
-
AMAQA: A Metadata-based QA Dataset for RAG Systems19 May 2025 0 repositories listed
-
ChartMuseum: Testing Visual Reasoning Capabilities of Large Vision-Language Models19 May 2025 0 repositories listed
-
KIT's Offline Speech Translation and Instruction Following Submission for IWSLT 202519 May 2025 0 repositories listed
-
ORQA: A Benchmark and Foundation Model for Holistic Operating Room Modeling19 May 2025 0 repositories listed
-
Q²Forge: Minting Competency Questions and SPARQL Queries for Question-Answering Over Knowledge Graphs19 May 2025 0 repositories listed
-
Rethinking Predictive Modeling for LLM Routing: When Simple kNN Beats Complex Learned Routers19 May 2025 0 repositories listed
-
SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models19 May 2025 0 repositories listed
-
The Hidden Structure -- Improving Legal Document Understanding Through Explicit Text Formatting19 May 2025 0 repositories listed
-
Tianyi: A Traditional Chinese Medicine all-rounder language model and its Real-World Clinical Practice19 May 2025 0 repositories listed
-
Understanding Complexity in VideoQA via Visual Program Generation19 May 2025 0 repositories listed
-
Disambiguation in Conversational Question Answering in the Era of LLM: A Survey18 May 2025 0 repositories listed
-
Enhancing Large Language Models with Reward-guided Tree Search for Knowledge Graph Question and Answering18 May 2025 0 repositories listed
-
GMSA: Enhancing Context Compression via Group Merging and Layer Semantic Alignment18 May 2025 0 repositories listed
-
RAGXplain: From Explainable Evaluation to Actionable Guidance of RAG Pipelines18 May 2025 0 repositories listed
-
Table-R1: Region-based Reinforcement Learning for Table Understanding18 May 2025 0 repositories listed
-
AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation17 May 2025 0 repositories listed
-
BELLE: A Bi-Level Multi-Agent Reasoning Framework for Multi-Hop Question Answering17 May 2025 0 repositories listed
-
Beyond Retrieval: Joint Supervision and Multimodal Document Ranking for Textbook Question Answering17 May 2025 0 repositories listed
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.