Browse State-of-the-Art › Question Answering › Papers, page 43
Question Answering
Papers archive 2025-07-28
archive papers tagged: 10,817 · with a code link: 4,171 · where Syntology ran a sample: 1,274 (1,073 with a run with no instrument failure, 201 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (1,274 of 10,817 tagged: 1,073 with a run with no instrument failure, 201 where every run was a failure of Syntology's instrument)
Page 43 of 109: papers 4,201 to 4,300 of 10,817, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Enhancing Biosecurity in Tamper-Resistant Large Language Models With Quantum Gradient Descent23 Jun 2025 0 repositories listed
-
Semantic similarity estimation for domain specific data using BERT and other techniques23 Jun 2025 0 repositories listed
-
GEMeX-ThinkVG: Towards Thinking with Visual Grounding in Medical VQA via Reinforcement Learning22 Jun 2025 0 repositories listed
-
Mental Health Equity in LLMs: Leveraging Multi-Hop Question Answering to Detect Amplified and Silenced Perspectives22 Jun 2025 0 repositories listed
-
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations21 Jun 2025 0 repositories listed
-
General-Purpose Robotic Navigation via LVLM-Orchestrated Perception, Reasoning, and Acting20 Jun 2025 0 repositories listed
-
Can Common VLMs Rival Medical VLMs? Evaluation and Strategic Insights19 Jun 2025 0 repositories listed
-
How Far Can Off-the-Shelf Multimodal Large Language Models Go in Online Episodic Memory Question Answering?19 Jun 2025 0 repositories listed
-
From RAG to Agentic: Validating Islamic-Medicine Responses with LLM Agents18 Jun 2025 0 repositories listed
-
MEGC2025: Micro-Expression Grand Challenge on Spot Then Recognize and Visual Question Answering18 Jun 2025 0 repositories listed
-
WikiMixQA: A Multimodal Benchmark for Question Answering over Tables and Charts18 Jun 2025 0 repositories listed
-
CAPO: Reinforcing Consistent Reasoning in Medical Decision-Making15 Jun 2025 0 repositories listed
-
AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making14 Jun 2025 0 repositories listed
-
MM-R5: MultiModal Reasoning-Enhanced ReRanker via Reinforcement Learning for Document Retrieval14 Jun 2025 0 repositories listed
-
A Fast, Reliable, and Secure Programming Language for LLM Agents with Code Actions13 Jun 2025 0 repositories listed
-
Benchmarking Multimodal LLMs on Recognition and Understanding over Chemical Tables13 Jun 2025 0 repositories listed
-
Instruction Tuning and CoT Prompting for Contextual Medical QA with LLMs13 Jun 2025 0 repositories listed
-
MTabVQA: Evaluating Multi-Tabular Reasoning of Language Models in Visual Space13 Jun 2025 0 repositories listed
-
AC/DC: LLM-based Audio Comprehension via Dialogue Continuation12 Jun 2025 0 repositories listed
-
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos12 Jun 2025 0 repositories listed
-
Can We Infer Confidential Properties of Training Data from LLMs?12 Jun 2025 0 repositories listed
-
CogStream: Context-guided Streaming Video Question Answering12 Jun 2025 0 repositories listed
-
Different Questions, Different Models: Fine-Grained Evaluation of Uncertainty and Calibration in Clinical QA with LLMs12 Jun 2025 0 repositories listed
-
EQA-RM: A Generative Embodied Reward Model with Test-time Scaling12 Jun 2025 0 repositories listed
-
HalLoc: Token-level Localization of Hallucinations for Vision Language Models12 Jun 2025 0 repositories listed
-
MedSeg-R: Reasoning Segmentation in Medical Images with Multimodal Large Language Models12 Jun 2025 0 repositories listed
-
Neural at ArchEHR-QA 2025: Agentic Prompt Optimization for Evidence-Grounded Clinical Question Answering12 Jun 2025 0 repositories listed
-
ChartReasoner: Code-Driven Modality Bridging for Long-Chain Reasoning in Chart Question Answering11 Jun 2025 0 repositories listed
-
CheckManual: A New Challenge and Benchmark for Manual-based Appliance Manipulation11 Jun 2025 0 repositories listed
-
Kvasir-VQA-x1: A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy11 Jun 2025 0 repositories listed
-
Provoking Multi-modal Few-Shot LVLM via Exploration-Exploitation In-Context Learning11 Jun 2025 0 repositories listed
-
mSTEB: Massively Multilingual Evaluation of LLMs on Speech and Text Tasks10 Jun 2025 0 repositories listed
-
WIP: Large Language Model-Enhanced Smart Tutor for Undergraduate Circuit Analysis10 Jun 2025 0 repositories listed
-
Efficient Context Selection for Long-Context QA: No Tuning, No Iteration, Just Adaptive-k10 Jun 2025 0 repositories listed
-
Improved LLM Agents for Financial Document Question Answering10 Jun 2025 0 repositories listed
-
PhyBlock: A Progressive Benchmark for Physical Understanding and Planning via 3D Block Assembly10 Jun 2025 0 repositories listed
-
VersaVid-R1: A Versatile Video Understanding and Reasoning Model from Question Answering to Captioning Tasks10 Jun 2025 0 repositories listed
-
LEANN: A Low-Storage Vector Index9 Jun 2025 0 repositories listed
-
Aligning Text, Images, and 3D Structure Token-by-Token9 Jun 2025 0 repositories listed
-
Federated In-Context Learning: Iterative Refinement for Improved Answer Quality9 Jun 2025 0 repositories listed
-
9 Jun 2025 0 repositories listed
-
Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning8 Jun 2025 0 repositories listed
-
Learning to Clarify by Reinforcement Learning Through Reward-Weighted Fine-Tuning8 Jun 2025 0 repositories listed
-
Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning8 Jun 2025 0 repositories listed
-
Meta-Adaptive Prompt Distillation for Few-Shot Visual Question Answering7 Jun 2025 0 repositories listed
-
The State-of-the-Art in Lifelog Retrieval: A Review of Progress at the ACM Lifelog Search Challenge Workshop 2022-247 Jun 2025 0 repositories listed
-
BioMol-MQA: A Multi-Modal Question Answering Dataset For LLM Reasoning Over Bio-Molecular Interactions6 Jun 2025 0 repositories listed
-
DynamicMind: A Tri-Mode Thinking System for Large Language Models6 Jun 2025 0 repositories listed
-
MAPLE: Multi-Agent Adaptive Planning with Long-Term Memory for Table Reasoning6 Jun 2025 0 repositories listed
-
Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques6 Jun 2025 0 repositories listed
-
Ontology-based knowledge representation for bone disease diagnosis: a foundation for safe and sustainable medical artificial intelligence systems5 Jun 2025 0 repositories listed
-
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs5 Jun 2025 0 repositories listed
-
Multiple-Choice Question Generation Using Large Language Models: Methodology and Educator Insights5 Jun 2025 0 repositories listed
-
TextVidBench: A Benchmark for Long Video Scene Text Understanding5 Jun 2025 0 repositories listed
-
Towards Efficient Speech-Text Jointly Decoding within One Speech Language Model4 Jun 2025 0 repositories listed
-
Plugging Schema Graph into Multi-Table QA: A Human-Guided Framework for Reducing LLM Reliance4 Jun 2025 0 repositories listed
-
ReXVQA: A Large-scale Visual Question Answering Benchmark for Generalist Chest X-ray Understanding4 Jun 2025 0 repositories listed
-
A Multi-Agent Framework for Mitigating Dialect Biases in Privacy Policy Question-Answering Systems3 Jun 2025 0 repositories listed
-
ExpertLongBench: Benchmarking Language Models on Expert-Level Long-Form Generation Tasks with Structured Checklists2 Jun 2025 0 repositories listed
-
Hanfu-Bench: A Multimodal Benchmark on Cross-Temporal Cultural Understanding and Transcreation2 Jun 2025 0 repositories listed
-
iQUEST: An Iterative Question-Guided Framework for Knowledge Base Question Answering2 Jun 2025 0 repositories listed
-
Learning Sparsity for Effective and Efficient Music Performance Question Answering2 Jun 2025 0 repositories listed
-
A Graph-Retrieval-Augmented Generation Framework Enhances Decision-Making in the Circular Economy1 Jun 2025 0 repositories listed
-
anyECG-chat: A Generalist ECG-MLLM for Flexible ECG Input and Multi-Task Understanding1 Jun 2025 0 repositories listed
-
Fast or Slow? Integrating Fast Intuition and Deliberate Thinking for Enhancing Visual Question Answering1 Jun 2025 0 repositories listed
-
A Simple Linear Patch Revives Layer-Pruned Large Language Models30 May 2025 0 repositories listed
-
ClinBench-HPB: A Clinical Benchmark for Evaluating LLMs in Hepato-Pancreato-Biliary Diseases30 May 2025 0 repositories listed
-
Exploring the Impact of Occupational Personas on Domain-Specific QA30 May 2025 0 repositories listed
-
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering30 May 2025 0 repositories listed
-
Improving Reliability and Explainability of Medical Question Answering through Atomic Fact Checking in Retrieval-Augmented LLMs30 May 2025 0 repositories listed
-
LaMP-QA: A Benchmark for Personalized Long-form Question Answering30 May 2025 0 repositories listed
-
Light as Deception: GPT-driven Natural Relighting Against Vision-Language Pre-training Models30 May 2025 0 repositories listed
-
MedOrch: Medical Diagnosis with Tool-Augmented Reasoning Agents for Flexible Extensibility30 May 2025 0 repositories listed
-
Pangu DeepDiver: Adaptive Search Intensity Scaling via Open-Web Reinforcement Learning30 May 2025 0 repositories listed
-
Vision LLMs Are Bad at Hierarchical Visual Understanding, and LLMs Are the Bottleneck30 May 2025 0 repositories listed
-
VUDG: A Dataset for Video Understanding Domain Generalization30 May 2025 0 repositories listed
-
ChartMind: A Comprehensive Benchmark for Complex Real-world Multimodal Chart Question Answering29 May 2025 0 repositories listed
-
Data-efficient Meta-models for Evaluation of Context-based Questions and Answers in LLMs29 May 2025 0 repositories listed
-
Diagnosing and Addressing Pitfalls in KG-RAG Datasets: Toward More Reliable Benchmarking29 May 2025 0 repositories listed
-
Differential Information: An Information-Theoretic Perspective on Preference Optimization29 May 2025 0 repositories listed
-
Fortune: Formula-Driven Reinforcement Learning for Symbolic Table Reasoning in Language Models29 May 2025 0 repositories listed
-
From Chat Logs to Collective Insights: Aggregative Question Answering29 May 2025 0 repositories listed
-
Let's Reason Formally: Natural-Formal Hybrid Reasoning Enhances LLM's Math Capability29 May 2025 0 repositories listed
-
MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering29 May 2025 0 repositories listed
-
mRAG: Elucidating the Design Space of Multi-modal Retrieval-Augmented Generation29 May 2025 0 repositories listed
-
Reinforcement Learning for Better Verbalized Confidence in Long-Form Generation29 May 2025 0 repositories listed
-
Spoken question answering for visual queries29 May 2025 0 repositories listed
-
TCM-Ladder: A Benchmark for Multimodal Question Answering on Traditional Chinese Medicine29 May 2025 0 repositories listed
-
3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language Model28 May 2025 0 repositories listed
-
Agent-UniRAG: A Trainable Open-Source LLM Agent Framework for Unified Retrieval-Augmented Generation Systems28 May 2025 0 repositories listed
-
ER-REASON: A Benchmark Dataset for LLM-Based Clinical Reasoning in the Emergency Room28 May 2025 0 repositories listed
-
EvolveSearch: An Iterative Self-Evolving Search Agent28 May 2025 0 repositories listed
-
Improving QA Efficiency with DistilBERT: Fine-Tuning and Inference on mobile Intel CPUs28 May 2025 0 repositories listed
-
NegVQA: Can Vision Language Models Understand Negation?28 May 2025 0 repositories listed
-
Read Your Own Mind: Reasoning Helps Surface Self-Confidence Signals in LLMs28 May 2025 0 repositories listed
-
StressTest: Can YOUR Speech LM Handle the Stress?28 May 2025 0 repositories listed
-
Structured Memory Mechanisms for Stable Context Representation in Large Language Models28 May 2025 0 repositories listed
-
DriveRX: A Vision-Language Reasoning Model for Cross-Task Autonomous Driving27 May 2025 0 repositories listed
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.