Browse State-of-the-Art › Question Answering › Papers, page 45
Question Answering
Papers archive 2025-07-28
archive papers tagged: 10,817 · with a code link: 4,171 · where Syntology ran a sample: 1,274 (1,073 with a run with no instrument failure, 201 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (1,274 of 10,817 tagged: 1,073 with a run with no instrument failure, 201 where every run was a failure of Syntology's instrument)
Page 45 of 109: papers 4,401 to 4,500 of 10,817, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
CCNU at SemEval-2025 Task 3: Leveraging Internal and External Knowledge of Large Language Models for Multilingual Hallucination Annotation17 May 2025 0 repositories listed
-
Recursive Question Understanding for Complex Question Answering over Heterogeneous Personal Data17 May 2025 0 repositories listed
-
TinyRS-R1: Compact Multimodal Language Model for Remote Sensing17 May 2025 0 repositories listed
-
Unveiling Knowledge Utilization Mechanisms in LLM-based Retrieval-Augmented Generation17 May 2025 0 repositories listed
-
𝒜LLM4ADD: Unlocking the Capabilities of Audio Large Language Models for Audio Deepfake Detection16 May 2025 0 repositories listed
-
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models16 May 2025 0 repositories listed
-
THELMA: Task Based Holistic Evaluation of Large Language Model Applications-RAG Question Answering16 May 2025 0 repositories listed
-
CAFE: Retrieval Head-based Coarse-to-Fine Information Seeking to Enhance Multi-Document QA Capability15 May 2025 0 repositories listed
-
DIF: A Framework for Benchmarking and Verifying Implicit Bias in LLMs15 May 2025 0 repositories listed
-
End-to-End Vision Tokenizer Tuning15 May 2025 0 repositories listed
-
Enhancing Multi-Image Question Answering via Submodular Subset Selection15 May 2025 0 repositories listed
-
Leveraging Graph Retrieval-Augmented Generation to Support Learners' Understanding of Knowledge Concepts in MOOCs15 May 2025 0 repositories listed
-
What Does Neuro Mean to Cardio? Investigating the Role of Clinical Specialty Data in Medical LLMs15 May 2025 0 repositories listed
-
Omni-R1: Do You Really Need Audio to Fine-Tune Your Audio LLM?14 May 2025 0 repositories listed
-
SafePath: Conformal Prediction for Safe LLM-Based Autonomous Navigation14 May 2025 0 repositories listed
-
The Impact of Large Language Models on Task Automation in Manufacturing Services14 May 2025 0 repositories listed
-
Variational Visual Question Answering14 May 2025 0 repositories listed
-
Fusing Bidirectional Chains of Thought and Reward Mechanisms A Method for Enhancing Question-Answering Capabilities of Large Language Models for Chinese Intangible Cultural Heritage13 May 2025 0 repositories listed
-
WixQA: A Multi-Dataset Benchmark for Enterprise Retrieval-Augmented Generation13 May 2025 0 repositories listed
-
Multi-Domain Audio Question Answering Toward Acoustic Content Reasoning in The DCASE 2025 Challenge12 May 2025 0 repositories listed
-
Private LoRA Fine-tuning of Open-Source LLMs with Homomorphic Encryption12 May 2025 0 repositories listed
-
Relative Overfitting and Accept-Reject Framework12 May 2025 0 repositories listed
-
Visually Interpretable Subtask Reasoning for Visual Question Answering12 May 2025 0 repositories listed
-
Building a Human-Verified Clinical Reasoning Dataset via a Human LLM Hybrid Pipeline for Trustworthy Medical AI11 May 2025 0 repositories listed
-
Multi-Modal Explainable Medical AI Assistant for Trustworthy Human-AI Collaboration11 May 2025 0 repositories listed
-
Overview of the NLPCC 2025 Shared Task 4: Multi-modal, Multilingual, and Multi-hop Medical Instructional Video Question Answering Challenge11 May 2025 0 repositories listed
-
PLHF: Prompt Optimization with Few-Shot Human Feedback11 May 2025 0 repositories listed
-
10 May 2025 0 repositories listed Syntology 4 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; every one of the 4 samples that ran constructed an object rather than computing a result (of 5 harvested samples) · 5 pointer-only (licence)
-
A Grounded Memory System For Smart Personal Assistants9 May 2025 0 repositories listed
-
Assessing Robustness to Spurious Correlations in Post-Training Language Models9 May 2025 0 repositories listed
-
CellVerse: Do Large Language Models Really Understand Cell Biology?9 May 2025 0 repositories listed
-
Document Attribution: Examining Citation Relationships using Large Language Models9 May 2025 0 repositories listed
-
Healthy LLMs? Benchmarking LLM Knowledge of UK Government Public Health Information9 May 2025 0 repositories listed
-
Natural Reflection Backdoor Attack on Vision Language Model for Autonomous Driving9 May 2025 0 repositories listed
-
Towards Developmentally Plausible Rewards: Communicative Success as a Learning Signal for Interactive Language Models9 May 2025 0 repositories listed
-
An Open-Source Dual-Loss Embedding Model for Semantic Retrieval in Higher Education8 May 2025 0 repositories listed
-
Lost in OCR Translation? Vision-Based Approaches to Robust Document Retrieval8 May 2025 0 repositories listed
-
SITE: towards Spatial Intelligence Thorough Evaluation8 May 2025 0 repositories listed
-
Fine-Tuning Large Language Models and Evaluating Retrieval Methods for Improved Question Answering on Building Codes7 May 2025 0 repositories listed
-
HiPerRAG: High-Performance Retrieval Augmented Generation for Scientific Insights7 May 2025 0 repositories listed
-
Q-Heart: ECG Question Answering via Knowledge-Informed Multimodal LLMs7 May 2025 0 repositories listed
-
A Reasoning-Focused Legal Retrieval Benchmark6 May 2025 0 repositories listed
-
VLM Q-Learning: Aligning Vision-Language Models for Interactive Decision-Making6 May 2025 0 repositories listed
-
Structure Causal Models and LLMs Integration in Medical Visual Question Answering5 May 2025 0 repositories listed
-
Task-Oriented Semantic Communication in Large Multimodal Models-based Vehicle Networks5 May 2025 0 repositories listed
-
Compositional Image-Text Matching and Retrieval by Grounding Entities4 May 2025 0 repositories listed
-
Adaptive Token Boundaries: Integrating Human Chunking Mechanisms into Multimodal LLMs3 May 2025 0 repositories listed
-
Knowledge-Augmented Language Models Interpreting Structured Chest X-Ray Findings3 May 2025 0 repositories listed
-
OODTE: A Differential Testing Engine for the ONNX Optimizer3 May 2025 0 repositories listed
-
Beyond Attention: Toward Machines with Intrinsic Higher Mental States2 May 2025 0 repositories listed
-
Grounding Task Assistance with Multimodal Cues from a Single Demonstration2 May 2025 0 repositories listed
-
Transferable Adversarial Attacks on Black-Box Vision-Language Models2 May 2025 0 repositories listed
-
TRAVELER: A Benchmark for Evaluating Temporal Reasoning across Vague, Implicit and Explicit References2 May 2025 0 repositories listed
-
HalluMix: A Task-Agnostic, Multi-Domain Benchmark for Real-World Hallucination Detection1 May 2025 0 repositories listed
-
Calibrating Uncertainty Quantification of Multi-Modal LLMs using Grounding30 Apr 2025 0 repositories listed
-
ConSens: Assessing context grounding in open-book question answering30 Apr 2025 0 repositories listed
-
Zoomer: Adaptive Image Focus Optimization for Black-box MLLM30 Apr 2025 0 repositories listed
-
LLM Enhancer: Merged Approach using Vector Embedding for Reducing Large Language Model Hallucinations with External Knowledge29 Apr 2025 0 repositories listed
-
LMME3DHF: Benchmarking and Evaluating Multimodal 3D Human Face Generation with LMMs29 Apr 2025 0 repositories listed
-
SetKE: Knowledge Editing for Knowledge Elements Overlap29 Apr 2025 0 repositories listed
-
Knowledge Distillation of Domain-adapted LLMs for Question-Answering in Telecom28 Apr 2025 0 repositories listed
-
m-KAILIN: Knowledge-Driven Agentic Scientific Corpus Distillation Framework for Biomedical Large Language Models Training28 Apr 2025 0 repositories listed
-
OpenTCM: A GraphRAG-Empowered LLM-based System for Traditional Chinese Medicine Knowledge Retrieval and Diagnosis28 Apr 2025 0 repositories listed
-
SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning28 Apr 2025 0 repositories listed
-
Pushing the boundary on Natural Language Inference25 Apr 2025 0 repositories listed
-
A Comprehensive Survey of Knowledge-Based Vision Question Answering Systems: The Lifecycle of Knowledge in Visual Reasoning Task24 Apr 2025 0 repositories listed
-
Data-Driven Calibration of Prediction Sets in Large Vision-Language Models Based on Inductive Conformal Prediction24 Apr 2025 0 repositories listed
-
TraveLLaMA: Facilitating Multi-modal Large Language Models to Understand Urban Scenes and Provide Travel Assistance23 Apr 2025 0 repositories listed
-
FinDER: Financial Dataset for Question Answering and Evaluating Retrieval-Augmented Generation22 Apr 2025 0 repositories listed
-
The Great Nugget Recall: Automating Fact Extraction and RAG Evaluation with Large Language Models21 Apr 2025 0 repositories listed
-
Towards Understanding Camera Motions in Any Video21 Apr 2025 0 repositories listed
-
A Hierarchical Framework for Measuring Scientific Paper Innovation via Large Language Models20 Apr 2025 0 repositories listed
-
CoLoTa: A Dataset for Entity-based Commonsense Reasoning over Long-Tail Knowledge20 Apr 2025 0 repositories listed
-
FairSteer: Inference Time Debiasing for LLMs with Dynamic Activation Steering20 Apr 2025 0 repositories listed
-
FinSage: A Multi-aspect RAG System for Financial Filings Question Answering20 Apr 2025 0 repositories listed
-
Neglected Risks: The Disturbing Reality of Children's Images in Datasets and the Urgent Call for Accountability20 Apr 2025 0 repositories listed
-
Bottom-Up Synthesis of Knowledge-Grounded Task-Oriented Dialogues with Iteratively Self-Refined Prompts19 Apr 2025 0 repositories listed
-
LegalRAG: A Hybrid RAG System for Multilingual Legal Information Retrieval19 Apr 2025 0 repositories listed
-
SConU: Selective Conformal Uncertainty in Large Language Models19 Apr 2025 0 repositories listed
-
Accommodate Knowledge Conflicts in Retrieval-augmented LLMs: Towards Reliable Response Generation in the Wild17 Apr 2025 0 repositories listed
-
ChartQA-X: Generating Explanations for Charts17 Apr 2025 0 repositories listed
-
Hadamard product in deep learning: Introduction, Advances and Challenges17 Apr 2025 0 repositories listed
-
WebLists: Extracting Structured Information From Complex Interactive Websites Using Executable LLM Agents17 Apr 2025 0 repositories listed
-
Bridging the Semantic Gaps: Improving Medical VQA Consistency with LLM-Augmented Question Sets16 Apr 2025 0 repositories listed
-
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching16 Apr 2025 0 repositories listed
-
Mitigating LLM Hallucinations with Knowledge Graphs: A Case Study16 Apr 2025 0 repositories listed
-
16 Apr 2025 0 repositories listed
-
AskQE: Question Answering as Automatic Evaluation for Machine Translation15 Apr 2025 0 repositories listed
-
Exploring the Role of Knowledge Graph-Based RAG in Japanese Medical Question Answering with Small-Scale LLMs15 Apr 2025 0 repositories listed
-
From Misleading Queries to Accurate Answers: A Three-Stage Fine-Tuning Method for LLMs15 Apr 2025 0 repositories listed
-
LVLM_CSP: Accelerating Large Vision Language Models via Clustering, Scattering, and Pruning for Reasoning Segmentation15 Apr 2025 0 repositories listed
-
Streamlining Biomedical Research with Specialized LLMs15 Apr 2025 0 repositories listed
-
Building Trustworthy Multimodal AI: A Review of Fairness, Transparency, and Ethics in Vision-Language Tasks14 Apr 2025 0 repositories listed
-
Constructing Micro Knowledge Graphs from Technical Support Documents14 Apr 2025 0 repositories listed
-
Hallucination Detection in LLMs via Topological Divergence on Attention Graphs14 Apr 2025 0 repositories listed
-
MMKB-RAG: A Multi-Modal Knowledge-Based Retrieval-Augmented Generation Framework14 Apr 2025 0 repositories listed
-
Reasoning Court: Combining Reasoning, Action, and Judgment for Multi-Hop Reasoning14 Apr 2025 0 repositories listed
-
See or Recall: A Sanity Check for the Role of Vision in Solving Visualization Question Answer Tasks with Multimodal LLMs14 Apr 2025 0 repositories listed
-
VDocRAG: Retrieval-Augmented Generation over Visually-Rich Documents14 Apr 2025 0 repositories listed
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.