Browse State-of-the-Art › Hallucination › Papers, page 14
Hallucination
Papers archive 2025-07-28
archive papers tagged: 1,816 · with a code link: 752 · where Syntology ran a sample: 276 (240 with a run with no instrument failure, 36 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (276 of 1,816 tagged: 240 with a run with no instrument failure, 36 where every run was a failure of Syntology's instrument)
Page 14 of 19: papers 1,301 to 1,400 of 1,816, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
WildHallucinations: Evaluating Long-form Factuality in LLMs with Real-World Entity Queries24 Jul 2024 0 repositories listed
-
Generation Constraint Scaling Can Mitigate Hallucination23 Jul 2024 0 repositories listed
-
LawLuo: A Multi-Agent Collaborative Framework for Multi-Round Chinese Legal Consultation23 Jul 2024 0 repositories listed
-
Shared Imagination: LLMs Hallucinate Alike23 Jul 2024 0 repositories listed
-
Developing a Reliable, Fast, General-Purpose Hallucination Detection and Mitigation Service22 Jul 2024 0 repositories listed
-
Multilingual Fine-Grained News Headline Hallucination Detection22 Jul 2024 0 repositories listed
-
Text2Place: Affordance-aware Text Guided Human Placement22 Jul 2024 0 repositories listed
-
Evaluating and Enhancing Trustworthiness of LLMs in Perception Tasks18 Jul 2024 0 repositories listed
-
BEAF: Observing BEfore-AFter Changes to Evaluate Hallucination in Vision-language Models18 Jul 2024 0 repositories listed
-
Black-Box Opinion Manipulation Attacks to Retrieval-Augmented Generation of Large Language Models18 Jul 2024 0 repositories listed
-
Retrieval-Augmented Generation for Natural Language Processing: A Survey18 Jul 2024 0 repositories listed
-
Addressing Image Hallucination in Text-to-Image Generation through Factual Image Retrieval15 Jul 2024 0 repositories listed
-
GraphEval: A Knowledge-Graph Based LLM Hallucination Evaluation Framework15 Jul 2024 0 repositories listed
-
Look Within, Why LLMs Hallucinate: A Causal Perspective14 Jul 2024 0 repositories listed
-
Cohesive Conversations: Enhancing Authenticity in Multi-Agent Simulated Dialogues13 Jul 2024 0 repositories listed
-
On Mitigating Code LLM Hallucinations with API Documentation13 Jul 2024 0 repositories listed
-
DAHRS: Divergence-Aware Hallucination-Remediated SRL Projection12 Jul 2024 0 repositories listed
-
The Two Sides of the Coin: Hallucination Generation and Detection with LLMs as Evaluators for LLMs12 Jul 2024 0 repositories listed
-
Lynx: An Open Source Hallucination Evaluation Model11 Jul 2024 0 repositories listed
-
Fuse, Reason and Verify: Geometry Problem Solving with Parsed Clauses from Diagram10 Jul 2024 0 repositories listed
-
Knowledge Overshadowing Causes Amalgamated Hallucination in Large Language Models10 Jul 2024 0 repositories listed
-
Learning with Instance-Dependent Noisy Labels by Anchor Hallucination and Hard Sample Label Correction10 Jul 2024 0 repositories listed
-
GTP-4o: Modality-prompted Heterogeneous Graph Learning for Omni-modal Biomedical Representation8 Jul 2024 0 repositories listed
-
Vision-Language Models under Cultural and Inclusive Considerations8 Jul 2024 0 repositories listed
-
VideoCoT: A Video Chain-of-Thought Dataset with Active Annotation Tool7 Jul 2024 0 repositories listed
-
Code Hallucination5 Jul 2024 0 repositories listed
-
Classification-Based Automatic HDL Code Generation Using LLMs4 Jul 2024 0 repositories listed
-
Hallucination Detection: Robustly Discerning Reliable Answers in Large Language Models4 Jul 2024 0 repositories listed
-
Query-Guided Self-Supervised Summarization of Nursing Notes4 Jul 2024 0 repositories listed
-
STOC-TOT: Stochastic Tree-of-Thought with Constrained Decoding for Complex Reasoning in Multi-Hop Question Answering4 Jul 2024 0 repositories listed
-
Zero-shot Persuasive Chatbots with LLM-Generated Strategies and Information Retrieval4 Jul 2024 0 repositories listed
-
A Comparative Study of DSL Code Generation: Fine-Tuning vs. Optimized Retrieval Augmentation3 Jul 2024 0 repositories listed
-
FSM: A Finite State Machine Based Zero-Shot Prompting Paradigm for Multi-Hop Question Answering3 Jul 2024 0 repositories listed
-
Pelican: Correcting Hallucination in Vision-LLMs via Claim Decomposition and Program of Thought Verification2 Jul 2024 0 repositories listed
-
Understanding Alignment in Multimodal LLMs: A Comprehensive Study2 Jul 2024 0 repositories listed
-
Free-text Rationale Generation under Readability Level Control1 Jul 2024 0 repositories listed
-
LLM Uncertainty Quantification through Directional Entailment Graph and Claim Level Response Augmentation1 Jul 2024 0 repositories listed
-
The Need for Guardrails with Large Language Models in Medical Safety-Critical Settings: An Artificial Intelligence Application in the Pharmacovigilance Ecosystem1 Jul 2024 0 repositories listed
-
Unveiling Glitches: A Deep Dive into Image Encoding Bugs within CLIP30 Jun 2024 0 repositories listed
-
A Study on Effect of Reference Knowledge Choice in Generating Technical Content Relevant to SAPPhIRE Model Using Large Language Model29 Jun 2024 0 repositories listed
-
PFME: A Modular Approach for Fine-grained Hallucination Detection and Editing of Large Language Models29 Jun 2024 0 repositories listed
-
Applying RLAIF for Code Generation with API-usage in Lightweight LLMs28 Jun 2024 0 repositories listed
-
Prompt-Consistency Image Generation (PCIG): A Unified Framework Integrating LLMs, Knowledge Graphs, and Controllable Diffusion Models24 Jun 2024 0 repositories listed
-
VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models24 Jun 2024 0 repositories listed
-
Does Object Grounding Really Reduce Hallucination of Large Vision-Language Models?20 Jun 2024 0 repositories listed
-
From Descriptive Richness to Bias: Unveiling the Dark Side of Generative Image Caption Enrichment20 Jun 2024 0 repositories listed
-
HIGHT: Hierarchical Graph Tokenization for Molecule-Language Alignment20 Jun 2024 0 repositories listed
-
Large Language Models are Skeptics: False Negative Problem of Input-conflicting Hallucination20 Jun 2024 0 repositories listed
-
Beyond Under-Alignment: Atomic Preference Enhanced Factuality Tuning for Large Language Models18 Jun 2024 0 repositories listed
-
Do More Details Always Introduce More Hallucinations in LVLM-based Image Captioning?18 Jun 2024 0 repositories listed
-
RichRAG: Crafting Rich Responses for Multi-faceted Queries in Retrieval-Augmented Generation18 Jun 2024 0 repositories listed
-
What Matters in Memorizing and Recalling Facts? Multifaceted Benchmarks for Knowledge Probing in Language Models18 Jun 2024 0 repositories listed
-
Hallucination Mitigation Prompts Long-term Video Understanding17 Jun 2024 0 repositories listed
-
InternalInspector I²: Robust Confidence Estimation in LLMs through Internal States17 Jun 2024 0 repositories listed
-
CoMT: Chain-of-Medical-Thought Reduces Hallucination in Medical Report Generation17 Jun 2024 0 repositories listed
-
Mitigating Large Language Model Hallucination with Faithful Finetuning17 Jun 2024 0 repositories listed
-
Teaching Large Language Models to Express Knowledge Boundary from Their Own Signals16 Jun 2024 0 repositories listed
-
Detecting and Evaluating Medical Hallucinations in Large Vision Language Models14 Jun 2024 0 repositories listed
-
Beyond Words: On Large Language Models Actionability in Mission-Critical Risk Analysis11 Jun 2024 0 repositories listed
-
Estimating the Hallucination Rate of Generative AI11 Jun 2024 0 repositories listed
-
HalluDial: A Large-Scale Benchmark for Automatic Dialogue-Level Hallucination Evaluation11 Jun 2024 0 repositories listed
-
Progressive Query Expansion for Retrieval Over Cost-constrained Data Sources11 Jun 2024 0 repositories listed
-
Investigating and Addressing Hallucinations of LLMs in Tasks Involving Negation8 Jun 2024 0 repositories listed
-
Robustness Assessment of Mathematical Reasoning in the Presence of Missing and Contradictory Conditions7 Jun 2024 0 repositories listed
-
ActionReasoningBench: Reasoning about Actions with and without Ramification Constraints6 Jun 2024 0 repositories listed
-
Chaos with Keywords: Exposing Large Language Models Sycophantic Hallucination to Misleading Keywords and Evaluating Defense Strategies6 Jun 2024 0 repositories listed
-
Confabulation: The Surprising Value of Large Language Model Hallucinations6 Jun 2024 0 repositories listed
-
Analyzing LLM Behavior in Dialogue Summarization: Unveiling Circumstantial Hallucination Trends5 Jun 2024 0 repositories listed
-
Towards Detecting LLMs Hallucination via Markov Chain-based Multi-agent Debate Framework5 Jun 2024 0 repositories listed
-
CODE: Contrasting Self-generated Description to Combat Hallucination in Large Multi-modal Models4 Jun 2024 0 repositories listed
-
Enhancing Trust in LLMs: Algorithms for Comparing and Interpreting LLMs4 Jun 2024 0 repositories listed
-
How to Explore with Belief: State Entropy Maximization in POMDPs4 Jun 2024 0 repositories listed
-
Ask-EDA: A Design Assistant Empowered by LLM, Hybrid RAG and Abbreviation De-hallucination3 Jun 2024 0 repositories listed
-
Decompose, Enrich, and Extract! Schema-aware Event Extraction using LLMs3 Jun 2024 0 repositories listed
-
Large Language Model Assisted Optimal Bidding of BESS in FCAS Market: An AI-agent based Approach3 Jun 2024 0 repositories listed
-
Luna: An Evaluation Foundation Model to Catch Language Model Hallucinations with High Accuracy and Low Cost3 Jun 2024 0 repositories listed
-
Comprehensive Evaluation of Large Language Models for Topic Modeling2 Jun 2024 0 repositories listed
-
Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools30 May 2024 0 repositories listed
-
Similarity is Not All You Need: Endowing Retrieval Augmented Generation with Multi Layered Thoughts30 May 2024 0 repositories listed
-
MASSIVE Multilingual Abstract Meaning Representation: A Dataset and Baselines for Hallucination Detection29 May 2024 0 repositories listed
-
MetaToken: Detecting Hallucination in Image Descriptions by Meta Classification29 May 2024 0 repositories listed
-
Two-Layer Retrieval-Augmented Generation Framework for Low-Resource Medical Question Answering Using Reddit Data: Proof-of-Concept Study29 May 2024 0 repositories listed
-
Conv-CoA: Improving Open-domain Question Answering in Large Language Models via Conversational Chain-of-Action28 May 2024 0 repositories listed
-
28 May 2024 0 repositories listed Syntology 6 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
RITUAL: Random Image Transformations as a Universal Anti-hallucination Lever in Large Vision Language Models28 May 2024 0 repositories listed
-
Laboratory-Scale AI: Open-Weight Models are Competitive with ChatGPT Even in Low-Resource Settings27 May 2024 0 repositories listed
-
GeneAgent: Self-verification Language Agent for Gene Set Knowledge Discovery using Domain Databases25 May 2024 0 repositories listed
-
CHARP: Conversation History AwaReness Probing for Knowledge-grounded Dialogue Systems24 May 2024 0 repositories listed
-
Large Language Model Pruning24 May 2024 0 repositories listed
-
Scaling Laws for Discriminative Classification in Large Language Models24 May 2024 0 repositories listed
-
CrossCheckGPT: Universal Hallucination Ranking for Multimodal Foundation Models22 May 2024 0 repositories listed
-
Less for More: Enhanced Feedback-aligned Mixed LLMs for Molecule Caption Generation and Fine-Grained NLI Evaluation22 May 2024 0 repositories listed
-
GameVLM: A Decision-making Framework for Robotic Task Planning Based on Visual Language Models and Zero-sum Games22 May 2024 0 repositories listed
-
Gradient Projection For Continual Parameter-Efficient Tuning22 May 2024 0 repositories listed
-
Presentations are not always linear! GNN meets LLM for Document-to-Presentation Transformation with Attribution21 May 2024 0 repositories listed
-
CT-Eval: Benchmarking Chinese Text-to-Table Performance in Large Language Models20 May 2024 0 repositories listed
-
Evaluating Text-to-Speech Synthesis from a Large Discrete Token-based Speech Language Model16 May 2024 0 repositories listed
-
A Comprehensive Survey of Hallucination in Large Language, Image, Video and Audio Foundation Models15 May 2024 0 repositories listed
-
Word Alignment as Preference for Machine Translation15 May 2024 0 repositories listed
-
ALMol: Aligned Language-Molecule Translation LLMs through Offline Preference Contrastive Optimisation14 May 2024 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.