Browse State-of-the-Art › Hallucination › Papers, page 12
Hallucination
Papers archive 2025-07-28
archive papers tagged: 1,816 · with a code link: 752 · where Syntology ran a sample: 276 (240 with a run with no instrument failure, 36 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (276 of 1,816 tagged: 240 with a run with no instrument failure, 36 where every run was a failure of Syntology's instrument)
Page 12 of 19: papers 1,101 to 1,200 of 1,816, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Steps are all you need: Rethinking STEM Education with Prompt Engineering6 Dec 2024 0 repositories listed
-
Verb Mirage: Unveiling and Assessing Verb Concept Hallucinations in Multimodal Large Language Models6 Dec 2024 0 repositories listed
-
Deep priors for satellite image restoration with accurate uncertainties5 Dec 2024 0 repositories listed
-
GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration5 Dec 2024 0 repositories listed
-
Reducing Tool Hallucination via Reliability Alignment5 Dec 2024 0 repositories listed
-
VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding4 Dec 2024 0 repositories listed
-
Who Brings the Frisbee: Probing Hidden Hallucination Factors in Large Vision-Language Model via Causality Analysis4 Dec 2024 0 repositories listed
-
An Evolutionary Large Language Model for Hallucination Mitigation3 Dec 2024 0 repositories listed
-
CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy3 Dec 2024 0 repositories listed
-
AI Benchmarks and Datasets for LLM Evaluation2 Dec 2024 0 repositories listed
-
Beyond Logit Lens: Contextual Embeddings for Robust Hallucination Detection & Grounding in VLMs28 Nov 2024 0 repositories listed
-
DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models27 Nov 2024 0 repositories listed
-
OPCap:Object-aware Prompting Captioning27 Nov 2024 0 repositories listed
-
A Topic-level Self-Correctional Approach to Mitigate Hallucinations in MLLMs26 Nov 2024 0 repositories listed
-
AI2T: Building Trustable AI Tutors by Interactively Teaching a Self-Aware Learning Agent26 Nov 2024 0 repositories listed
-
Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach26 Nov 2024 0 repositories listed
-
Meaningless is better: hashing bias-inducing words in LLM prompts improves performance in logical reasoning and statistical learning26 Nov 2024 0 repositories listed
-
26 Nov 2024 0 repositories listed
-
Enhancing Multi-Agent Consensus through Third-Party LLM Integration: Analyzing Uncertainty and Mitigating Hallucinations in Large Language Models25 Nov 2024 0 repositories listed
-
Detecting Hallucinations in Virtual Histology with Neural Precursors22 Nov 2024 0 repositories listed
-
ICT: Image-Object Cross-Level Trusted Intervention for Mitigating Object Hallucination in Large Vision-Language Models22 Nov 2024 0 repositories listed
-
Leveraging LLMs for Legacy Code Modernization: Challenges and Opportunities for LLM-Generated Documentation22 Nov 2024 0 repositories listed
-
Sycophancy in Large Language Models: Causes and Mitigations22 Nov 2024 0 repositories listed
-
CATCH: Complementary Adaptive Token-level Contrastive Decoding to Mitigate Hallucinations in LVLMs19 Nov 2024 0 repositories listed
-
Can Open-source LLMs Enhance Data Synthesis for Toxic Detection?: An Experimental Study18 Nov 2024 0 repositories listed
-
Mitigating Knowledge Conflicts in Language Model-Driven Question Answering18 Nov 2024 0 repositories listed
-
Enabling Explainable Recommendation in E-commerce with LLM-powered Product Knowledge Graph17 Nov 2024 0 repositories listed
-
INVARLLM: LLM-assisted Physical Invariant Extraction for Cyber-Physical Systems Anomaly Detection17 Nov 2024 0 repositories listed
-
Understanding Multimodal LLMs: the Mechanistic Interpretability of Llava in Visual Question Answering17 Nov 2024 0 repositories listed
-
A Novel Approach to Eliminating Hallucinations in Large Language Model-Assisted Causal Discovery16 Nov 2024 0 repositories listed
-
Chain-of-Programming (CoP) : Empowering Large Language Models for Geospatial Code Generation16 Nov 2024 0 repositories listed
-
ViBe: A Text-to-Video Benchmark for Evaluating Hallucination in Large Multimodal Models16 Nov 2024 0 repositories listed
-
Layer Importance and Hallucination Analysis in Large Language Models via Enhanced Activation Variance-Sparsity15 Nov 2024 0 repositories listed
-
Mitigating Hallucination in Multimodal Large Language Model via Hallucination-targeted Direct Preference Optimization15 Nov 2024 0 repositories listed
-
Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs15 Nov 2024 0 repositories listed
-
LLM Hallucination Reasoning with Zero-shot Knowledge Test14 Nov 2024 0 repositories listed
-
On the Limits of Language Generation: Trade-Offs Between Hallucination and Mode Collapse14 Nov 2024 0 repositories listed
-
SHARP: Unlocking Interactive Hallucination via Stance Transfer in Role-Playing Agents12 Nov 2024 0 repositories listed
-
Trustful LLMs: Customizing and Grounding Text Generation with Knowledge Bases and Dual Decoders12 Nov 2024 0 repositories listed
-
Evaluating the Accuracy of Chatbots in Financial Literature11 Nov 2024 0 repositories listed
-
Invar-RAG: Invariant LLM-aligned Retrieval for Better Generation11 Nov 2024 0 repositories listed
-
Prompt-Efficient Fine-Tuning for GPT-like Deep Models to Reduce Hallucination and to Improve Reproducibility in Scientific Text Generation Using Stochastic Optimisation Techniques10 Nov 2024 0 repositories listed
-
Mitigating Hallucination with ZeroG: An Advanced Knowledge Management Engine8 Nov 2024 0 repositories listed
-
Seeing Through the Fog: A Cost-Effectiveness Analysis of Hallucination Detection Systems8 Nov 2024 0 repositories listed
-
AMSnet-KG: A Netlist Dataset for LLM-based AMS Circuit Auto-Design Using Knowledge Graph RAG7 Nov 2024 0 repositories listed
-
LLM-R: A Framework for Domain-Adaptive Maintenance Scheme Generation Combining Hierarchical Agents and RAG7 Nov 2024 0 repositories listed
-
Prompt-Guided Internal States for Hallucination Detection of Large Language Models7 Nov 2024 0 repositories listed
-
Fine-Grained Guidance for Retrievers: Leveraging LLMs' Feedback in Retrieval-Augmented Generation6 Nov 2024 0 repositories listed
-
Fine-Tuning Vision-Language Model for Automated Engineering Drawing Information Extraction6 Nov 2024 0 repositories listed
-
H-POPE: Hierarchical Polling-based Probing Evaluation of Hallucinations in Large Vision-Language Models6 Nov 2024 0 repositories listed
-
Automated, LLM enabled extraction of synthesis details for reticular materials from scientific literature5 Nov 2024 0 repositories listed
-
Leveraging Vision-Language Models for Manufacturing Feature Recognition in CAD Designs5 Nov 2024 0 repositories listed
-
VERITAS: A Unified Approach to Reliability Evaluation5 Nov 2024 0 repositories listed
-
CleAR: Robust Context-Guided Generative Lighting Estimation for Mobile Augmented Reality4 Nov 2024 0 repositories listed
-
Improving Scientific Hypothesis Generation with Knowledge Grounded Large Language Models4 Nov 2024 0 repositories listed
-
Robust plug-and-play methods for highly accelerated non-Cartesian MRI reconstruction4 Nov 2024 0 repositories listed
-
RadFlag: A Black-Box Hallucination Detection Method for Medical Vision Language Models1 Nov 2024 0 repositories listed
-
Towards Multi-Source Retrieval-Augmented Generation via Synergizing Reasoning and Preference-Driven Retrieval1 Nov 2024 0 repositories listed
-
Exploring the Knowledge Mismatch Hypothesis: Hallucination Propensity in Small Models Fine-tuned on Data from Larger Models31 Oct 2024 0 repositories listed
-
Improbable Bigrams Expose Vulnerabilities of Incomplete Tokens in Byte-Level Tokenizers31 Oct 2024 0 repositories listed
-
EF-LLM: Energy Forecasting LLM with AI-assisted Automation, Enhanced Sparse Prediction, Hallucination Detection30 Oct 2024 0 repositories listed
-
VisAidMath: Benchmarking Visual-Aided Mathematical Reasoning30 Oct 2024 0 repositories listed
-
FactBench: A Dynamic Benchmark for In-the-Wild Language Model Factuality Evaluation29 Oct 2024 0 repositories listed
-
MARCO: Multi-Agent Real-time Chat Orchestration29 Oct 2024 0 repositories listed
-
A Perspective for Adapting Generalist AI to Specialized Medical AI Applications and Their Challenges28 Oct 2024 0 repositories listed
-
A Debate-Driven Experiment on LLM Hallucinations and Accuracy25 Oct 2024 0 repositories listed
-
Conditional Hallucinations for Image Compression25 Oct 2024 0 repositories listed
-
Investigating the Role of Prompting and External Tools in Hallucination Rates of Large Language Models25 Oct 2024 0 repositories listed
-
MaCTG: Multi-Agent Collaborative Thought Graph for Automatic Programming25 Oct 2024 0 repositories listed
-
23 Oct 2024 0 repositories listed Syntology 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Leveraging the Domain Adaptation of Retrieval Augmented Generation Models for Question Answering and Reducing Hallucination23 Oct 2024 0 repositories listed
-
Multilingual Hallucination Gaps in Large Language Models23 Oct 2024 0 repositories listed
-
Do Robot Snakes Dream like Electric Sheep? Investigating the Effects of Architectural Inductive Biases on Hallucination22 Oct 2024 0 repositories listed
-
Fine-Tuning Large Language Models to Appropriately Abstain with Semantic Entropy22 Oct 2024 0 repositories listed
-
GeoCode-GPT: A Large Language Model for Geospatial Code Generation Tasks22 Oct 2024 0 repositories listed
-
IPL: Leveraging Multimodal Large Language Models for Intelligent Product Listing22 Oct 2024 0 repositories listed
-
Privacy-hardened and hallucination-resistant synthetic data generation with logic-solvers22 Oct 2024 0 repositories listed
-
SG-FSM: A Self-Guiding Zero-Shot Prompting Paradigm for Multi-Hop Question Answering Based on Finite State Machine22 Oct 2024 0 repositories listed
-
Large language models enabled multiagent ensemble method for efficient EHR data labeling21 Oct 2024 0 repositories listed
-
Learning to Generate and Evaluate Fact-checking Explanations with Transformers21 Oct 2024 0 repositories listed
-
Mitigating Hallucinations of Large Language Models in Medical Information Extraction via Contrastive Decoding21 Oct 2024 0 repositories listed
-
NetSafe: Exploring the Topological Safety of Multi-agent Networks21 Oct 2024 0 repositories listed
-
Towards a Reliable Offline Personal AI Assistant for Long Duration Spaceflight21 Oct 2024 0 repositories listed
-
A Survey of Hallucination in Large Visual Language Models20 Oct 2024 0 repositories listed
-
Hallucination Detox: Sensitivity Dropout (SenD) for Large Language Model Training20 Oct 2024 0 repositories listed
-
Coarse-to-Fine Highlighting: Reducing Knowledge Hallucination in Large Language Models19 Oct 2024 0 repositories listed
-
Good Parenting is all you need -- Multi-agentic LLM Hallucination Mitigation18 Oct 2024 0 repositories listed
-
ETF: An Entity Tracing Framework for Hallucination Detection in Code Summaries17 Oct 2024 0 repositories listed
-
Utilizing Large Language Models in an iterative paradigm with domain feedback for zero-shot molecule optimization17 Oct 2024 0 repositories listed
-
Controlled Automatic Task-Specific Synthetic Data Generation for Hallucination Detection16 Oct 2024 0 repositories listed
-
Iter-AHMCL: Alleviate Hallucination for Large Language Model via Iterative Model-level Contrastive Learning16 Oct 2024 0 repositories listed
-
On A Scale From 1 to 5: Quantifying Hallucination in Faithfulness Evaluation16 Oct 2024 0 repositories listed
-
What Do LLMs Need to Understand Graphs: A Survey of Parametric Representation of Graphs16 Oct 2024 0 repositories listed
-
RosePO: Aligning LLM-based Recommenders with Human Values16 Oct 2024 0 repositories listed
-
When Not to Answer: Evaluating Prompts on GPT Models for Effective Abstention in Unanswerable Math Word Problems16 Oct 2024 0 repositories listed
-
AGENTiGraph: An Interactive Knowledge Graph Platform for LLM-based Chatbots Utilizing Private Data15 Oct 2024 0 repositories listed
-
Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs15 Oct 2024 0 repositories listed
-
LargePiG: Your Large Language Model is Secretly a Pointer Generator15 Oct 2024 0 repositories listed
-
Magnifier Prompt: Tackling Multimodal Hallucination via Extremely Simple Instructions15 Oct 2024 0 repositories listed
-
On the Capacity of Citation Generation by Large Language Models15 Oct 2024 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.