Browse State-of-the-Art › Image Captioning › Papers, page 10
Image Captioning
Papers archive 2025-07-28
archive papers tagged: 1,878 · with a code link: 774 · where Syntology ran a sample: 243 (201 with a run with no instrument failure, 42 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (243 of 1,878 tagged: 201 with a run with no instrument failure, 42 where every run was a failure of Syntology's instrument)
Page 10 of 19: papers 901 to 1,000 of 1,878, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularity23 Nov 2024 0 repositories listed
-
Uterine Ultrasound Image Captioning Using Deep Learning Techniques21 Nov 2024 0 repositories listed
-
AI Flow at the Network Edge19 Nov 2024 0 repositories listed
-
Mitigating Perception Bias: A Training-Free Approach to Enhance LMM for Image Quality Assessment19 Nov 2024 0 repositories listed
-
17 Nov 2024 0 repositories listed
-
Cross-Modal Consistency in Multimodal Large Language Models14 Nov 2024 0 repositories listed
-
BLIP3-KALE: Knowledge Augmented Large-Scale Dense Captions12 Nov 2024 0 repositories listed
-
Grounded Video Caption Generation12 Nov 2024 0 repositories listed
-
ViTOC: Vision Transformer and Object-aware Captioner9 Nov 2024 0 repositories listed
-
Image2Text2Image: A Novel Framework for Label-Free Evaluation of Image-to-Text Generation with Text-to-Image Diffusion Models8 Nov 2024 0 repositories listed
-
Seeing is Deceiving: Exploitation of Visual Pathways in Multi-Modal Language Models7 Nov 2024 0 repositories listed
-
RS-MoE: Mixture of Experts for Remote Sensing Image Captioning and Visual Question Answering3 Nov 2024 0 repositories listed
-
Designing a Robust Radiology Report Generation System2 Nov 2024 0 repositories listed
-
Aggregate-and-Adapt Natural Language Prompts for Downstream Generalization of CLIP31 Oct 2024 0 repositories listed
-
Large Language Model Benchmarks in Medical Tasks28 Oct 2024 0 repositories listed
-
Image Generation from Image Captioning -- Invertible Approach26 Oct 2024 0 repositories listed
-
Decoding Diffusion: A Scalable Framework for Unsupervised Analysis of Latent Space Biases and Representations Using Natural Language Prompts25 Oct 2024 0 repositories listed
-
Backdoor in Seconds: Unlocking Vulnerabilities in Large Pre-trained Models via Model Editing23 Oct 2024 0 repositories listed
-
MI-VisionShot: Few-shot adaptation of vision-language models for slide-level classification of histopathological images21 Oct 2024 0 repositories listed
-
VipAct: Visual-Perception Enhancement via Specialized VLM Agent Collaboration and Tool-use21 Oct 2024 0 repositories listed
-
Hiding-in-Plain-Sight (HiPS) Attack on CLIP for Targetted Object Removal from Images16 Oct 2024 0 repositories listed
-
MMCFND: Multimodal Multilingual Caption-aware Fake News Detection for Low-resource Indic Languages14 Oct 2024 0 repositories listed
-
12 Oct 2024 0 repositories listed
-
AnyAttack: Towards Large-scale Self-supervised Adversarial Attacks on Vision-language Models7 Oct 2024 0 repositories listed
-
AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark4 Oct 2024 0 repositories listed
-
Backdooring Vision-Language Models with Out-Of-Distribution Data2 Oct 2024 0 repositories listed
-
Quantifying the Gaps Between Translation and Native Perception in Training for Multimodal, Multilingual Retrieval2 Oct 2024 0 repositories listed
-
DENEB: A Hallucination-Robust Automatic Evaluation Metric for Image Captioning28 Sep 2024 0 repositories listed
-
TrojVLM: Backdoor Attack Against Vision Language Models28 Sep 2024 0 repositories listed
-
A TextGCN-Based Decoding Approach for Improving Remote Sensing Image Captioning27 Sep 2024 0 repositories listed
-
Enhancing Explainability in Multimodal Large Language Models Using Ontological Context27 Sep 2024 0 repositories listed
-
Brotherhood at WMT 2024: Leveraging LLM-Generated Contextual Conversations for Cross-Lingual Image Captioning23 Sep 2024 0 repositories listed
-
Effectively Enhancing Vision Language Large Models by Prompt Augmentation and Caption Utilization22 Sep 2024 0 repositories listed
-
@Bench: Benchmarking Vision-Language Models for Human-centered Assistive Technology21 Sep 2024 0 repositories listed
-
FullAnno: A Data Engine for Enhancing Image Comprehension of MLLMs20 Sep 2024 0 repositories listed
-
LLMs Can Check Their Own Results to Mitigate Hallucinations in Traffic Understanding Tasks19 Sep 2024 0 repositories listed
-
Vision Language Models Can Parse Floor Plan Maps19 Sep 2024 0 repositories listed
-
Decoding Style: Efficient Fine-Tuning of LLMs for Image-Guided Outfit Recommendation with Preference18 Sep 2024 0 repositories listed
-
NEVLP: Noise-Robust Framework for Efficient Vision-Language Pre-training15 Sep 2024 0 repositories listed
-
Evaluating authenticity and quality of image captions via sentiment and semantic analyses14 Sep 2024 0 repositories listed
-
BLens: Contrastive Captioning of Binary Functions using Ensemble Embedding12 Sep 2024 0 repositories listed
-
Bridging Paintings and Music -- Exploring Emotion based Music Generation through Paintings12 Sep 2024 0 repositories listed
-
Securing Vision-Language Models with a Robust Encoder Against Jailbreak and Adversarial Attacks11 Sep 2024 0 repositories listed
-
Mitigating Hallucination in Visual-Language Models via Re-Balancing Contrastive Decoding10 Sep 2024 0 repositories listed
-
PoseEmbroider: Towards a 3D, Visual, Semantic-aware Human Pose Representation10 Sep 2024 0 repositories listed
-
MLLM-LLaVA-FL: Multimodal Large Language Model Assisted Federated Learning9 Sep 2024 0 repositories listed
-
FODA-PG for Enhanced Medical Imaging Narrative Generation: Adaptive Differentiation of Normal and Abnormal Attributes6 Sep 2024 0 repositories listed
-
No Detail Left Behind: Revisiting Self-Retrieval for Fine-Grained Image Captioning4 Sep 2024 0 repositories listed
-
Fluent and Accurate Image Captioning with a Self-Trained Reward Model29 Aug 2024 0 repositories listed
-
Hand1000: Generating Realistic Hands from Text with Only 1,000 Images28 Aug 2024 0 repositories listed
-
Pixels to Prose: Understanding the art of Image Captioning28 Aug 2024 0 repositories listed
-
Bidirectional Awareness Induction in Autoregressive Seq2Seq Models25 Aug 2024 0 repositories listed
-
Shifted Window Fourier Transform And Retention For Image Captioning25 Aug 2024 0 repositories listed
-
The Brittleness of AI-Generated Image Watermarking Techniques: Examining Their Robustness Against Visual Paraphrasing Attacks19 Aug 2024 0 repositories listed
-
PathInsight: Instruction Tuning of Multimodal Datasets and Models for Intelligence Assisted Diagnosis in Histopathology13 Aug 2024 0 repositories listed
-
Surveying the Landscape of Image Captioning Evaluation: A Comprehensive Taxonomy and Novel Ensemble Method9 Aug 2024 0 repositories listed
-
Enhancing Journalism with AI: A Study of Contextualized Image Captioning for News Articles using LLMs and LMMs8 Aug 2024 0 repositories listed
-
One Framework to Rule Them All: Unifying Multimodal Tasks with LLM Neural-Tuning6 Aug 2024 0 repositories listed
-
Modelling Visual Semantics via Image Captioning to extract Enhanced Multi-Level Cross-Modal Semantic Incongruity Representation with Attention for Multimodal Sarcasm Detection5 Aug 2024 0 repositories listed
-
Dataset Scale and Societal Consistency Mediate Facial Impression Bias in Vision-Language AI4 Aug 2024 0 repositories listed
-
A Novel Evaluation Framework for Image2Text Generation3 Aug 2024 0 repositories listed
-
The Phantom Menace: Unmasking Privacy Leakages in Vision-Language Models2 Aug 2024 0 repositories listed
-
AI Safety in Practice: Enhancing Adversarial Robustness in Multimodal Image Captioning30 Jul 2024 0 repositories listed
-
VolDoGer: LLM-assisted Datasets for Domain Generalization in Vision-Language Tasks29 Jul 2024 0 repositories listed
-
HICEScore: A Hierarchical Metric for Image Captioning Evaluation26 Jul 2024 0 repositories listed
-
Imperfect Vision Encoders: Efficient and Robust Tuning for Vision-Language Models23 Jul 2024 0 repositories listed
-
VideoGameBunny: Towards vision assistants for video games21 Jul 2024 0 repositories listed
-
Downstream-Pretext Domain Knowledge Traceback for Active Learning20 Jul 2024 0 repositories listed
-
EVLM: An Efficient Vision-Language Model for Visual Understanding19 Jul 2024 0 repositories listed
-
LookupViT: Compressing visual information to a limited number of tokens17 Jul 2024 0 repositories listed
-
Resampled Datasets Are Not Enough: Mitigating Societal Bias Beyond Single Attributes4 Jul 2024 0 repositories listed
-
BACON: Improving Clarity of Image Captions via Bag-of-Concept Graphs3 Jul 2024 0 repositories listed
-
Certainly Uncertain: A Benchmark and Metric for Multimodal Epistemic and Aleatoric Awareness2 Jul 2024 0 repositories listed
-
Assistive Image Annotation Systems with Deep Learning and Natural Language Capabilities: A Review28 Jun 2024 0 repositories listed
-
Auto Cherry-Picker: Learning from High-quality Generative Data Driven by Language28 Jun 2024 0 repositories listed
-
Explainable Image Captioning using CNN- CNN architecture and Hierarchical Attention28 Jun 2024 0 repositories listed
-
RAVEN: Multitask Retrieval Augmented Vision-Language Learning27 Jun 2024 0 repositories listed
-
Towards Temporal Change Explanations from Bi-Temporal Satellite Images27 Jun 2024 0 repositories listed
-
MUMU: Bootstrapping Multimodal Image Generation from Text-to-Image Data26 Jun 2024 0 repositories listed
-
Enhancing Scientific Figure Captioning Through Cross-modal Learning24 Jun 2024 0 repositories listed
-
Does Object Grounding Really Reduce Hallucination of Large Vision-Language Models?20 Jun 2024 0 repositories listed
-
Reinforcing Pre-trained Models Using Counterfactual Images19 Jun 2024 0 repositories listed
-
Do More Details Always Introduce More Hallucinations in LVLM-based Image Captioning?18 Jun 2024 0 repositories listed
-
LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning17 Jun 2024 0 repositories listed
-
From Pixels to Prose: A Large Dataset of Dense Image Captions14 Jun 2024 0 repositories listed
-
OSPC: Detecting Harmful Memes with Large Language Model as a Catalyst14 Jun 2024 0 repositories listed
-
DS@BioMed at ImageCLEFmedical Caption 2024: Enhanced Attention Mechanisms in Medical Caption Generation through Concept Detection Integration1 Jun 2024 0 repositories listed
-
Image captioning in different languages31 May 2024 0 repositories listed
-
30 May 2024 0 repositories listed
-
MetaToken: Detecting Hallucination in Image Descriptions by Meta Classification29 May 2024 0 repositories listed
-
Multi-Modal Generative Embedding Model29 May 2024 0 repositories listed
-
Text-only Synthesis for Image Captioning28 May 2024 0 repositories listed
-
How Culturally Aware are Vision-Language Models?24 May 2024 0 repositories listed
-
LG-VQ: Language-Guided Codebook Learning23 May 2024 0 repositories listed
-
CrossCheckGPT: Universal Hallucination Ranking for Multimodal Foundation Models22 May 2024 0 repositories listed
-
Towards Retrieval-Augmented Architectures for Image Captioning21 May 2024 0 repositories listed
-
Contextual Emotion Recognition using Large Vision Language Models14 May 2024 0 repositories listed
-
Using Machine Translation to Augment Multilingual Classification9 May 2024 0 repositories listed
-
A Toolchain for Comprehensive Audio/Video Analysis Using Deep Learning Based Multimodal Approach (A use case of riot or violent context detection)2 May 2024 0 repositories listed
-
Beyond Human Vision: The Role of Large Vision Language Models in Microscope Image Analysis1 May 2024 0 repositories listed