Browse State-of-the-Art › Image Captioning › Papers, page 11
Image Captioning
Papers archive 2025-07-28
archive papers tagged: 1,878 · with a code link: 774 · where Syntology ran a sample: 243 (201 with a run with no instrument failure, 42 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (243 of 1,878 tagged: 201 with a run with no instrument failure, 42 where every run was a failure of Syntology's instrument)
Page 11 of 19: papers 1,001 to 1,100 of 1,878, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
What Makes for Good Image Captions?1 May 2024 0 repositories listed
-
Compressed Image Captioning using CNN-based Encoder-Decoder Framework28 Apr 2024 0 repositories listed
-
Semi-supervised Text-based Person Search28 Apr 2024 0 repositories listed
-
Learning text-to-video retrieval from image captioning26 Apr 2024 0 repositories listed
-
MM-PhyRLHF: Reinforcement Learning Framework for Multimodal Physics Question-Answering19 Apr 2024 0 repositories listed
-
The Solution for the CVPR2024 NICE Image Captioning Challenge19 Apr 2024 0 repositories listed
-
On Speculative Decoding for Multimodal Large Language Models13 Apr 2024 0 repositories listed
-
Panoptic Perception: A Novel Task and Fine-grained Dataset for Universal Remote Sensing Image Interpretation6 Apr 2024 0 repositories listed
-
Would Deep Generative Models Amplify Bias in Future Models?4 Apr 2024 0 repositories listed
-
Jump Self-attention: Capturing High-order Statistics in Transformers3 Apr 2024 0 repositories listed
-
VLRM: Vision-Language Models act as Reward Models for Image Captioning2 Apr 2024 0 repositories listed
-
LLaMA-Excitor: General Instruction Tuning via Indirect Feature Interaction1 Apr 2024 0 repositories listed
-
A Review of Multi-Modal Large Language and Vision Models28 Mar 2024 0 repositories listed
-
Text Data-Centric Image Captioning with Interactive Prompts28 Mar 2024 0 repositories listed
-
A Survey on Large Language Models from Concept to Implementation27 Mar 2024 0 repositories listed
-
Automated Report Generation for Lung Cytological Images Using a CNN Vision Classifier and Multiple-Transformer Text Decoders: Preliminary Study26 Mar 2024 0 repositories listed
-
Semi-Supervised Image Captioning Considering Wasserstein Graph Matching26 Mar 2024 0 repositories listed
-
The Solution for the ICCV 2023 1st Scientific Figure Captioning Challenge26 Mar 2024 0 repositories listed
-
Visual Hallucination: Definition, Quantification, and Prescriptive Remediations26 Mar 2024 0 repositories listed
-
Image Captioning in news report scenario24 Mar 2024 0 repositories listed
-
A Multimodal Approach for Cross-Domain Image Retrieval22 Mar 2024 0 repositories listed
-
MyVLM: Personalizing VLMs for User-Specific Queries21 Mar 2024 0 repositories listed
-
Improved Baselines for Data-efficient Perceptual Augmentation of LLMs20 Mar 2024 0 repositories listed
-
Inserting Faces inside Captions: Image Captioning with Attention Guided Merging20 Mar 2024 0 repositories listed
-
As Firm As Their Foundations: Can open-sourced foundation models be used to create adversarial examples for downstream tasks?19 Mar 2024 0 repositories listed
-
Entity6K: A Large Open-Domain Evaluation Dataset for Real-World Entity Recognition19 Mar 2024 0 repositories listed
-
Towards Multimodal In-Context Learning for Vision & Language Models19 Mar 2024 0 repositories listed
-
TARN-VIST: Topic Aware Reinforcement Network for Visual Storytelling18 Mar 2024 0 repositories listed
-
Few-Shot VQA with Frozen LLMs: A Tale of Two Approaches17 Mar 2024 0 repositories listed
-
A Comprehensive Survey of 3D Dense Captioning: Localizing and Describing Objects in 3D Scenes12 Mar 2024 0 repositories listed
-
Synth²: Boosting Visual-Language Models with Synthetic Captions and Image Embeddings12 Mar 2024 0 repositories listed
-
Transformer based Multitask Learning for Image Captioning and Object Detection10 Mar 2024 0 repositories listed
-
The Case for Evaluating Multimodal Translation Models on Text Datasets5 Mar 2024 0 repositories listed
-
EAMA : Entity-Aware Multimodal Alignment Based Approach for News Image Captioning29 Feb 2024 0 repositories listed
-
Vision Language Model-based Caption Evaluation Method Leveraging Visual Context Extraction28 Feb 2024 0 repositories listed
-
ArcSin: Adaptive ranged cosine Similarity injected noise for Language-Driven Visual Tasks27 Feb 2024 0 repositories listed
-
23 Feb 2024 0 repositories listed Syntology 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Exploring the Frontier of Vision-Language Models: A Survey of Current Methodologies and Future Directions20 Feb 2024 0 repositories listed
-
19 Feb 2024 0 repositories listed
-
Model Tailor: Mitigating Catastrophic Forgetting in Multi-modal Large Language Models19 Feb 2024 0 repositories listed
-
Cobra Effect in Reference-Free Image Captioning Metrics18 Feb 2024 0 repositories listed
-
Captions Are Worth a Thousand Words: Enhancing Product Retrieval with Pretrained Image-to-Text Models13 Feb 2024 0 repositories listed
-
Learning How To Ask: Cycle-Consistency Refines Prompts in Multimodal Foundation Models13 Feb 2024 0 repositories listed
-
Multimodal Learned Sparse Retrieval for Image Suggestion12 Feb 2024 0 repositories listed
-
Consistency Model is an Effective Posterior Sample Approximation for Diffusion Inverse Solvers9 Feb 2024 0 repositories listed
-
Large Language Models for Captioning and Retrieving Remote Sensing Images9 Feb 2024 0 repositories listed
-
CIC: A Framework for Culturally-Aware Image Captioning8 Feb 2024 0 repositories listed
-
Exploring Visual Culture Awareness in GPT-4V: A Comprehensive Probing8 Feb 2024 0 repositories listed
-
Image captioning for Brazilian Portuguese using GRIT model7 Feb 2024 0 repositories listed
-
Text or Image? What is More Important in Cross-Domain Generalization Capabilities of Hate Meme Detection Models?7 Feb 2024 0 repositories listed
-
PICS: Pipeline for Image Captioning and Search1 Feb 2024 0 repositories listed
-
SCO-VIST: Social Interaction Commonsense Knowledge-based Visual Storytelling1 Feb 2024 0 repositories listed
-
17 Jan 2024 0 repositories listed
-
KTVIC: A Vietnamese Image Captioning Dataset on the Life Domain16 Jan 2024 0 repositories listed
-
Jewelry Recognition via Encoder-Decoder Models15 Jan 2024 0 repositories listed
-
What Else Would I Like? A User Simulator using Alternatives for Improved Evaluation of Fashion Conversational Recommendation Systems11 Jan 2024 0 repositories listed
-
Let's Go Shopping (LGS) -- Web-Scale Image-Text Dataset for Visual Concept Understanding9 Jan 2024 0 repositories listed
-
MAMI: Multi-Attentional Mutual-Information for Long Sequence Neuron Captioning5 Jan 2024 0 repositories listed
-
Object-oriented backdoor attack against image captioning5 Jan 2024 0 repositories listed
-
4 Jan 2024 0 repositories listed
-
Social Media Ready Caption Generation for Brands3 Jan 2024 0 repositories listed
-
Cycle-Consistency Learning for Captioning and Grounding23 Dec 2023 0 repositories listed
-
LLM4VG: Large Language Models Evaluation for Video Grounding21 Dec 2023 0 repositories listed
-
Dietary Assessment with Multimodal ChatGPT: A Systematic Analysis14 Dec 2023 0 repositories listed
-
Improving Cross-modal Alignment with Synthetic Pairs for Text-only Image Captioning14 Dec 2023 0 repositories listed
-
Synocene, Beyond the Anthropocene: De-Anthropocentralising Human-Nature-AI Interaction13 Dec 2023 0 repositories listed
-
Filter & Align: Leveraging Human Knowledge to Curate Image-Text Data11 Dec 2023 0 repositories listed
-
8 Dec 2023 0 repositories listed
-
User-Aware Prefix-Tuning is a Good Learner for Personalized Image Captioning8 Dec 2023 0 repositories listed
-
On the Robustness of Large Multimodal Models Against Image Adversarial Attacks6 Dec 2023 0 repositories listed
-
Towards More Unified In-context Visual Understanding5 Dec 2023 0 repositories listed
-
CLAMP: Contrastive LAnguage Model Prompt-tuning4 Dec 2023 0 repositories listed
-
Enhancing Image Captioning with Neural Models1 Dec 2023 0 repositories listed
-
1 Dec 2023 0 repositories listed
-
InstructSeq: Unifying Vision Tasks with Instruction-conditioned Multi-modal Sequence Generation30 Nov 2023 0 repositories listed
-
A natural language processing-based approach: mapping human perception by understanding deep semantic features in street view images29 Nov 2023 0 repositories listed
-
EVCap: Retrieval-Augmented Image Captioning with External Visual-Name Memory for Open-World Comprehension27 Nov 2023 0 repositories listed
-
DECap: Towards Generalized Explicit Caption Editing via Diffusion Mechanism25 Nov 2023 0 repositories listed
-
Violet: A Vision-Language Model for Arabic Image Captioning with Gemini Decoder15 Nov 2023 0 repositories listed
-
Improving Image Captioning via Predicting Structured Concepts14 Nov 2023 0 repositories listed
-
Holistic Evaluation of GPT-4V for Biomedical Imaging10 Nov 2023 0 repositories listed
-
How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model10 Nov 2023 0 repositories listed
-
Visual Analytics for Efficient Image Exploration and User-Guided Image Captioning2 Nov 2023 0 repositories listed
-
Enhanced Knowledge Injection for Radiology Report Generation1 Nov 2023 0 repositories listed
-
What a Whole Slide Image Can Tell? Subtype-guided Masked Transformer for Pathological Image Captioning31 Oct 2023 0 repositories listed
-
Improving Medical Visual Representations via Radiology Report Generation30 Oct 2023 0 repositories listed
-
CropCap: Embedding Visual Cross-Partition Dependency for Image Captioning27 Oct 2023 0 repositories listed
-
Impressions: Understanding Visual Semiotics and Aesthetic Impact27 Oct 2023 0 repositories listed
-
A Picture is Worth a Thousand Words: Principled Recaptioning Improves Image Generation25 Oct 2023 0 repositories listed
-
Semantic and Expressive Variation in Image Captions Across Languages22 Oct 2023 0 repositories listed
-
CLAIR: Evaluating Image Captions with Large Language Models19 Oct 2023 0 repositories listed
-
Lost in Translation: When GPT-4V(ision) Can't See Eye to Eye with Text. A Vision-Language-Consistency Analysis of VLLMs and Beyond19 Oct 2023 0 repositories listed
-
Towards Automatic Satellite Images Captions Generation Using Large Language Models17 Oct 2023 0 repositories listed
-
Ziya-Visual: Bilingual Large Vision-Language Model via Multi-Task Instruction Tuning12 Oct 2023 0 repositories listed
-
A Comparative Study of Pre-trained CNNs and GRU-Based Attention for Image Caption Generation11 Oct 2023 0 repositories listed
-
Improving mitosis detection on histopathology images using large vision-language models11 Oct 2023 0 repositories listed
-
LangNav: Language as a Perceptual Representation for Navigation11 Oct 2023 0 repositories listed
-
The Solution for the CVPR2023 NICE Image Captioning Challenge10 Oct 2023 0 repositories listed
-
ViCor: Bridging Visual Understanding and Commonsense Reasoning with Large Language Models9 Oct 2023 0 repositories listed
-
Lightweight In-Context Tuning for Multimodal Unified Models8 Oct 2023 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.