Browse State-of-the-Art › Image Captioning › Papers, page 13
Image Captioning
Papers archive 2025-07-28
archive papers tagged: 1,878 · with a code link: 774 · where Syntology ran a sample: 243 (201 with a run with no instrument failure, 42 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (243 of 1,878 tagged: 201 with a run with no instrument failure, 42 where every run was a failure of Syntology's instrument)
Page 13 of 19: papers 1,201 to 1,300 of 1,878, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
An Empirical Investigation into the Use of Image Captioning for Automated Software Documentation3 Jan 2023 0 repositories listed
-
Crossing the Gap: Domain Generalization for Image Captioning1 Jan 2023 0 repositories listed
-
Image as a Foreign Language: BEiT Pretraining for Vision and Vision-Language Tasks1 Jan 2023 0 repositories listed
-
PromptCap: Prompt-Guided Image Captioning for VQA with GPT-31 Jan 2023 0 repositories listed
-
Do DALL-E and Flamingo Understand Each Other?23 Dec 2022 0 repositories listed
-
Efficient Image Captioning for Edge Devices18 Dec 2022 0 repositories listed
-
Cross-Modal Similarity-Based Curriculum Learning for Image Captioning14 Dec 2022 0 repositories listed
-
NLIP: Noise-robust Language-Image Pre-training14 Dec 2022 0 repositories listed
-
Cap2Aug: Caption guided Image to Image data Augmentation11 Dec 2022 0 repositories listed
-
7 Dec 2022 0 repositories listed
-
Dataset vs Reality: Understanding Model Performance from the Perspective of Information Need6 Dec 2022 0 repositories listed
-
Controllable Image Captioning via Prompting4 Dec 2022 0 repositories listed
-
Focus! Relevant and Sufficient Context Selection for News Image Captioning1 Dec 2022 0 repositories listed
-
Weakly Supervised Annotations for Multi-modal Greeting Cards Dataset1 Dec 2022 0 repositories listed
-
Uncertainty-Aware Image Captioning30 Nov 2022 0 repositories listed
-
Predictive linguistic cues for fake news: a societal artificial intelligence problem26 Nov 2022 0 repositories listed
-
Can Machines Imitate Humans? Integrative Turing Tests for Vision and Language Demonstrate a Narrowing Gap23 Nov 2022 0 repositories listed
-
22 Nov 2022 0 repositories listed
-
A survey on knowledge-enhanced multimodal learning19 Nov 2022 0 repositories listed
-
18 Nov 2022 0 repositories listed
-
Feedback is Needed for Retakes: An Explainable Poor Image Notification Framework for the Visually Impaired17 Nov 2022 0 repositories listed
-
Zero-shot Image Captioning by Anchor-augmented Vision-Language Space Alignment14 Nov 2022 0 repositories listed
-
Investigations in Audio Captioning: Addressing Vocabulary Imbalance and Evaluating Suitability of Language-Centric Performance Metrics12 Nov 2022 0 repositories listed
-
VieCap4H-VLSP 2021: ObjectAoA-Enhancing performance of Object Relation Transformer with Attention on Attention for Vietnamese image captioning10 Nov 2022 0 repositories listed
-
Understanding Cross-modal Interactions in V&L Models that Generate Scene Descriptions9 Nov 2022 0 repositories listed
-
Image Caption Generation for Low-Resource Assamese Language1 Nov 2022 0 repositories listed
-
DiMBERT: Learning Vision-Language Grounded Representations with Disentangled Multimodal-Attention28 Oct 2022 0 repositories listed
-
Bloom Library: Multimodal Datasets in 300+ Languages for a Variety of Downstream Tasks26 Oct 2022 0 repositories listed
-
FaD-VLP: Fashion Vision-and-Language Pre-training towards Unified Retrieval and Captioning26 Oct 2022 0 repositories listed
-
Image-Text Retrieval with Binary and Continuous Label Supervision20 Oct 2022 0 repositories listed
-
Prophet Attention: Predicting Attention with Future Attention for Image Captioning19 Oct 2022 0 repositories listed
-
Aligning MAGMA by Few-Shot Learning and Finetuning18 Oct 2022 0 repositories listed
-
Probing Cross-modal Semantics Alignment Capability from the Textual Perspective18 Oct 2022 0 repositories listed
-
Generating image captions with external encyclopedic knowledge10 Oct 2022 0 repositories listed
-
Text-to-Audio Grounding Based Novel Metric for Evaluating Audio Caption Similarity3 Oct 2022 0 repositories listed
-
DeltaNet: Conditional Medical Report Generation for COVID-19 Diagnosis1 Oct 2022 0 repositories listed
-
JPG - Jointly Learn to Align: Automated Disease Prediction and Radiology Report Generation1 Oct 2022 0 repositories listed
-
Multi-view and Cross-view Brain Decoding1 Oct 2022 0 repositories listed
-
On the Effects of Video Grounding on Language Models1 Oct 2022 0 repositories listed
-
Medical Image Captioning via Generative Pretrained Transformers28 Sep 2022 0 repositories listed
-
DRAMA: Joint Risk Localization and Captioning in Driving22 Sep 2022 0 repositories listed
-
Show, Interpret and Tell: Entity-aware Contextualised Image Captioning in Wikipedia21 Sep 2022 0 repositories listed
-
Toward 3D Spatial Reasoning for Human-like Text-based Visual Question Answering21 Sep 2022 0 repositories listed
-
15 Sep 2022 0 repositories listed
-
PreSTU: Pre-Training for Scene-Text Understanding12 Sep 2022 0 repositories listed
-
Every picture tells a story: Image-grounded controllable stylistic story generation4 Sep 2022 0 repositories listed
-
vieCap4H-VLSP 2021: Vietnamese Image Captioning for Healthcare Domain using Swin Transformer and Attention-based LSTM3 Sep 2022 0 repositories listed
-
A Medical Semantic-Assisted Transformer for Radiographic Report Generation22 Aug 2022 0 repositories listed
-
Aesthetic Attributes Assessment of Images with AMANv2 and DPC-CaptionsV29 Aug 2022 0 repositories listed
-
Distinctive Image Captioning via CLIP Guided Group Optimization8 Aug 2022 0 repositories listed
-
RadTex: Learning Efficient Radiograph Representations from Text Reports5 Aug 2022 0 repositories listed
-
Neuro-Symbolic Learning: Principles and Applications in Ophthalmology31 Jul 2022 0 repositories listed
-
Retrieval-Augmented Transformer for Image Captioning26 Jul 2022 0 repositories listed
-
A Baseline for Detecting Out-of-Distribution Examples in Image Captioning12 Jul 2022 0 repositories listed
-
Adaptive Fine-Grained Predicates Learning for Scene Graph Generation11 Jul 2022 0 repositories listed
-
Exploring the sequence length bottleneck in the Transformer for Image Captioning7 Jul 2022 0 repositories listed
-
Predicting Word Learning in Children from the Performance of Computer Vision Systems7 Jul 2022 0 repositories listed
-
Are metrics measuring what they should? An evaluation of image captioning task metrics4 Jul 2022 0 repositories listed
-
American == White in Multimodal Language-and-Image AI1 Jul 2022 0 repositories listed
-
Competence-based Multimodal Curriculum Learning for Medical Report Generation24 Jun 2022 0 repositories listed
-
DALL-E for Detection: Language-driven Compositional Image Synthesis for Object Detection20 Jun 2022 0 repositories listed
-
19 Jun 2022 0 repositories listed
-
A Self-Guided Framework for Radiology Report Generation19 Jun 2022 0 repositories listed
-
Image Captioning based on Feature Refinement and Reflective Decoding16 Jun 2022 0 repositories listed
-
Measuring Representational Harms in Image Captioning14 Jun 2022 0 repositories listed
-
Improving Image Captioning with Control Signal of Sentence Quality7 Jun 2022 0 repositories listed
-
Intra-agent speech permits zero-shot task acquisition7 Jun 2022 0 repositories listed
-
Examining the Effects of Language-and-Vision Data Augmentation for Generation of Descriptions of Human Faces1 Jun 2022 0 repositories listed
-
Visual Transformer for Object Detection1 Jun 2022 0 repositories listed
-
Prompt-based Learning for Unpaired Image Captioning26 May 2022 0 repositories listed
-
25 May 2022 0 repositories listed
-
On Advances in Text Generation from Images Beyond Captioning: A Case Study in Self-Rationalization24 May 2022 0 repositories listed
-
Reassessing Evaluation Practices in Visual Question Answering: A Case Study on Out-of-Distribution Generalization24 May 2022 0 repositories listed
-
It Isn't Sh!tposting, It's My CAT Posting18 May 2022 0 repositories listed
-
Understanding Transfer Learning for Chest Radiograph Clinical Report Generation with Modified Transformer Architectures5 May 2022 0 repositories listed
-
Answer-Me: Multi-Task Open-Vocabulary Visual Question Answering2 May 2022 0 repositories listed
-
Combine to Describe: Evaluating Compositional Generalization in Image Captioning1 May 2022 0 repositories listed
-
Molecular Identification from AFM images using the IUPAC Nomenclature and Attribute Multimodal Recurrent Neural Networks1 May 2022 0 repositories listed
-
Controllable Image Captioning28 Apr 2022 0 repositories listed
-
CapOnImage: Context-driven Dense-Captioning on Image27 Apr 2022 0 repositories listed
-
Cross-view Brain Decoding18 Apr 2022 0 repositories listed
-
Guiding Attention using Partial-Order Relationships for Image Captioning15 Apr 2022 0 repositories listed
-
10 Apr 2022 0 repositories listed
-
8 Apr 2022 0 repositories listed
-
On Distinctive Image Captioning via Comparing and Reweighting8 Apr 2022 0 repositories listed
-
Semantic Exploration from Language Abstractions and Pretrained Representations8 Apr 2022 0 repositories listed
-
1 Apr 2022 0 repositories listed
-
WuDaoMM: A large-scale Multi-Modal Dataset for Pre-training models22 Mar 2022 0 repositories listed
-
AlignTransformer: Hierarchical Alignment of Visual Regions and Disease Tags for Medical Report Generation18 Mar 2022 0 repositories listed
-
DU-VLG: Unifying Vision-and-Language Generation via Dual Sequence-to-Sequence Pre-training17 Mar 2022 0 repositories listed
-
Contrastive Visual Semantic Pretraining Magnifies the Semantics of Natural Language Representations14 Mar 2022 0 repositories listed
-
Leveraging Visual Knowledge in Language Tasks: An Empirical Study on Intermediate Pre-training for Cross-modal Knowledge Transfer14 Mar 2022 0 repositories listed
-
Enabling Multimodal Generation on CLIP via Vision-Language Knowledge Distillation12 Mar 2022 0 repositories listed
-
Taking an Emotional Look at Video Paragraph Captioning12 Mar 2022 0 repositories listed
-
Semantic Distillation Guided Salient Object Detection8 Mar 2022 0 repositories listed
-
Unpaired Image Captioning by Image-level Weakly-Supervised Visual Concept Recognition7 Mar 2022 0 repositories listed
-
A Deep Neural Framework for Image Caption Generation Using GRU-Based Attention Mechanism3 Mar 2022 0 repositories listed
-
Interactive Machine Learning for Image Captioning28 Feb 2022 0 repositories listed
-
I-Tuning: Tuning Frozen Language Models with Image for Lightweight Image Captioning14 Feb 2022 0 repositories listed
-
Bench-Marking And Improving Arabic Automatic Image Captioning Through The Use Of Multi-Task Learning Paradigm11 Feb 2022 0 repositories listed