Browse State-of-the-Art › Image Captioning › Papers, page 12
Image Captioning
Papers archive 2025-07-28
archive papers tagged: 1,878 · with a code link: 774 · where Syntology ran a sample: 243 (201 with a run with no instrument failure, 42 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (243 of 1,878 tagged: 201 with a run with no instrument failure, 42 where every run was a failure of Syntology's instrument)
Page 12 of 19: papers 1,101 to 1,200 of 1,878, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Module-wise Adaptive Distillation for Multimodality Foundation Models6 Oct 2023 0 repositories listed
-
5 Oct 2023 0 repositories listed
-
On the Performance of Multimodal Language Models4 Oct 2023 0 repositories listed
-
Self-Supervised Open-Ended Classification with Small Visual Language Models30 Sep 2023 0 repositories listed
-
Targeted Image Data Augmentation Increases Basic Skills Captioning Robustness27 Sep 2023 0 repositories listed
-
Aligning Large Multimodal Models with Factually Augmented RLHF25 Sep 2023 0 repositories listed
-
FaceGemma: Enhancing Image Captioning with Facial Attributes for Portrait Images24 Sep 2023 0 repositories listed
-
Contextual Emotion Estimation from Image Captions22 Sep 2023 0 repositories listed
-
Auto-ACD: A Large-scale Dataset for Audio-Language Representation Learning20 Sep 2023 0 repositories listed
-
Prefix-diffusion: A Lightweight Diffusion Model for Diverse Image Captioning10 Sep 2023 0 repositories listed
-
NICE: CVPR 2023 Challenge on Zero-shot Image Captioning5 Sep 2023 0 repositories listed
-
Physically Grounded Vision-Language Models for Robotic Manipulation5 Sep 2023 0 repositories listed
-
RSDiff: Remote Sensing Image Generation from Text Using Diffusion Model3 Sep 2023 0 repositories listed
-
Can Prompt Learning Benefit Radiology Report Generation?30 Aug 2023 0 repositories listed
-
Finding-Aware Anatomical Tokens for Chest X-Ray Automated Reporting30 Aug 2023 0 repositories listed
-
Towards Real Time Egocentric Segment Captioning for The Blind and Visually Impaired in RGB-D Theatre Images26 Aug 2023 0 repositories listed
-
DLIP: Distilling Language-Image Pre-training24 Aug 2023 0 repositories listed
-
Explore and Tell: Embodied Visual Captioning in 3D Environments21 Aug 2023 0 repositories listed
-
Generic Attention-model Explainability by Weighted Relevance Accumulation20 Aug 2023 0 repositories listed
-
UniBrain: Unify Image Reconstruction and Captioning All in One Diffusion Model from Human Brain Activity14 Aug 2023 0 repositories listed
-
IIHT: Medical Report Generation with Image-to-Indicator Hierarchical Transformer10 Aug 2023 0 repositories listed
-
Informative Scene Graph Generation via Debiasing10 Aug 2023 0 repositories listed
-
Asynchronous Evolution of Deep Neural Network Architectures8 Aug 2023 0 repositories listed
-
Building Safe and Reliable AI systems for Safety Critical Tasks with Vision-Language Processing6 Aug 2023 0 repositories listed
-
A Comprehensive Analysis of Real-World Image Captioning and Scene Identification5 Aug 2023 0 repositories listed
-
Improving Generalization of Image Captioning with Unsupervised Prompt Learning5 Aug 2023 0 repositories listed
-
Multimodal Neurons in Pretrained Text-Only Transformers3 Aug 2023 0 repositories listed
-
Guiding Image Captioning Models Toward More Specific Captions31 Jul 2023 0 repositories listed
-
Visual Captioning at Will: Describing Images and Videos Guided by a Few Stylized Sentences31 Jul 2023 0 repositories listed
-
Causal reasoning in typical computer vision tasks26 Jul 2023 0 repositories listed
-
Enhancing image captioning with depth information using a Transformer-based framework24 Jul 2023 0 repositories listed
-
OxfordTVG-HIC: Can Machine Make Humorous Captions from Images?21 Jul 2023 0 repositories listed
-
Embedded Heterogeneous Attention Transformer for Cross-lingual Image Captioning19 Jul 2023 0 repositories listed
-
Improving Multimodal Datasets with Image Captioning19 Jul 2023 0 repositories listed
-
AIC-AB NET: A Neural Network for Image Captioning with Spatial Attention and Text Attributes14 Jul 2023 0 repositories listed
-
Reading Radiology Imaging Like The Radiologist12 Jul 2023 0 repositories listed
-
Multimodal Prompt Learning for Product Title Generation with Extremely Limited Labels5 Jul 2023 0 repositories listed
-
More for Less: Compact Convolutional Transformers Enable Robust Medical Image Classification with Limited Data1 Jul 2023 0 repositories listed
-
Benchmarking Zero-Shot Recognition with Vision-Language Models: Challenges on Granularity and Specificity28 Jun 2023 0 repositories listed
-
Self-Supervised Image Captioning with CLIP26 Jun 2023 0 repositories listed
-
Improving Reference-based Distinctive Image Captioning with Contrastive Rewards25 Jun 2023 0 repositories listed
-
Improving Image Captioning Descriptiveness by Ranking and LLM-based Fusion20 Jun 2023 0 repositories listed
-
Generation of Radiology Findings in Chest X-Ray by Leveraging Collaborative Knowledge18 Jun 2023 0 repositories listed
-
A Survey of Vision-Language Pre-training from the Lens of Multimodal Machine Translation12 Jun 2023 0 repositories listed
-
Putting Humans in the Image Captioning Loop6 Jun 2023 0 repositories listed
-
Towards Adaptable and Interactive Image Captioning with Data Augmentation and Episodic Memory6 Jun 2023 0 repositories listed
-
Cheap-fake Detection with LLM using Prompt Engineering5 Jun 2023 0 repositories listed
-
"Let's not Quote out of Context": Unified Vision-Language Pretraining for Context Assisted Image Captioning1 Jun 2023 0 repositories listed
-
Image Captioning with Multi-Context Synthetic Data29 May 2023 0 repositories listed
-
Green Runner: A tool for efficient model selection from model repositories26 May 2023 0 repositories listed
-
Mindstorms in Natural Language-Based Societies of Mind26 May 2023 0 repositories listed
-
HAAV: Hierarchical Aggregation of Augmented Views for Image Captioning25 May 2023 0 repositories listed
-
EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought24 May 2023 0 repositories listed
-
Exploring Affordance and Situated Meaning in Image Captions: A Multimodal Analysis24 May 2023 0 repositories listed
-
Alt-Text with Context: Improving Accessibility for Images on Twitter24 May 2023 0 repositories listed
-
Text-based Person Search without Parallel Image-Text Data22 May 2023 0 repositories listed
-
Cross2StrA: Unpaired Cross-lingual Image Captioning with Cross-lingual Cross-modal Structure-pivoted Alignment20 May 2023 0 repositories listed
-
DiffCap: Exploring Continuous Diffusion on Image Captioning20 May 2023 0 repositories listed
-
Semantic Composition in Visually Grounded Language Models15 May 2023 0 repositories listed
-
11 May 2023 0 repositories listed
-
Towards L-System Captioning for Tree Reconstruction10 May 2023 0 repositories listed
-
Exploiting Pseudo Image Captions for Multimodal Summarization9 May 2023 0 repositories listed
-
UIT-OpenViIC: A Novel Benchmark for Evaluating Image Captioning in Vietnamese7 May 2023 0 repositories listed
-
Image Captioners Sometimes Tell More Than Images They See4 May 2023 0 repositories listed
-
Fairness in AI Systems: Mitigating gender bias from language-vision models3 May 2023 0 repositories listed
-
Making the Most of What You Have: Adapting Pre-trained Visual Language Models in the Low-data Regime3 May 2023 0 repositories listed
-
Quality-agnostic Image Captioning to Safely Assist People with Vision Impairment28 Apr 2023 0 repositories listed
-
27 Apr 2023 0 repositories listed
-
A-CAP: Anticipation Captioning with Commonsense Knowledge13 Apr 2023 0 repositories listed
-
Advancing Medical Imaging with Language Models: A Journey from N-grams to ChatGPT11 Apr 2023 0 repositories listed
-
Boosting Cross-task Transferability of Adversarial Patches with Visual Relations11 Apr 2023 0 repositories listed
-
ImageCaptioner²: Image Captioner for Image Captioning Bias Amplification Assessment10 Apr 2023 0 repositories listed
-
Towards Self-Explainability of Deep Neural Networks with Heatmap Captioning and Large-Language Models5 Apr 2023 0 repositories listed
-
Scalable and Accurate Self-supervised Multimodal Representation Learning without Aligned Video and Text Data4 Apr 2023 0 repositories listed
-
Mask-free OVIS: Open-Vocabulary Instance Segmentation without Manual Mask Annotations29 Mar 2023 0 repositories listed
-
Variational Distribution Learning for Unsupervised Text-to-Image Generation28 Mar 2023 0 repositories listed
-
Open-Vocabulary Object Detection using Pseudo Caption Labels23 Mar 2023 0 repositories listed
-
Sketch2Saliency: Learning to Detect Salient Objects from Human Drawings20 Mar 2023 0 repositories listed
-
Multi-modal reward for visual relationships-based image captioning19 Mar 2023 0 repositories listed
-
Visual Information Matters for ASR Error Correction16 Mar 2023 0 repositories listed
-
PR-MCS: Perturbation Robust Metric for MultiLingual Image Captioning15 Mar 2023 0 repositories listed
-
Breaking Common Sense: WHOOPS! A Vision-and-Language Benchmark of Synthetic and Compositional Images13 Mar 2023 0 repositories listed
-
Learning Combinatorial Prompts for Universal Controllable Image Captioning11 Mar 2023 0 repositories listed
-
Interpretable Visual Question Answering Referring to Outside Knowledge8 Mar 2023 0 repositories listed
-
Graph Neural Networks in Vision-Language Image Understanding: A Survey7 Mar 2023 0 repositories listed
-
Comparative study of Transformer and LSTM Network with attention mechanism on Image Captioning5 Mar 2023 0 repositories listed
-
See Your Heart: Psychological states Interpretation through Visual Creations11 Feb 2023 0 repositories listed
-
Nemesis: Neural Mean Teacher Learning-Based Emotion-Centric Speaker9 Feb 2023 0 repositories listed
-
Re-ViLM: Retrieval-Augmented Visual Language Model for Zero and Few-Shot Image Captioning9 Feb 2023 0 repositories listed
-
Stacked Cross-modal Feature Consolidation Attention Networks for Image Captioning8 Feb 2023 0 repositories listed
-
KENGIC: KEyword-driven and N-Gram Graph based Image Captioning7 Feb 2023 0 repositories listed
-
DEVICE: DEpth and VIsual ConcEpts Aware Transformer for TextCaps3 Feb 2023 0 repositories listed
-
PromptMix: Text-to-image diffusion models enhance the performance of lightweight networks30 Jan 2023 0 repositories listed
-
Exploring External Knowledge for Accurate modeling of Visual and Language Problems27 Jan 2023 0 repositories listed
-
Semi-Supervised Image Captioning by Adversarially Propagating Labeled Data26 Jan 2023 0 repositories listed
-
Style-Aware Contrastive Learning for Multi-Style Image Captioning26 Jan 2023 0 repositories listed
-
Summarize the Past to Predict the Future: Natural Language Descriptions of Context Boost Multimodal Object Interaction Anticipation22 Jan 2023 0 repositories listed
-
Towards Models that Can See and Read18 Jan 2023 0 repositories listed
-
Embodied Agents for Efficient Exploration and Smart Scene Description17 Jan 2023 0 repositories listed
-
An Image captioning algorithm based on the Hybrid Deep Learning Technique (CNN+GRU)6 Jan 2023 0 repositories listed