Browse State-of-the-Art › Image Captioning
Image Captioning
774 papers with code · 33 benchmarks · 79 datasets archive 2025-07-28
Image Captioning is the task of describing the content of an image in words. This task lies at the intersection of computer vision and natural language processing. Most image captioning systems use an encoder-decoder framework, where an input image is encoded into an intermediate representation of the information in the image, and then decoded into a descriptive text sequence. The most popular benchmarks are nocaps and COCO, and models are typically evaluated according to a BLEU or CIDER metric.
( Image credit: Reflective Decoding Network for Image Captioning, ICCV'19)
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
34 leaderboard tables shown for this task, 33 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 34 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
79 datasets whose archive record lists this task, ordered by the archive's paper count. 30 shown of 79 until expanded.
Subtasks archive 2025-07-28
8 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 774 papers with code (1,878 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
10 Feb 2015 92 repositories listed Syntology ran 15 of 84 samples · 69 unverified · 7 pointer-only (licence)Inspired by recent work in machine translation and object detection, we introduce an attention based model that automatically learns to describe the content of images.
-
17 Nov 2014 74 repositories listed Syntology ran 13 of 34 samples · 21 unverified · 6 pointer-only (licence)Experiments on several datasets show the accuracy of the model and the fluency of the language it learns solely from image descriptions.
-
25 Jul 2017 65 repositories listed Syntology ran 9 of 9 samples · 0 unverified · 6 pointer-only (licence)Top-down visual attention mechanisms have been used extensively in image captioning and visual question answering (VQA) to enable deeper image understanding through fine-grained analysis and even multiple steps of…
-
2 Dec 2016 31 repositories listed Syntology ran 8 of 13 samples · 5 unverified · 3 pointer-only (licence)In this paper we consider the problem of optimizing image captioning systems using reinforcement learning, and show that by carefully optimizing our systems using the test metrics of the MSCOCO task, significant gains…
-
13 May 2019 30 repositories listed Syntology ran 17 of 24 samples · 7 unverified · 5 pointer-only (licence)Regional dropout strategies have been proposed to enhance the performance of convolutional neural network classifiers.
-
7 Oct 2016 25 repositories listed Syntology ran 7 of 14 samples · 7 unverified · 14 pointer-only (licence)We observe that our method consistently outperforms BS and previously proposed techniques for diverse decoding from neural sequence models.
-
20 Nov 2014 24 repositories listed Syntology ran 12 of 32 samples · 20 unverified · 27 pointer-only (licence)We propose a novel paradigm for evaluating image descriptions that uses human consensus.
-
3 May 2015 21 repositories listed Syntology ran 6 of 7 samples · 1 unverified · 6 pointer-only (licence)Given an image and a natural language question about the image, the task is to provide an accurate natural language answer.
-
8 Sep 2014 21 repositories listed Syntology ran 2 of 6 samples · 4 unverified · 6 pointer-only (licence)We present a simple regularization technique for Recurrent Neural Networks (RNNs) with Long Short-Term Memory (LSTM) units.
-
21 Apr 2019 20 repositories listed Syntology ran 44 of 76 samples · 32 unverified · 33 pointer-only (licence)We propose BERTScore, an automatic evaluation metric for text generation.
-
21 Sep 2016 20 repositories listed Syntology ran 4 of 15 samples · 11 unverified · 7 pointer-only (licence)Automatically describing the content of an image is a fundamental problem in artificial intelligence that connects computer vision and natural language processing.
-
30 Jan 2023 17 repositories listed Syntology ran 4 of 8 samples · 4 unverified · 1 pointer-only (licence)The cost of vision-and-language pre-training has become increasingly prohibitive due to end-to-end training of large-scale models.
-
6 Dec 2016 17 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)The model decides whether to attend to the image and where, in order to extract meaningful information for sequential word generation.
-
19 Jun 2018 13 repositories listed Syntology ran 1 of 34 samples · 33 unverified · 10 pointer-only (licence)We compare our approach to state-of-the-art importance extraction methods using both an automatic deletion/insertion metric and a pointing metric based on human-annotated object segments.
-
27 Jul 2016 12 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)In this paper, we design a benchmark task and provide the associated datasets for recognizing face images and link them to corresponding entity keys in a knowledge base.
-
29 Jul 2016 11 repositories listedThere is considerable interest in the task of automatically generating image captions.
-
12 Dec 2016 10 repositories listed Syntology ran 1 of 29 samples · 28 unverifiedTo understand stuff and things in context we introduce COCO-Stuff, which augments all 164K images of the COCO 2017 dataset with pixel-wise annotations for 91 stuff classes.
-
28 Jan 2022 9 repositories listedFurthermore, performance improvement has been largely achieved by scaling up the dataset with noisy image-text pairs collected from the web, which is a suboptimal source of supervision.
-
9 Jun 2015 9 repositories listed Syntology ran 0 of 8 samples · 8 unverifiedRecurrent Neural Networks can be trained to produce sequences of tokens given some input, as exemplified by recent results in machine translation and image captioning.
-
27 Aug 2018 8 repositories listed Syntology ran 2 of 3 samples · 1 unverified · 2 pointer-only (licence)This paper introduces a convolutional recurrent network with attention for speech command recognition.
-
8 Aug 2022 7 repositories listed Syntology ran 13 of 20 samples · 7 unverified · 4 pointer-only (licence)The main idea behind our approach is to first represent the discrete data as binary bits, and then train a continuous diffusion model to model these bits as real numbers which we call analog bits.
-
2 Jan 2021 7 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 1 pointer-only (licence)In our experiments we feed the visual features generated by the new object detection model into a Transformer-based VL fusion model \oscar \cite{li2020oscar}, and utilize an improved approach \short\ to pre-train the VL…
-
4 May 2022 6 repositories listed Syntology ran 9 of 17 samples · 8 unverifiedWe apply a contrastive loss between unimodal image and text embeddings, in addition to a captioning loss on the multimodal decoder outputs which predicts text tokens autoregressively.
-
17 Apr 2018 6 repositories listed Syntology ran 0 of 8 samples · 8 unverifiedWe consider the task of text attribute transfer: transforming a sentence to alter a specific attribute (e.
-
27 May 2024 5 repositories listed Syntology ran 14 of 22 samples · 8 unverified · 8 pointer-only (licence)Traditional feedback learning for hallucination reduction relies on labor-intensive manual labeling or expensive proprietary models.
-
19 Aug 2019 5 repositories listed Syntology ran 4 of 14 samples · 10 unverifiedIn this paper, we propose an Attention on Attention (AoA) module, which extends the conventional attention mechanisms to determine the relevance between attention results and queries.
-
1 Dec 2023 4 repositories listedMultimodal Large Language Models (MLLMs) have recently demonstrated impressive capabilities in multimodal understanding, reasoning, and interaction.
-
1 Nov 2022 4 repositories listed Syntology ran 4 of 5 samples · 1 unverified · 1 pointer-only (licence)We consider the task of image-captioning using only the CLIP model and additional text data at training time, and no additional captioned images.
-
7 Oct 2022 4 repositories listed Syntology ran 0 of 5 samples · 5 unverifiedVisually-situated language is ubiquitous -- sources range from textbooks with diagrams to web pages with images and tables, to mobile apps with buttons and forms.
-
7 Feb 2022 4 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedIn this work, we pursue a unified paradigm for multimodal pretraining to break the scaffolds of complex task/modality-specific customization.
Syntology lines on 27 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections