Browse State-of-the-Art › Image to text
Image to text
103 papers with code · 0 benchmarks · 3 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
3 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 103 papers with code (246 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
30 Jan 2023 17 repositories listed Syntology ran 4 of 8 samples · 4 unverified · 1 pointer-only (licence)The cost of vision-and-language pre-training has become increasingly prohibitive due to end-to-end training of large-scale models.
-
7 Oct 2022 4 repositories listed Syntology ran 0 of 5 samples · 5 unverifiedVisually-situated language is ubiquitous -- sources range from textbooks with diagrams to web pages with images and tables, to mobile apps with buttons and forms.
-
16 Dec 2021 4 repositories listedWe propose a cross-modal attention distillation framework to train a dual-encoder model for vision-language understanding tasks, such as visual reasoning and visual question answering.
-
1 Dec 2014 4 repositories listedConvolutional neural network (CNN) is a neural network that can make use of the internal structure of data such as the 2D structure of image data.
-
29 May 2024 3 repositories listedWe present Cephalo, a series of multimodal vision large language models (V-LLMs) designed for materials science applications, integrating visual and linguistic data for enhanced understanding.
-
1 Apr 2024 3 repositories listedFor instance, the widely-used CLIPScore measures the alignment between a (generated) image and text prompt, but it fails to produce reliable scores for complex prompts involving compositions of objects, attributes, and…
-
12 Mar 2023 3 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)Inspired by the unified view, UniDiffuser learns all distributions simultaneously with a minimal modification to the original diffusion model -- perturbs data in all modalities instead of a single modality, inputs…
-
15 Nov 2022 3 repositories listedIn this work, we expand the existing single-flow diffusion pipeline into a multi-task multimodal network, dubbed Versatile Diffusion (VD), that handles multiple flows of text-to-image, image-to-text, and variations in…
-
20 Oct 2020 3 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedWe further show via a human evaluation and a qualitative analysis that our system leads to generations that are more factually complete and consistent compared to the baselines.
-
11 Apr 2025 2 repositories listedBased on EvalMi-50K, we propose LMM4LMM, an LMM-based metric for evaluating large multimodal T2I generation from multiple dimensions including perception, text-image correspondence, and task-specific accuracy.
-
18 Feb 2025 2 repositories listed Syntology ran 2 of 16 samples · 14 unverifiedWe present Magma, a foundation model that serves multimodal AI agentic tasks in both the digital and physical worlds.
-
22 Dec 2024 2 repositories listedThis survey examines the state of the art in text summarization models, with a specific focus on the abstractive summarization approach.
-
28 Oct 2024 2 repositories listedTo address this limitation, this paper proposes a training-free method called Semantic Editing Increment for ZS-CIR (SEIZE) to retrieve the target image based on the query image and text without training.
-
11 Jul 2024 2 repositories listedTo conduct ZS-CIR, the prevailing methods employ pre-trained image-to-text models to transform the query image and text into a single text, which is then projected into the common feature space by CLIP to retrieve the…
-
4 Apr 2024 2 repositories listed Syntology ran 5 of 9 samples · 4 unverified · 9 pointer-only (licence)We further attribute this phenomenon to the diffusion model's insufficient condition utilization, which is caused by its training paradigm.
-
23 Aug 2023 2 repositories listed Syntology ran 6 of 7 samples · 1 unverified · 7 pointer-only (licence)Building a competitive counterpart in other languages is highly challenging due to the low-resource nature of non-English multimodal data (i.
-
15 Aug 2023 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)In this work, we design the first vision-language dataset distillation method, building on the idea of trajectory matching.
-
11 Jul 2023 2 repositories listed Syntology ran 0 of 2 samples · 2 unverifiedWe present Emu, a Transformer-based multimodal foundation model, which can seamlessly generate images and texts in multimodal context.
-
11 Mar 2023 2 repositories listedWe present the results of extensive experiments on twelve NLG tasks, showing that, without using any labeled downstream pairs for training, ZeroNLG generates high-quality and believable outputs and significantly…
-
9 Nov 2022 2 repositories listedText-conditioned image generation models have recently achieved astonishing results in image quality and text alignment and are consequently employed in a fast-growing number of applications.
-
30 Sep 2022 2 repositories listed Syntology ran 5 of 14 samples · 9 unverifiedPrior work has shown that pretrained LMs can be taught to caption images when a vision model's parameters are optimized to encode images in the language space.
-
31 Dec 2021 2 repositories listedTo explore the landscape of large-scale pre-training for bidirectional text-image generation, we train a 10-billion parameter ERNIE-ViLG model on a large-scale dataset of 145 million (Chinese) image-text pairs which…
-
1 Dec 2021 2 repositories listed Syntology ran 13 of 14 samples · 1 unverifiedThis is an exploratory study that discovers the current image quantization (vector quantization) do not satisfy translation equivariance in the quantized space due to aliasing.
-
14 Mar 2019 2 repositories listedGenerating an image from a given text description has two goals: visual realism and semantic consistency.
-
14 Aug 2018 2 repositories listedText-to-Image translation has been an active area of research in the recent past.
-
10 Jun 2025 1 repository listedALTA achieves superior performance in vision-language matching tasks like retrieval and zero-shot classification by adapting the pretrained vision model from masked record modeling.
-
17 May 2025 1 repository listedAdditionally, we develop a specialized training strategy to align embeddings from both original and modality-completed inputs, ensuring consistency within the embedding space.
-
25 Mar 2025 1 repository listedWe propose a novel vision-language foundation model, LRSCLIP, and a multimodal dataset, LRS2M.
-
19 Mar 2025 1 repository listedResults: On the n2c2 dataset, our method achieved a new state-of-the-art criterion-level accuracy of 93\%.
-
13 Mar 2025 1 repository listed Syntology ran 6 of 9 samples · 3 unverifiedWhile conventional approaches treat the text modality as a conditioning signal that gradually guides the denoising process from Gaussian noise to the target image modality, we explore a much simpler paradigm-directly…
Syntology lines on 12 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections