Methods › Computer Vision › Vision and Language Pre-Trained Models › VL-T5

VL-T5

5 papers tagged archive 2025-07-28

Introduced by Jaemin Cho et al. in Unifying Vision-and-Language Tasks via Text Generation

archive 2025-07-28 Description, source and code snippet are the archive's method entry.

VL-T5 is a unified framework that learns different tasks in a single architecture with the same language modeling objective, i.e., multimodal conditional text generation. The model learns to generate labels in text based on the visual and textual inputs. In contrast to other existing methods, the framework unifies tasks as generating text labels conditioned on multimodal inputs. This allows the model to tackle vision-and-language tasks with unified text generation objective. The models use text prefixes to adapt to different tasks.

PaperSource

Papers archive 2025-07-28

5 shown of 5, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.

Tasks archive 2025-07-28

20 shown of 22 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.

TaskPapers
Image Captioning3
Visual Question Answering (VQA)3
Decoder2
Question Answering2
Referring Expression Comprehension2
Text Generation2
Visual Question Answering2
Conditional Text Generation1
Diversity1
Human-Object Interaction Detection1
Image Description1
Image Retrieval1
Image to text1
Language Modeling1
Language Modelling1
Multi-Task Learning1
Object Categorization1
Object Localization1
Referring Expression1
Retrieval1

Usage over time archive 2025-07-28

Papers per year tagged with VL-T5: 2021 to 2024, peak 2 2 0 2021: 2 papers 2021 2022: 2 papers 2022 2023: 0 papers 2023 2024: 1 paper 2024
Papers per year the archive tags with this method, by the paper's archive date (5 dated). Bars are counts, not a trend claim.

Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).

Categories archive 2025-07-28

Vision and Language Pre-Trained Models

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections