Methods › Computer Vision › Vision and Language Pre-Trained Models › Florence

Florence

12 papers tagged archive 2025-07-28

Introduced by Lu Yuan et al. in Florence: A New Foundation Model for Computer Vision

archive 2025-07-28 Description, source and code snippet are the archive's method entry.

Florence is a computer vision foundation model aiming to learn universal visual-language representations that be adapted to various computer vision tasks, visual question answering, image captioning, video retrieval, among other tasks. Florence's workflow consists of data curation, unified learning, Transformer architectures and adaption. Florence is pre-trained in an image-label-description space, utilizing a unified image-text contrastive learning. It involves a two-tower architecture: 12-layer Transformer for the language encoder, and a Vision Transformer for the image encoder. Two linear projection layers are added on top of the image encoder and language encoder to match the dimensions of image and language features. Compared to previous methods for cross-modal shared representations, Florence expands beyond simple classification and retrieval capabilities to advanced representations that support object level, multiple modality, and videos respectively.

PaperSource

Papers archive 2025-07-28

12 shown of 12, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.

Tasks archive 2025-07-28

20 shown of 36 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.

TaskPapers
Image Classification3
Object Detection3
object-detection3
Benchmarking2
image-classification2
3D Face Reconstruction1
Action Classification1
Action Recognition1
Action Recognition In Videos1
Cross-Modal Retrieval1
Data Integration1
Face Alignment1
Face Model1
Face Reconstruction1
GPU1
Image Captioning1
Language Modeling1
Language Modelling1
Object1
Optical Character Recognition1

Usage over time archive 2025-07-28

Papers per year tagged with Florence: 2021 to 2025, peak 6 6 0 2021: 1 paper 2021 2022: 1 paper 2022 2023: 4 papers 2023 2024: 0 papers 2024 2025: 6 papers 2025
Papers per year the archive tags with this method, by the paper's archive date (12 dated). Bars are counts, not a trend claim.

Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).

Categories archive 2025-07-28

Vision and Language Pre-Trained Models

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections