Methods › Computer Vision › Vision and Language Pre-Trained Models › Florence
Florence
Introduced by Lu Yuan et al. in Florence: A New Foundation Model for Computer Vision
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
Florence is a computer vision foundation model aiming to learn universal visual-language representations that be adapted to various computer vision tasks, visual question answering, image captioning, video retrieval, among other tasks. Florence's workflow consists of data curation, unified learning, Transformer architectures and adaption. Florence is pre-trained in an image-label-description space, utilizing a unified image-text contrastive learning. It involves a two-tower architecture: 12-layer Transformer for the language encoder, and a Vision Transformer for the image encoder. Two linear projection layers are added on top of the image encoder and language encoder to match the dimensions of image and language features. Compared to previous methods for cross-modal shared representations, Florence expands beyond simple classification and retrieval capabilities to advanced representations that support object level, multiple modality, and videos respectively.
Papers archive 2025-07-28
12 shown of 12, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
FRED: The Florence RGB-Event Drone Dataset 5 Jun 2025 · 0 repositories · arXiv:2506.05163
-
PsOCR: Benchmarking Large Multimodal Models for Optical Character Recognition in Low-resource Pashto Language 15 May 2025 · 1 repository · arXiv:2505.10055
-
Quantifying walkable accessibility to urban services: An application to Florence, Italy 17 Apr 2025 · 0 repositories · arXiv:2504.12934
-
Fine-Tuning Florence2 for Enhanced Object Detection in Un-constructed Environments: Vision-Language Model Approach 6 Mar 2025 · 0 repositories · arXiv:2503.04918
-
TexLiDAR: Automated Text Understanding for Panoramic LiDAR Data 5 Feb 2025 · 1 repository · arXiv:2502.04385
-
Learning Nonverbal Cues in Multiparty Social Interactions for Robotic Facilitators 18 Jan 2025 · 0 repositories · arXiv:2501.10857
-
Smart City Digital Twin Framework for Real-Time Multi-Data Integration and Wide Public Distribution 23 Sep 2023 · 0 repositories · arXiv:2309.13394
-
ASM: Adaptive Skinning Model for High-Quality 3D Face Modeling 19 Apr 2023 · 0 repositories · arXiv:2304.09423
-
DIME-FM: DIstilling Multimodal and Efficient Foundation Models 31 Mar 2023 · 0 repositories · arXiv:2303.18232
-
DIME-FM : DIstilling Multimodal and Efficient Foundation Models 1 Jan 2023 · 0 repositories
-
The Florence 4D Facial Expression Dataset 30 Oct 2022 · 0 repositories · arXiv:2210.16807
-
Florence: A New Foundation Model for Computer Vision 22 Nov 2021 · 2 repositories · arXiv:2111.11432
Tasks archive 2025-07-28
20 shown of 36 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections