{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/florence-2-advancing-a-unified-representation","title":"Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks","arxiv_id":"2311.06242","date":"2023-11-10","proceeding":"CVPR 2024 1","authors":["Bin Xiao","Haiping Wu","Weijian Xu","Xiyang Dai","Houdong Hu","Yumao Lu","Michael Zeng","Ce Liu","Lu Yuan"],"abstract":"We introduce Florence-2, a novel vision foundation model with a unified, prompt-based representation for a variety of computer vision and vision-language tasks. While existing large vision models excel in transfer learning, they struggle to perform a diversity of tasks with simple instructions, a capability that implies handling the complexity of various spatial hierarchy and semantic granularity. Florence-2 was designed to take text-prompt as task instructions and generate desirable results in text forms, whether it be captioning, object detection, grounding or segmentation. This multi-task learning setup demands large-scale, high-quality annotated data. To this end, we co-developed FLD-5B that consists of 5.4 billion comprehensive visual annotations on 126 million images, using an iterative strategy of automated image annotation and model refinement. We adopted a sequence-to-sequence structure to train Florence-2 to perform versatile and comprehensive vision tasks. Extensive evaluations on numerous tasks demonstrated Florence-2 to be a strong vision foundation model contender with unprecedented zero-shot and fine-tuning capabilities.","url_abs":"https://arxiv.org/abs/2311.06242v1","url_pdf":"https://arxiv.org/pdf/2311.06242v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"florence-2-advancing-a-unified-representation","repo_url":"https://github.com/retkowsky/florence-2","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"diversity","task_name":"Diversity"},{"task_slug":"multi-task-learning","task_name":"Multi-Task Learning"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"},{"task_slug":"visual-grounding","task_name":"Visual Grounding"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/visual-grounding-on-refcoco-test-b","task":"Visual Grounding","dataset":"RefCOCO+ test B","model":"Florence-2-large-ft","rank_in_archive_order":1,"of":6,"metrics":{"Accuracy (%)":"92.0"},"uses_additional_data":true},{"leaderboard":"/sota/visual-grounding-on-refcoco-testa","task":"Visual Grounding","dataset":"RefCOCO+ testA","model":"Florence-2-large-ft","rank_in_archive_order":1,"of":7,"metrics":{"Accuracy (%)":" 95.3"},"uses_additional_data":true},{"leaderboard":"/sota/visual-grounding-on-refcoco-val","task":"Visual Grounding","dataset":"RefCOCO+ val","model":"Florence-2-large-ft","rank_in_archive_order":1,"of":6,"metrics":{"Accuracy (%)":"93.4"},"uses_additional_data":true}],"syntology":{"syntology_url":"https://syntology.ai/paper/2311.06242","atlas_url":"https://app.syntology.ai/?focus=2311.06242","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}