{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/u2-bench-benchmarking-large-vision-language","title":"U2-BENCH: Benchmarking Large Vision-Language Models on Ultrasound Understanding","arxiv_id":"2505.17779","date":"2025-05-23","proceeding":null,"authors":["Anjie Le","Henan Liu","Yue Wang","Zhenyu Liu","Rongkun Zhu","Taohan Weng","Jinze Yu","Boyang Wang","Yalun Wu","Kaiwen Yan","Quanlin Sun","Meirui Jiang","Jialun Pei","Siya Liu","Haoyun Zheng","Zhoujun Li","Alison Noble","Jacques Souquet","Xiaoqing Guo","Manxi Lin","Hongcheng Guo"],"abstract":"Ultrasound is a widely-used imaging modality critical to global healthcare, yet its interpretation remains challenging due to its varying image quality on operators, noises, and anatomical structures. Although large vision-language models (LVLMs) have demonstrated impressive multimodal capabilities across natural and medical domains, their performance on ultrasound remains largely unexplored. We introduce U2-BENCH, the first comprehensive benchmark to evaluate LVLMs on ultrasound understanding across classification, detection, regression, and text generation tasks. U2-BENCH aggregates 7,241 cases spanning 15 anatomical regions and defines 8 clinically inspired tasks, such as diagnosis, view recognition, lesion localization, clinical value estimation, and report generation, across 50 ultrasound application scenarios. We evaluate 20 state-of-the-art LVLMs, both open- and closed-source, general-purpose and medical-specific. Our results reveal strong performance on image-level classification, but persistent challenges in spatial reasoning and clinical language generation. U2-BENCH establishes a rigorous and unified testbed to assess and accelerate LVLM research in the uniquely multimodal domain of medical ultrasound imaging.","url_abs":"https://arxiv.org/abs/2505.17779v1","url_pdf":"https://arxiv.org/pdf/2505.17779v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"benchmarking","task_name":"Benchmarking"},{"task_slug":"spatial-reasoning","task_name":"Spatial Reasoning"},{"task_slug":"text-generation","task_name":"Text Generation"}],"methods":[],"datasets_introduced":[{"slug":"u2-bench","name":"U2-Bench","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}