{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/towards-holistic-surgical-scene-understanding","title":"Towards Holistic Surgical Scene Understanding","arxiv_id":"2212.04582","date":"2022-12-08","proceeding":null,"authors":["Natalia Valderrama","Paola Ruiz Puentes","Isabela Hernández","Nicolás Ayobi","Mathilde Verlyk","Jessica Santander","Juan Caicedo","Nicolás Fernández","Pablo Arbeláez"],"abstract":"Most benchmarks for studying surgical interventions focus on a specific challenge instead of leveraging the intrinsic complementarity among different tasks. In this work, we present a new experimental framework towards holistic surgical scene understanding. First, we introduce the Phase, Step, Instrument, and Atomic Visual Action recognition (PSI-AVA) Dataset. PSI-AVA includes annotations for both long-term (Phase and Step recognition) and short-term reasoning (Instrument detection and novel Atomic Action recognition) in robot-assisted radical prostatectomy videos. Second, we present Transformers for Action, Phase, Instrument, and steps Recognition (TAPIR) as a strong baseline for surgical scene understanding. TAPIR leverages our dataset's multi-level annotations as it benefits from the learned representation on the instrument detection task to improve its classification capacity. Our experimental results in both PSI-AVA and other publicly available databases demonstrate the adequacy of our framework to spur future research on holistic surgical scene understanding.","url_abs":"https://arxiv.org/abs/2212.04582v4","url_pdf":"https://arxiv.org/pdf/2212.04582v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"towards-holistic-surgical-scene-understanding","repo_url":"https://github.com/bcv-uniandes/tapir","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"atomic-action-recognition","task_name":"Atomic action recognition"},{"task_slug":"scene-understanding","task_name":"Scene Understanding"},{"task_slug":"surgical-phase-recognition","task_name":"Surgical phase recognition"}],"methods":[],"datasets_introduced":[{"slug":"psi-ava","name":"PSI-AVA","full_name":""}],"methods_introduced":[],"results":[{"leaderboard":"/sota/surgical-phase-recognition-on-misaw","task":"Surgical phase recognition","dataset":"MISAW","model":"TAPIR","rank_in_archive_order":3,"of":3,"metrics":{"mAP":"94.24"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2212.04582","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}