Browse State-of-the-Art › Surgical phase recognition
Surgical phase recognition
31 papers with code · 4 benchmarks · 8 datasets archive 2025-07-28
The first 40 videos are used for training, the last 40 videos are used for testing.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
4 leaderboard tables shown for this task, 4 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| Cholec80 (6 rows) | LoViT | LoViT: Long Video Transformer for Surgical Phase Recognition | code | — | Compare |
| HeiChole Benchmark (5 rows) | MuST | MuST: Multi-Scale Transformers for Surgical Phase Recognition | code | — | Compare |
| MISAW (3 rows) | MuST | MuST: Multi-Scale Transformers for Surgical Phase Recognition | code | — | Compare |
| GraSP (2 rows) | MuST | MuST: Multi-Scale Transformers for Surgical Phase Recognition | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
8 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
2 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 31 papers with code (69 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
20 Jan 2024 3 repositories listed Syntology ran 7 of 13 samples · 6 unverified · 1 pointer-only (licence)To exploit our proposed benchmark, we introduce the Transformers for Actions, Phases, Steps, and Instrument Segmentation (TAPIS) model, a general architecture that combines a global video feature extractor with…
-
16 May 2024 2 repositories listedBy disentangling embedding spaces of different hierarchical levels, the learned multi-modal representations encode short-term and long-term surgical concepts in the same model.
-
10 Jul 2021 2 repositories listedTo address the problem, we propose a new non end-to-end training strategy and explore different designs of multi-stage architecture for surgical phase recognition task.
-
24 Mar 2020 2 repositories listedAutomatic surgical phase recognition is a challenging and crucial task with the potential to improve patient safety and become an integral part of intra-operative decision-support systems.
-
25 Mar 2025 1 repository listedOur experimental results show that SurgFM outperforms state-of-the-art models in multiple downstream tasks, including significant gains in surgical phase recognition (+8.
-
26 Feb 2025 1 repository listedFirst, to mitigate computational inefficiencies, we propose the EndoMamba backbone, optimized for real-time inference.
-
29 Jan 2025 1 repository listedAccurate surgical phase recognition is crucial for advancing computer-assisted interventions, yet the scarcity of labeled data hinders training reliable deep learning models.
-
19 Sep 2024 1 repository listedTo ensure a global understanding of the surgical procedure, we devise a phase localization strategy for SurgPLAN++ to predict phase segments across the entire video through phase proposals.
-
18 Sep 2024 1 repository listedEndoscopic Submucosal Dissection (ESD) is a minimally invasive procedure initially developed for early gastric cancer treatment and has expanded to address diverse gastrointestinal lesions.
-
10 Sep 2024 1 repository listedIn the other stream, DACAT fine-tunes a new frame encoder to extract the frame-wise feature at the current moment.
-
Surgformer: Surgical Transformer with Hierarchical Temporal Attention for Surgical Phase Recognition7 Aug 2024 1 repository listedMoreover, we propose a novel Hierarchical Temporal Attention (HTA) to capture both global and local information within varied temporal resolutions from a target frame-centric perspective.
-
24 Jul 2024 1 repository listedPhase recognition in surgical videos is crucial for enhancing computer-aided surgical systems as it enables automated understanding of sequential procedural stages.
-
11 Jul 2024 1 repository listedInspired by the recent success of Mamba, a state space model with linear scalability in sequence length, this paper presents SR-Mamba, a novel attention-free model specifically tailored to meet the challenges of…
-
30 May 2024 1 repository listedTo address this issue, we introduce a new egocentric open surgery video dataset for phase recognition, named EgoSurgery-Phase.
-
11 Dec 2023 1 repository listedRecently, spatiotemporal graphs have emerged as a concise and elegant manner of representing video clips in an object-centric fashion, and have shown to be useful for downstream tasks such as action recognition.
-
23 Aug 2023 1 repository listedTo fully exploit the power of SSL, we create sizable unlabeled endoscopic video datasets for training MSNs.
-
19 Jul 2023 1 repository listedIn addition, we propose to train the feature extractor, a standard CNN, together with an LSTM on preferably long video segments, i.
-
23 May 2023 1 repository listedSurgical phase recognition is a basic component for different context-aware applications in computer- and robot-assisted surgery.
-
15 May 2023 1 repository listedOur results demonstrate the effectiveness of our approach in achieving state-of-the-art performance of surgical phase recognition on two datasets of different surgical procedures and temporal sequencing characteristics…
-
30 Mar 2023 1 repository listedIn this work, we investigate the need for endoscopy domain-specific pretraining based on downstream objectives.
-
1 Jan 2023 1 repository listedWe highlight that the inference time of SKiT is constant, and independent from the input length, making it a stable choice for keeping a record of important global information, that appears on long surgical videos,…
-
8 Dec 2022 1 repository listedSecond, we present Transformers for Action, Phase, Instrument, and steps Recognition (TAPIR) as a strong baseline for surgical scene understanding.
-
1 Jul 2022 1 repository listedCorrect transfer of these methods to surgery, as described and conducted in this work, leads to substantial performance gains over generic uses of SSL - up to 7.
-
19 May 2022 1 repository listedOur key insight is to distill knowledge from publicly available models trained on large generic datasets4 to facilitate the self-supervised learning of surgical videos.
-
16 Feb 2022 1 repository listedOur study uncovers unique insights of surgical phase recognition with timestamp supervisions: 1) timestamp annotation can reduce 74% annotation time compared with the full annotation, and surgeons tend to annotate those…
-
22 Nov 2021 1 repository listedAutomatic surgical phase recognition plays a vital role in robot-assisted surgeries.
-
2 Jul 2021 1 repository listedIn particular, we propose (I) an end-to-end recurrent neural network to recognize the lens-implantation phase and (II) a novel semantic segmentation network to segment the lens and pupil after the implantation phase.
-
17 Mar 2021 1 repository listedIn this paper, we introduce, for the first time in surgical workflow analysis, Transformer to reconsider the ignored complementary effects of spatial and temporal features for accurate surgical phase recognition.
-
13 Jul 2019 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)Mutually leveraging both low-level feature sharing and high-level prediction correlating, our MTRCNet-CL method can encourage the interactions between the two tasks to a large extent, and hence can bring about benefits…
-
30 Nov 2018 1 repository listedVision algorithms capable of interpreting scenes from a real-time video stream are necessary for computer-assisted surgery systems to achieve context-aware behavior.
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections