Papers › Dissecting Self-Supervised Learning Methods for Surgical Computer Vision

Dissecting Self-Supervised Learning Methods for Surgical Computer Vision

1 Jul 2022arXiv:2207.00449archive 2025-07-28

Sanat Ramesh, Vinkle Srivastav, Deepak Alapatt, Tong Yu, Aditya Murali, Luca Sestini, Chinedu Innocent Nwoye, Idris Hamoud, Saurav Sharma, Antoine Fleurentin, Georgios Exarchakis, Alexandros Karargyris, Nicolas Padoy

The field of surgical computer vision has undergone considerable breakthroughs in recent years with the rising popularity of deep neural network-based methods. However, standard fully-supervised approaches for training such models require vast amounts of annotated data, imposing a prohibitively high cost; especially in the clinical domain. Self-Supervised Learning (SSL) methods, which have begun to gain traction in the general computer vision community, represent a potential solution to these annotation costs, allowing to learn useful representations from only unlabeled data. Still, the effectiveness of SSL methods in more complex and impactful domains, such as medicine and surgery, remains limited and unexplored. In this work, we address this critical need by investigating four state-of-the-art SSL methods (MoCo v2, SimCLR, DINO, SwAV) in the context of surgical computer vision. We present an extensive analysis of the performance of these methods on the Cholec80 dataset for two fundamental and popular tasks in surgical context understanding, phase recognition and tool presence detection. We examine their parameterization, then their behavior with respect to training data quantities in semi-supervised settings. Correct transfer of these methods to surgery, as described and conducted in this work, leads to substantial performance gains over generic uses of SSL - up to 7.4% on phase recognition and 20% on tool presence detection - as well as state-of-the-art semi-supervised phase recognition approaches by up to 14%. Further results obtained on a highly diverse selection of surgical datasets exhibit strong generalization properties. The code is available at https://github.com/CAMMA-public/SelfSupSurg.

PaperPDFCode

Code

camma-public/selfsupsurg officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action Triplet RecognitionSelf-Supervised LearningSemantic SegmentationSurgical phase recognitionSurgical tool detection

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Action Triplet Recognition CholecT50 (Challenge) MoCo V2 Surg SSL - Rendezvous head mAP 35.7 #4 of 27 Archive leaderboard report
Semantic Segmentation Endoscapes MoCo V2 Surg SSL - DeepLabv3+ head Mean F1 73.2 #1 of 2 Archive leaderboard report
Surgical phase recognition Cholec80 MoCo V2 Surg SSL - TCN head F1 81.6 #4 of 6 Archive leaderboard report
Surgical phase recognition HeiChole Benchmark MoCo V2 Surg SSL - TCN head F1 64.7 #5 of 5 Archive leaderboard report
Surgical tool detection Cholec80 MoCo V2 Surg SSL - FCN head mAP 93.5 #1 of 6 Archive leaderboard report
Surgical tool detection HeiChole Benchmark MoCo V2 Surg SSL - FCN head mAP 66.9 #1 of 1 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

1x1 ConvolutionAttentionAverage PoolingBatch NormalizationBottleneck Residual BlockColorJitterConvolutionDense ConnectionsFeedforward NetworkGlobal Average PoolingKaiming InitializationLayer NormalizationLinear LayerMax PoolingMulti-Head AttentionNT-XentRandom Gaussian BlurRandom Resized CropReLUResidual BlockResidual ConnectionSimCLRSoftmaxVision Transformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections