Browse State-of-the-Art › Scene Segmentation
Scene Segmentation
141 papers with code · 6 benchmarks · 10 datasets archive 2025-07-28
Scene segmentation is the task of splitting a scene into its various object components.
Image adapted from Temporally coherent 4D reconstruction of complex dynamic scenes.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
6 leaderboard tables shown for this task, 6 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| SUN-RGBD (5 rows) | ICM | Scene Parsing via Integrated Classification Model and... | code | — | Compare |
| ScanNet (3 rows) | 3DMV | 3DMV: Joint 3D-Multi-View Prediction for 3D Semantic Scene Segmentation | code | — | Compare |
| StreetHazards (3 rows) | Mask2Anomaly | Unmasking Anomalies in Road-Scene Segmentation | code | — | Compare |
| MovieNet (2 rows) | NeighborNet | Neighbor Relations Matter in Video Scene Detection | code | — | Compare |
| NYU Depth v2 (1 row) | Dilated FCN-2s RGB | Efficient Yet Deep Convolutional Neural Networks for Semantic Segmentation | code | — | Compare |
| UAVid (1 row) | UNetFormer | UNetFormer: A UNet-like Transformer for Efficient Semantic... | code | Syntology ran 2 of 2 samples · 0 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
10 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 141 papers with code (283 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
2 Dec 2016 110 repositories listed Syntology ran 89 of 164 samples · 75 unverified · 90 pointer-only (licence)Point cloud is an important type of geometric data structure.
-
2 Nov 2015 74 repositories listed Syntology ran 9 of 44 samples · 35 unverified · 10 pointer-only (licence)We show that SegNet provides good performance with competitive inference time and more efficient inference memory-wise as compared to other architectures.
-
20 May 2016 37 repositories listedConvolutional networks are powerful visual models that yield hierarchies of features.
-
16 Dec 2020 24 repositories listedFor example, on the challenging S3DIS dataset for large-scale semantic scene segmentation, the Point Transformer attains an mIoU of 70.
-
9 Sep 2018 12 repositories listed Syntology ran 0 of 7 samples · 7 unverified · 5 pointer-only (licence)Specifically, we append two types of attention modules on top of traditional dilated FCN, which model the semantic interdependencies in spatial and channel dimensions respectively.
-
18 Apr 2019 10 repositories listed Syntology ran 5 of 12 samples · 7 unverified · 3 pointer-only (licence)Furthermore, these locations are continuous in space and can be learned by the network.
-
3 Jan 2018 9 repositories listedWe propose and study a task we name panoptic segmentation (PS).
-
11 Aug 2019 6 repositories listedBy viewing the indices as a function of the feature map, we introduce the concept of "learning to index", and present a novel index-guided encoder-decoder framework where indices are self-learned adaptively from data…
-
3 May 2019 5 repositories listed Syntology ran 1 of 2 samples · 1 unverifiedIn this work we introduce a novel, CNN-based architecture that can be trained end-to-end to deliver seamless scene segmentation results.
-
6 Apr 2020 4 repositories listed Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)Scene, as the crucial unit of storytelling in movies, contains complex activities of actors and their interactions in a physical environment.
-
8 Jul 2019 4 repositories listed Syntology ran 5 of 8 samples · 3 unverified · 5 pointer-only (licence)The computation cost and memory footprints of the voxel-based models grow cubically with the input resolution, making it memory-prohibitive to scale up the resolution.
-
5 Apr 2024 3 repositories listed Syntology ran 8 of 14 samples · 6 unverified · 6 pointer-only (licence)To reduce the reliance on large-scale datasets, recent works in 3D segmentation resort to few-shot learning.
-
23 Mar 2021 3 repositories listedCurrent scene segmentation methods suffer from cumbersome model structures and high computational complexity, impeding their applications to real-world scenarios that require real-time processing.
-
30 Jan 2020 3 repositories listedIn 2015 we began a sub-challenge at the EndoVis workshop at MICCAI in Munich using endoscope images of ex-vivo tissue with automatically generated annotations from robot forward kinematics and instrument CAD models.
-
2 Dec 2016 3 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)Colorectal cancer (CRC) is the third cause of cancer death worldwide.
-
21 Nov 2023 2 repositories listedBased on such observation, we propose a depth-aware framework to explicitly leverage depth estimation to mix the categories and facilitate the two complementary tasks, i.
-
15 Sep 2023 2 repositories listedReal-time transportation surveillance is an essential part of the intelligent transportation system (ITS).
-
20 Mar 2023 2 repositories listedIn this work, we present a zero-shot volumetric open-vocabulary semantic scene segmentation method.
-
26 Nov 2022 2 repositories listed Syntology ran 3 of 10 samples · 7 unverifiedSemantic segmentation models classify pixels into a set of known (``in-distribution'') visual classes.
-
4 Jul 2022 2 repositories listedWith DIAL-Filters, we design both unsupervised and supervised frameworks for nighttime driving-scene segmentation, which can be trained in an end-to-end manner.
-
23 Jun 2022 2 repositories listedOur empirical analysis suggests that without the high cost of data collection and annotation, we can achieve decent surgical instrument segmentation performance.
-
26 May 2022 2 repositories listed Syntology ran 3 of 10 samples · 7 unverifiedMulti-sensor fusion is essential for an accurate and reliable autonomous driving system.
-
4 Apr 2022 2 repositories listedRobust visual recognition under adverse weather conditions is of great importance in real-world applications.
-
3 Dec 2021 2 repositories listedIn this paper, we propose a series of modular operations for effective geometric feature learning from 3D triangle meshes.
-
21 Sep 2021 2 repositories listedThe last layer of FCN is typically a global classifier (1x1 convolution) to recognize each pixel to a semantic label.
-
29 Mar 2021 2 repositories listedDomain adaptation is to transfer the shared knowledge learned from the source domain to a new environment, i.
-
29 Mar 2021 2 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedEnhancing the generalization capability of deep neural networks to unseen domains is crucial for safety-critical applications in the real world such as autonomous driving.
-
7 Jul 2020 2 repositories listed Syntology ran 1 of 8 samples · 7 unverified · 8 pointer-only (licence)Learning to infer graph representations and performing spatial reasoning in a complex surgical environment can play a vital role in surgical scene understanding in robotic surgery.
-
3 Apr 2020 2 repositories listedGiven an input image and corresponding ground truth, Affinity Loss constructs an ideal affinity map to supervise the learning of Context Prior.
-
4 Dec 2016 2 repositories listedRobust perception-action models should be learned from training data with diverse visual appearances and realistic behaviors, yet current approaches to deep visuomotor policy learning have been generally limited to…
Syntology lines on 13 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections