Browse State-of-the-Art › Scene Graph Generation
Scene Graph Generation
151 papers with code · 7 benchmarks · 11 datasets archive 2025-07-28
A scene graph is a structured representation of an image, where nodes in a scene graph correspond to object bounding boxes with their object categories, and edges correspond to their pairwise relationships between objects. The task of Scene Graph Generation is to generate a visually-grounded scene graph that most accurately correlates with an image.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
7 leaderboard tables shown for this task, 7 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| Visual Genome (19 rows) | SpeaQ (without reweighting) | Groupwise Query Specialization and Quality-Aware Multi-Assignment... | code | Syntology ran 1 of 3 samples · 2 unverified | Compare |
| 4D-OR (5 rows) | ORacle | ORacle: Large Vision-Language Models for Knowledge-Guided Holistic... | code | — | Compare |
| 3R-Scan (2 rows) | SceneGraphFusion | SceneGraphFusion: Incremental 3D Scene Graph Prediction from RGB-D... | code | — | Compare |
| VRD (2 rows) | FactorizableNet | Factorizable Net: An Efficient Subgraph-based Framework for Scene... | code | — | Compare |
| GQA (1 row) | KnowZRel | KnowZRel: Common Sense Knowledge-based Zero-Shot Relationship... | code | — | Compare |
| MM-OR (1 row) | MM2SG | MM-OR: A Large Multimodal Operating Room Dataset for Semantic... | code | — | Compare |
| MS-COCO (1 row) | NeuSyRE | NeuSyRE: Neuro-Symbolic Visual Understanding and Reasoning... | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
11 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
2 subtasks in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 151 papers with code (318 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
27 Feb 2020 6 repositories listed Syntology ran 3 of 6 samples · 3 unverified · 6 pointer-only (licence)Today's scene graph generation (SGG) task is still far from practical, mainly due to the severe training bias, e.
-
5 Dec 2018 6 repositories listed Syntology ran 4 of 14 samples · 10 unverifiedWe propose to compose dynamic tree structures that place the objects in an image into a visual context, helping visual reasoning tasks such as scene graph generation and visual Q&A.
-
10 Jan 2017 5 repositories listed Syntology ran 3 of 5 samples · 2 unverified · 3 pointer-only (licence)In this work, we explicitly model the objects and their relationships using scene graphs, a visually-grounded graphical structure of an image.
-
21 Jun 2021 4 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)The key to our method is a set of learnable triplet queries and a structured triplet detector which could be jointly optimized from the training set in an end-to-end manner.
-
1 Apr 2021 4 repositories listedScene graph generation is an important visual understanding task with a broad range of vision applications.
-
13 Jun 2024 3 repositories listedThis paper constructs a large-scale dataset for SGG in large-size VHR SAI with image sizes ranging from 512 x 768 to 27, 860 x 31, 096 pixels, named STAR (Scene graph generaTion in lArge-size satellite imageRy),…
-
16 May 2024 3 repositories listed Syntology ran 13 of 14 samples · 1 unverified · 10 pointer-only (licence)To facilitate research in this new area, we build a richly annotated PSG-4D dataset consisting of 3K RGB-D videos with a total of 1M frames, each of which is labeled with 4D panoptic segmentation masks as well as…
-
28 Nov 2023 3 repositories listed Syntology ran 11 of 11 samples · 0 unverified · 3 pointer-only (licence)PVSG relates to the existing video scene graph generation (VidSGG) problem, which focuses on temporal interactions between humans and objects grounded with bounding boxes in videos.
-
18 Aug 2023 3 repositories listed Syntology ran 22 of 30 samples · 8 unverifiedIn this paper, we propose RLIPv2, a fast converging model that enables the scaling of relational pre-training to large-scale pseudo-labelled scene graph data.
-
30 Nov 2022 3 repositories listedHowever, it is difficult to draw a proper scene graph for image retrieval, image generation, and multi-modal applications.
-
8 Mar 2019 3 repositories listed Syntology ran 3 of 8 samples · 5 unverifiedMore specifically, we show that the statistical correlations between objects appearing in images and their relationships, can be explicitly represented by a structured knowledge graph, and a routing mechanism is learned…
-
7 Mar 2019 3 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedThe first, Entity Instance Confusion, occurs when the model confuses multiple instances of the same type of entity (e.
-
15 Nov 2018 3 repositories listedIn this paper, we present a method that improves scene graph generation by explicitly modeling inter-dependency among the entire object instances.
-
1 Aug 2018 3 repositories listedWe propose a novel scene graph generation model called Graph R-CNN, that is both effective and efficient at detecting objects and their relations in images.
-
22 Jun 2017 3 repositories listedGraphs are a useful abstraction of image content.
-
30 Jun 2023 2 repositories listedFor further understanding of comics, an automated approach is needed to link text in comics to characters speaking the words.
-
22 Mar 2022 2 repositories listedScene graph generation (SGG) is designed to extract (subject, predicate, object) triplets in images.
-
26 Jul 2021 2 repositories listed Syntology ran 7 of 14 samples · 7 unverified · 5 pointer-only (licence)Compared to the task of scene graph generation from images, it is more challenging because of the dynamic relationships between objects and the temporal dependencies between frames allowing for a richer semantic…
-
27 Mar 2021 2 repositories listedScene graphs are a compact and explicit representation successfully used in a variety of 2D scene understanding tasks.
-
7 Jul 2020 2 repositories listed Syntology ran 1 of 8 samples · 7 unverified · 8 pointer-only (licence)Learning to infer graph representations and performing spatial reasoning in a complex surgical environment can play a vital role in surgical scene understanding in robotic surgery.
-
17 Jun 2020 2 repositories listed Syntology ran 3 of 11 samples · 8 unverifiedScene graph generation models understand the scene through object and predicate recognition, but are prone to mistakes due to the challenges of perception in the wild.
-
16 Jul 2018 2 repositories listedRecent approaches on visual scene understanding attempt to build a scene graph -- a computational representation of objects and their pairwise relationships.
-
9 Jun 2025 1 repository listedScene-Graph Generation (SGG) seeks to recognize objects in an image and distill their salient pairwise relationships.
-
30 May 2025 1 repository listedThis new dataset and benchmark set a new foundation for OR perception, offering a rich, multimodal resource for next-generation clinical perception.
-
26 May 2025 1 repository listedIn this work, we introduce Text-Scene Graph (TSG) Bench, a benchmark designed to systematically assess LLMs' ability to (1) understand scene graphs and (2) generate them from textual narratives.
-
21 May 2025 1 repository listedDespite the dominance of convolutional and transformer-based architectures in image-to-image retrieval, these models are prone to biases arising from low-level visual features, such as color.
-
18 Mar 2025 1 repository listedThis embedding then serves as the input to task-specific heads for object classification, scene graph generation, etc.
-
4 Mar 2025 1 repository listedOperating rooms (ORs) are complex, high-stakes environments requiring precise understanding of interactions among medical staff, tools, and equipment for enhancing surgical assistance, situational awareness, and patient…
-
21 Feb 2025 1 repository listedThe generalisability of Scene Graph Generation (SGG) methods is crucial for reliable reasoning and real-world applicability.
-
21 Feb 2025 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)Existing Video Scene Graph Generation (VidSGG) studies are trained in a fully supervised manner, which requires all frames in a video to be annotated, thereby incurring high annotation cost compared to Image Scene Graph…
Syntology lines on 13 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections