Papers › Complete 3d relationships extraction modality alignment network for 3d dense captioning

Complete 3d relationships extraction modality alignment network for 3d dense captioning

1 Aug 2024IEEE Transactions on Visualization and Computer Graphics. 2024 8archive 2025-07-28

Aihua Mao, Zhi Yang, Wanxin Chen, Ran Yi, Yong-Jin Liu

3D dense captioning aims to semantically describe each object detected in a 3D scene, which plays a significant role in 3D scene understanding. Previous works lack a complete definition of 3D spatial relationships and the directly integrate visual and language modalities, thus ignoring the discrepancies between the two modalities. To address these issues, we propose a novel complete 3D relationship extraction modality alignment network, which consists of three steps: 3D object detection, complete 3D relationships extraction, and modality alignment caption. To comprehensively capture the 3D spatial relationship features, we define a complete set of 3D spatial relationships, including the local spatial relationship between objects and the global spatial relationship between each object and the entire scene. To this end, we propose a complete 3D relationships extraction module based on message passing and self-attention to mine multi-scale spatial relationship features and inspect the transformation to obtain features in different views. In addition, we propose the modality alignment caption module to fuse multi-scale relationship features and generate descriptions to bridge the semantic gap from the visual space to the language space with the prior information in the word embedding, and help generate improved descriptions for the 3D scene. Extensive experiments demonstrate that the proposed model outperforms the state-of-the-art methods on the ScanRefer and Nr3D datasets.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

3D Object Detection3D dense captioningDense CaptioningObject DetectionScene Understandingobject-detection

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
3D dense captioning Nr3D REMAN BLEU-4 20.37 #7 of 10 Archive leaderboard report
3D dense captioning Nr3D REMAN CIDEr 34.81 #7 of 10 Archive leaderboard report
3D dense captioning Nr3D REMAN METEOR 23.01 #7 of 10 Archive leaderboard report
3D dense captioning Nr3D REMAN ROUGE-L 50.99 #7 of 10 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

SET

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections