Papers › Context-Aware Alignment and Mutual Masking for 3D-Language Pre-Training

Context-Aware Alignment and Mutual Masking for 3D-Language Pre-Training

1 Jan 2023CVPR 2023 1archive 2025-07-28

Zhao Jin, Munawar Hayat, Yuwei Yang, Yulan Guo, Yinjie Lei

3D visual language reasoning plays an important role in effective human-computer interaction. The current approaches for 3D visual reasoning are task-specific, and lack pre-training methods to learn generic representations that can transfer across various tasks. Despite the encouraging progress in vision-language pre-training for image-text data, 3D-language pre-training is still an open issue due to limited 3D-language paired data, highly sparse and irregular structure of point clouds and ambiguities in spatial relations of 3D objects with viewpoint changes. In this paper, we present a generic 3D-language pre-training approach, that tackles multiple facets of 3D-language reasoning by learning universal representations. Our learning objective constitutes two main parts. 1) Context aware spatial-semantic alignment to establish fine-grained correspondence between point clouds and texts. It reduces relational ambiguities by aligning 3D spatial relationships with textual semantic context. 2) Mutual 3D-Language Masked modeling to enable cross-modality information exchange. Instead of reconstructing sparse 3D points for which language can hardly provide cues, we propose masked proposal reasoning to learn semantic class and mask-invariant representations. Our proposed 3D-language pre-training method achieves promising results once adapted to various downstream tasks, including 3D visual grounding, 3D dense captioning and 3D question answering. Our codes are available at https://github.com/leolyj/3D-VLP

PaperPDFCode

Code

leolyj/3d-vlp officialmentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

3D dense captioning3D visual groundingDense CaptioningQuestion AnsweringVisual GroundingVisual Reasoning

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
3D dense captioning ScanRefer Dataset 3D-VLP BLEU-4 31.87 #11 of 12 Archive leaderboard report
3D dense captioning ScanRefer Dataset 3D-VLP CIDEr 50.02 #11 of 12 Archive leaderboard report
3D dense captioning ScanRefer Dataset 3D-VLP METEOR 24.53 #11 of 12 Archive leaderboard report
3D dense captioning ScanRefer Dataset 3D-VLP ROUGE-L 51.17 #11 of 12 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

AWARE

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections