Papers › PLA: Language-Driven Open-Vocabulary 3D Scene Understanding

PLA: Language-Driven Open-Vocabulary 3D Scene Understanding

29 Nov 2022CVPR 2023 1arXiv:2211.16312archive 2025-07-28

Runyu Ding, Jihan Yang, Chuhui Xue, Wenqing Zhang, Song Bai, Xiaojuan Qi

Open-vocabulary scene understanding aims to localize and recognize unseen categories beyond the annotated label space. The recent breakthrough of 2D open-vocabulary perception is largely driven by Internet-scale paired image-text data with rich vocabulary concepts. However, this success cannot be directly transferred to 3D scenarios due to the inaccessibility of large-scale 3D-text pairs. To this end, we propose to distill knowledge encoded in pre-trained vision-language (VL) foundation models through captioning multi-view images from 3D, which allows explicitly associating 3D and semantic-rich captions. Further, to foster coarse-to-fine visual-semantic representation learning from captions, we design hierarchical 3D-caption pairs, leveraging geometric constraints between 3D scenes and multi-view images. Finally, by employing contrastive learning, the model learns language-aware embeddings that connect 3D and text for open-vocabulary tasks. Our method not only remarkably outperforms baseline methods by 25.8% ∼ 44.7% hIoU and 14.5% ∼ 50.4% hAP₅₀ in open-vocabulary semantic and instance segmentation, but also shows robust transferability on challenging zero-shot domain transfer tasks. See the project website at https://dingry.github.io/projects/PLA.

PaperPDFConference PDFCode

In Syntology View this paper on Syntology: its repositories, every harvested function with whether it ran, its licence and the call to fetch it.

Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

cvmi-lab/pla officialmentioned in papermentioned on GitHubpytorchApache-2.0 report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

3D Open-Vocabulary Instance SegmentationContrastive LearningInstance SegmentationRepresentation LearningScene UnderstandingSemantic Segmentation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
3D Open-Vocabulary Instance Segmentation S3DIS PLA AP50 Base B6/N6 46.9 #2 of 4 Archive leaderboard report
3D Open-Vocabulary Instance Segmentation S3DIS PLA AP50 Base B8/N4 59.0 #2 of 4 Archive leaderboard report
3D Open-Vocabulary Instance Segmentation S3DIS PLA AP50 Novel B6/N6 9.8 #2 of 4 Archive leaderboard report
3D Open-Vocabulary Instance Segmentation S3DIS PLA AP50 Novel B8/N4 8.6 #2 of 4 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections