Papers › GROUNDHOG: Grounding Large Language Models to Holistic Segmentation

GROUNDHOG: Grounding Large Language Models to Holistic Segmentation

26 Feb 2024CVPR 2024 1arXiv:2402.16846archive 2025-07-28

Yichi Zhang, Ziqiao Ma, Xiaofeng Gao, Suhaila Shakiah, Qiaozi Gao, Joyce Chai

Most multimodal large language models (MLLMs) learn language-to-object grounding through causal language modeling where grounded objects are captured by bounding boxes as sequences of location tokens. This paradigm lacks pixel-level representations that are important for fine-grained visual understanding and diagnosis. In this work, we introduce GROUNDHOG, an MLLM developed by grounding Large Language Models to holistic segmentation. GROUNDHOG incorporates a masked feature extractor and converts extracted features into visual entity tokens for the MLLM backbone, which then connects groundable phrases to unified grounding masks by retrieving and merging the entity masks. To train GROUNDHOG, we carefully curated M3G2, a grounded visual instruction tuning dataset with Multi-Modal Multi-Grained Grounding, by harvesting a collection of segmentation-grounded datasets with rich annotations. Our experimental results show that GROUNDHOG achieves superior performance on various language grounding tasks without task-specific fine-tuning, and significantly reduces object hallucination. GROUNDHOG also demonstrates better grounding towards complex forms of visual input and provides easy-to-understand diagnosis in failure cases.

PaperPDFConference PDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Generalized Referring Expression SegmentationHallucinationLanguage ModelingLanguage ModellingObject HallucinationReferring Expression Segmentation

1 archive task tag without a task page not shown.

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Generalized Referring Expression Segmentation gRefCOCO GROUNDHOG gIoU 66.70 #6 of 13 Archive leaderboard report
Referring Expression Segmentation PhraseCut GROUNDHOG Mean IoU 54.5 #2 of 6 Archive leaderboard report
Referring Expression Segmentation RefCOCO+ test B GROUNDHOG Overall IoU 64.9 #9 of 30 Archive leaderboard report
Referring Expression Segmentation RefCOCO+ testA GROUNDHOG Overall IoU 75.0 #11 of 30 Archive leaderboard report
Referring Expression Segmentation RefCOCO+ val GROUNDHOG Overall IoU 70.5 #13 of 33 Archive leaderboard report
Referring Expression Segmentation RefCOCOg-test GROUNDHOG Overall IoU 74.6 #8 of 18 Archive leaderboard report
Referring Expression Segmentation RefCOCOg-val GROUNDHOG Overall IoU 74.1 #9 of 23 Archive leaderboard report
Referring Expression Segmentation RefCoCo val GROUNDHOG Overall IoU 78.5 #14 of 37 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections