Papers › GROUNDHOG: Grounding Large Language Models to Holistic Segmentation
GROUNDHOG: Grounding Large Language Models to Holistic Segmentation
Yichi Zhang, Ziqiao Ma, Xiaofeng Gao, Suhaila Shakiah, Qiaozi Gao, Joyce Chai
Most multimodal large language models (MLLMs) learn language-to-object grounding through causal language modeling where grounded objects are captured by bounding boxes as sequences of location tokens. This paradigm lacks pixel-level representations that are important for fine-grained visual understanding and diagnosis. In this work, we introduce GROUNDHOG, an MLLM developed by grounding Large Language Models to holistic segmentation. GROUNDHOG incorporates a masked feature extractor and converts extracted features into visual entity tokens for the MLLM backbone, which then connects groundable phrases to unified grounding masks by retrieving and merging the entity masks. To train GROUNDHOG, we carefully curated M3G2, a grounded visual instruction tuning dataset with Multi-Modal Multi-Grained Grounding, by harvesting a collection of segmentation-grounded datasets with rich annotations. Our experimental results show that GROUNDHOG achieves superior performance on various language grounding tasks without task-specific fine-tuning, and significantly reduces object hallucination. GROUNDHOG also demonstrates better grounding towards complex forms of visual input and provides easy-to-understand diagnosis in failure cases.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
1 archive task tag without a task page not shown.
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Generalized Referring Expression Segmentation | gRefCOCO | GROUNDHOG | gIoU | 66.70 | #6 of 13 | Archive leaderboard | report |
| Referring Expression Segmentation | PhraseCut | GROUNDHOG | Mean IoU | 54.5 | #2 of 6 | Archive leaderboard | report |
| Referring Expression Segmentation | RefCOCO+ test B | GROUNDHOG | Overall IoU | 64.9 | #9 of 30 | Archive leaderboard | report |
| Referring Expression Segmentation | RefCOCO+ testA | GROUNDHOG | Overall IoU | 75.0 | #11 of 30 | Archive leaderboard | report |
| Referring Expression Segmentation | RefCOCO+ val | GROUNDHOG | Overall IoU | 70.5 | #13 of 33 | Archive leaderboard | report |
| Referring Expression Segmentation | RefCOCOg-test | GROUNDHOG | Overall IoU | 74.6 | #8 of 18 | Archive leaderboard | report |
| Referring Expression Segmentation | RefCOCOg-val | GROUNDHOG | Overall IoU | 74.1 | #9 of 23 | Archive leaderboard | report |
| Referring Expression Segmentation | RefCoCo val | GROUNDHOG | Overall IoU | 78.5 | #14 of 37 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections