Papers › TextRegion: Text-Aligned Region Tokens from Frozen Image-Text Models
TextRegion: Text-Aligned Region Tokens from Frozen Image-Text Models
Yao Xiao, Qiqian Fu, Heyi Tao, Yuqun Wu, Zhen Zhu, Derek Hoiem
Image-text models excel at image-level tasks but struggle with detailed visual understanding. While these models provide strong visual-language alignment, segmentation models like SAM2 offer precise spatial boundaries for objects. To this end, we propose TextRegion, a simple, effective, and training-free framework that combines the strengths of image-text models and SAM2 to generate powerful text-aligned region tokens. These tokens enable detailed visual understanding while preserving open-vocabulary capabilities. They can be directly applied to various downstream tasks, including open-world semantic segmentation, referring expression comprehension, and grounding. We conduct extensive evaluations and consistently achieve superior or competitive performance compared to state-of-the-art training-free methods. Additionally, our framework is compatible with many image-text models, making it highly practical and easily extensible as stronger models emerge. Code is available at: https://github.com/avaxiao/TextRegion.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Unsupervised Semantic Segmentation with Language-image Pre-training | ADE20K | TextRegion | Mean IoU (val) | 27.3 | #2 of 13 | Archive leaderboard | report |
| Unsupervised Semantic Segmentation with Language-image Pre-training | COCO-Stuff-171 | TextRegion | mIoU | 31.2 | #2 of 12 | Archive leaderboard | report |
| Unsupervised Semantic Segmentation with Language-image Pre-training | PASCAL Context-59 | TextRegion | mIoU | 46.1 | #2 of 12 | Archive leaderboard | report |
| Unsupervised Semantic Segmentation with Language-image Pre-training | PASCAL Context-60 | TextRegion | mIoU | 41.2 | #2 of 4 | Archive leaderboard | report |
| Unsupervised Semantic Segmentation with Language-image Pre-training | PASCAL VOC | TextRegion | mIoU | 73.1 | #2 of 10 | Archive leaderboard | report |
| Unsupervised Semantic Segmentation with Language-image Pre-training | PascalVOC-20 | TextRegion | mIoU | 89.5 | #2 of 10 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections