Papers › HPT++: Hierarchically Prompting Vision-Language Models with Multi-Granularity...

HPT++: Hierarchically Prompting Vision-Language Models with Multi-Granularity Knowledge Generation and Improved Structure Modeling

27 Aug 2024arXiv:2408.14812archive 2025-07-28

Yubin Wang, Xinyang Jiang, De Cheng, Wenli Sun, Dongsheng Li, Cairong Zhao

Prompt learning has become a prevalent strategy for adapting vision-language foundation models (VLMs) such as CLIP to downstream tasks. With the emergence of large language models (LLMs), recent studies have explored the potential of using category-related descriptions to enhance prompt effectiveness. However, conventional descriptions lack explicit structured information necessary to represent the interconnections among key elements like entities or attributes with relation to a particular category. Since existing prompt tuning methods give little consideration to managing structured knowledge, this paper advocates leveraging LLMs to construct a graph for each description to prioritize such structured knowledge. Consequently, we propose a novel approach called Hierarchical Prompt Tuning (HPT), enabling simultaneous modeling of both structured and conventional linguistic knowledge. Specifically, we introduce a relationship-guided attention module to capture pair-wise associations among entities and attributes for low-level prompt learning. In addition, by incorporating high-level and global-level prompts modeling overall semantics, the proposed hierarchical structure forges cross-level interlinks and empowers the model to handle more complex and long-term relationships. Finally, by enhancing multi-granularity knowledge generation, redesigning the relationship-driven attention re-weighting module, and incorporating consistent constraints on the hierarchical text encoder, we propose HPT++, which further improves the performance of HPT. Our experiments are conducted across a wide range of evaluation settings, including base-to-new generalization, cross-dataset evaluation, and domain generalization. Extensive results and ablation studies demonstrate the effectiveness of our methods, which consistently outperform existing SOTA methods.

PaperPDFCode

In Syntology View this paper on Syntology: its repositories, every harvested function with whether it ran, its licence and the call to fetch it.

Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

ThomasWangY/2024-AAAI-HPT mentioned on GitHubpytorch report
vill-lab/2024-aaai-hpt mentioned on GitHubpytorchMIT report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Domain GeneralizationPrompt EngineeringPrompt Learning

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Prompt Engineering Caltech-101 HPT++ Harmonic mean 96.96 #2 of 14 Archive leaderboard report
Prompt Engineering DTD HPT++ Harmonic mean 74.23 #3 of 14 Archive leaderboard report
Prompt Engineering EuroSAT HPT++ Harmonic mean 87.36 #3 of 14 Archive leaderboard report
Prompt Engineering FGVC-Aircraft HPT++ Harmonic mean 41.33 #3 of 14 Archive leaderboard report
Prompt Engineering Food-101 HPT++ Harmonic mean 91.09 #9 of 13 Archive leaderboard report
Prompt Engineering ImageNet HPT++ Harmonic mean 74.24 #6 of 15 Archive leaderboard report
Prompt Engineering ImageNet V2 HPT++ Top-1 accuracy % 65.31 #1 of 8 Archive leaderboard report
Prompt Engineering ImageNet-A HPT++ Top-1 accuracy % 51.18 #3 of 9 Archive leaderboard report
Prompt Engineering ImageNet-R HPT++ Top-1 accuracy % 77.52 #4 of 9 Archive leaderboard report
Prompt Engineering ImageNet-S HPT++ Top-1 accuracy % 49.28 #5 of 9 Archive leaderboard report
Prompt Engineering Oxford 102 Flower HPT++ Harmonic mean 85.85 #8 of 14 Archive leaderboard report
Prompt Engineering Oxford-IIIT Pet Dataset HPT++ Harmonic mean 96.91 #2 of 14 Archive leaderboard report
Prompt Engineering SUN397 HPT++ Harmonic mean 81.11 #5 of 14 Archive leaderboard report
Prompt Engineering Stanford Cars HPT++ Harmonic mean 75.59 #8 of 14 Archive leaderboard report
Prompt Engineering UCF101 HPT++ Harmonic mean 83.81 #3 of 14 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

AttentionCLIPSoftmax

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections