Papers › CAILA: Concept-Aware Intra-Layer Adapters for Compositional Zero-Shot Learning

CAILA: Concept-Aware Intra-Layer Adapters for Compositional Zero-Shot Learning

26 May 2023arXiv:2305.16681archive 2025-07-28

Zhaoheng Zheng, Haidong Zhu, Ram Nevatia

In this paper, we study the problem of Compositional Zero-Shot Learning (CZSL), which is to recognize novel attribute-object combinations with pre-existing concepts. Recent researchers focus on applying large-scale Vision-Language Pre-trained (VLP) models like CLIP with strong generalization ability. However, these methods treat the pre-trained model as a black box and focus on pre- and post-CLIP operations, which do not inherently mine the semantic concept between the layers inside CLIP. We propose to dive deep into the architecture and insert adapters, a parameter-efficient technique proven to be effective among large language models, into each CLIP encoder layer. We further equip adapters with concept awareness so that concept-specific features of "object", "attribute", and "composition" can be extracted. We assess our method on four popular CZSL datasets, MIT-States, C-GQA, UT-Zappos, and VAW-CZSL, which shows state-of-the-art performance compared to existing methods on all of them.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

zhaohengz/caila officialmentioned in papermentioned on GitHubpytorch report
zhaohengz/llamp mentioned on GitHubpytorchMIT report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

AttributeCompositional Zero-Shot LearningZero-Shot Learning

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Compositional Zero-Shot Learning MIT-States, generalized split CAILA H-Mean 39.9 #1 of 2 Archive leaderboard report
Compositional Zero-Shot Learning MIT-States, generalized split CAILA Seen accuracy 51.0 #1 of 2 Archive leaderboard report
Compositional Zero-Shot Learning MIT-States, generalized split CAILA Test AUC top 1 23.4 #1 of 2 Archive leaderboard report
Compositional Zero-Shot Learning MIT-States, generalized split CAILA Test AUC top 2 - #1 of 2 Archive leaderboard report
Compositional Zero-Shot Learning MIT-States, generalized split CAILA Test AUC top 3 - #1 of 2 Archive leaderboard report
Compositional Zero-Shot Learning MIT-States, generalized split CAILA Unseen accuracy 53.9 #1 of 2 Archive leaderboard report
Compositional Zero-Shot Learning MIT-States, generalized split CAILA Val AUC top 1 - #1 of 2 Archive leaderboard report
Compositional Zero-Shot Learning MIT-States, generalized split CAILA Val AUC top 2 - #1 of 2 Archive leaderboard report
Compositional Zero-Shot Learning MIT-States, generalized split CAILA Val AUC top 3 - #1 of 2 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

CLIPFocus

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections