Datasets › CUB-200-2011

CUB-200-2011 (Caltech-UCSD Birds-200-2011)

Introduced in The Caltech-UCSD Birds-200-2011 Dataset archive 2025-07-28

The Caltech-UCSD Birds-200-2011 (CUB-200-2011) dataset is the most widely-used dataset for fine-grained visual categorization task. It contains 11,788 images of 200 subcategories belonging to birds, 5,994 for training and 5,794 for testing. Each image has detailed annotations: 1 subcategory label, 15 part locations, 312 binary attributes and 1 bounding box. The textual information comes from Reed et al.. They expand the CUB-200-2011 dataset by collecting fine-grained natural language descriptions. Ten single-sentence descriptions are collected for each image. The natural language descriptions are collected through the Amazon Mechanical Turk (AMT) platform, and are required at least 10 words, without any information of subcategories and actions.

Source: Fine-grained Visual-textual Representation Learning Image Source: http://www.vision.caltech.edu/visipedia/CUB-200-2011.html

Benchmarks archive 2025-07-28

All 47 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Few-Shot Image Classification CUB 200 5-way 1-shot PT+MAP+SF+SOT (transductive) Accuracy 95.80 The Self-Optimal-Transport Feature Transform danielshalam/bpa 36 Compare
Few-Shot Image Classification CUB 200 5-way 5-shot CAML [Laion-2b] Accuracy 98.7 Context-Aware Meta-Learning cfifty/CAML 32 Compare
Fine-Grained Image Classification CUB-200-2011 HERBS Accuracy 93.1% Fine-grained Visual Classification with High-temperature... chou141253/FGVC-HERBS 30 Compare
Metric Learning CUB-200-2011 Unicom+ViT-L@336px R@1 90.1 Unicom: Universal and Compact Representation Learning... OML-Team/open-metric-learning +2 30 Compare
Text-to-Image Generation CUB RAT-Diffusion FID 6.36 Data Extrapolation for Text-to-image Generation on Small Datasets senmaoy/RAT-Diffusion 20 Compare
Zero-Shot Learning CUB-200-2011 ZeroDiff average top-1 classification accuracy 87.5 Exploring Data Efficiency in Zero-Shot Learning with... — 14 Compare
Fine-Grained Image Classification CUB-200-2011 TransFG Accuracy 91.7% TransFG: A Transformer Architecture for Fine-grained Recognition TACJu/TransFG +1 12 Compare
Weakly-Supervised Object Localization CUB-200-2011 DiPS MaxBoxAccV2 90.9 Discriminative Sampling of Proposals in Self-Supervised... — 10 Compare
Cross-Domain Few-Shot CUB MSENet 5 shot 71.59 Enhancing Few-Shot Image Classification through... FatemehAskari/MSENet 9 Compare
Image Attribution CUB-200-2011 SMDL-Attribution (ICLR version) Insertion AUC score (ResNet-101) 0.7262 Less is More: Fewer Interpretable Region via Submodular... ruoyuchen10/smdl-attribution 8 Compare
Image Retrieval CUB-200-2011 CGD (MG/SG) R@1 79.2 Combination of Multiple Global Descriptors for Image Retrieval naver/cgd +6 8 Compare
Point-interactive Image Colorization CUB-200-2011 iColoriT PSNR@10 30.595 iColoriT: Towards Propagating Local Hint to the Right... pmh9960/iColoriT 7 Compare
Few-Shot Class-Incremental Learning CUB-200-2011 CoACT Last Accuracy 81.19 Few-shot Tuning of Foundation Models for... shuvenduroy/coact-fscil 6 Compare
Few-Shot Image Classification CUB-200-2011 - 0-Shot Word CNN-RNN (DS-SJE Embedding) Top-1 Accuracy 56.8% Learning Deep Representations of Fine-grained Visual Descriptions hanzhanggit/StackGAN-v2 +8 5 Compare
Generalized Few-Shot Learning CUB MVCN Per-Class Accuracy (1-shot) 57.3 Better Generalized Few-Shot Learning Even Without Base Data bigdata-inha/zero-base-gfsl 5 Compare
Long-tail learning with class descriptors CUB-LT DRAGON + Bal'Loss Per-Class Accuracy 60.1 From Generalized zero-shot learning to long-tail with... dvirsamuel/DRAGON 5 Compare
Error Understanding CUB-200-2011 SMDL-Attribution (ICLR version) Average highest confidence (ResNet-101) 0.4513 Less is More: Fewer Interpretable Region via Submodular... ruoyuchen10/smdl-attribution 4 Compare
Few-Shot Image Classification CUB 200 50-way (0-shot) Prototypical Networks Accuracy 54.6 Prototypical Networks for Few-shot Learning learnables/learn2learn +42 4 Compare
Graph Matching CUB URL F1 score 0.951 Universe Points Representation Learning for Partial... — 4 Compare
Image Classification CUB Entropy-based Logic Explained Network Classification Accuracy 0.9295 Entropy-based Logic Explanations of Neural Networks pietrobarbiero/pytorch_explain +2 4 Compare
Image Clustering CUB Birds FineGAN Accuracy 0.126 FineGAN: Unsupervised Hierarchical Disentanglement for... kkanshul/finegan 4 Compare
Image Generation CUB 128 x 128 Projected GAN FID 2.79 Projected GANs Converge Faster autonomousvision/projected_gan +2 4 Compare
Small Data Image Classification CUB-200-2011, 30 samples per class GLICO Accuracy 77.75 Generative Latent Implicit Conditional Optimization when... IdanAzuri/glico-learning-small-sample 4 Compare
Few-Shot Image Classification CUB-200 - 0-Shot Learning TAFE-Net Accuracy 56.9% TAFE-Net: Task-Aware Feature Embeddings for Low Shot Learning ucbdrive/tafe-net 3 Compare
Generalized Zero-Shot Learning CUB-200-2011 ZeroDiff Harmonic mean 81.6 Exploring Data Efficiency in Zero-Shot Learning with... — 3 Compare
Weakly-Supervised Object Localization CUB-200-2011 FALcon GT-known localization accuracy 88.30 Exploring Foveation and Saccade for Improved... TimurIbrayev/FALcon 3 Compare
Concept-based Classification CUB-200-2011 CGEM (ResNet-34) Task Accuracy (%) 79.68 Concept Graph Embedding Models for Enhanced Accuracy and... jumpsnack/cgem 2 Compare
Fine-Grained Image Recognition CUB-200-2011 PIM Accuracy 92.8 A Novel Plug-in Module for Fine-Grained Visual Classification chou141253/fgvc-pim 2 Compare
Fine-Grained Image Recognition CUB Birds HOI-Net 1:1 Accuracy 90.02% High-Order-Interaction for weakly supervised... puallee/HOI-Net 2 Compare
Image Classification CUB-200-2011 Sparse-CBM Accuracy 80.02 Sparse Concept Bottleneck Models: Gumbel Tricks in... andron00e/sparsecbm +1 2 Compare
Image Classification Imbalanced CUB-200-2011 Multi-task Accuracy 99.67 A New Periocular Dataset Collected by Mobile Devices in... — 2 Compare
Interpretable Machine Learning CUB-200-2011 Q-SENN Top 1 Accuracy 85.9 Q-SENN: Quantized Self-Explaining Neural Networks thomasnorr/q-senn 2 Compare
Metric Learning CUB-200-2011 Hyp-DINO R@1 80.9 Hyperbolic Vision Transformers: Combining Improvements... OML-Team/open-metric-learning +1 2 Compare
Document Text Classification CUB-200-2011 Bert Accuracy 65.0 Are These Birds Similar: Learning Branched Networks for... nicolalandro/ntsnet-cub200 +2 1 Compare
Few-Shot Image Classification CUB-200-2011 5-way (1-shot) MATANet Accuracy 67.33 Multi-scale Adaptive Task Attention Network for Few-Shot Learning — 1 Compare
Few-Shot Image Classification CUB-200-2011 5-way (5-shot) MATANet Accuracy 83.92 Multi-scale Adaptive Task Attention Network for Few-Shot Learning — 1 Compare
Fine-Grained Image Classification Imbalanced CUB-200-2011 PC-Softmax Accuracy 89.73 Rethinking Softmax with Cross-Entropy: Neural Network... ZhenyueQin/Research-Softmax-with-Mutual-Information 1 Compare
Fine-Grained Visual Recognition CUB-200-2011 Selfsynthx Accuracy (%) 85.02 Enhancing Cognition and Explainability of Multimodal... sycny/selfsynthx 1 Compare
Image Clustering CUB-200-2011 MES-Loss NMI 73.35 MES-Loss: Mutually equidistant separation metric... — 1 Compare
Multimodal Deep Learning CUB-200-2011 Two Branch Network (Text - Bert + Image - Nts-Net) Accuracy 96.81 Are These Birds Similar: Learning Branched Networks for... nicolalandro/ntsnet-cub200 +2 1 Compare
Multimodal Text and Image Classification CUB-200-2011 Two Branch Network (Text - Bert + Image - Nts-Net) Accuracy 96.81 Are These Birds Similar: Learning Branched Networks for... nicolalandro/ntsnet-cub200 +2 1 Compare
Semantic correspondence CUB-200-2011 LDM Correspondences Mean PCK@0.05 61.6 Unsupervised Semantic Correspondence Using Stable Diffusion ubc-vision/LDM_correspondences 1 Compare
Single-View 3D Reconstruction CUB-200-2011 3D Magic Mirror FID 63.5 3D Magic Mirror: Clothing Reconstruction from a Single... layumi/3D-Magic-Mirror 1 Compare
Small Data Image Classification CUB-200-2011, 5 samples per class GLICO Accuracy 51.52 Generative Latent Implicit Conditional Optimization when... IdanAzuri/glico-learning-small-sample 1 Compare
Transductive Zero-Shot Classification CUB-200-2011 ZLaP Accuracy 64.1 Label Propagation for Zero-shot Classification with... vladan-stojnic/zlap 1 Compare
Weakly-Supervised Object Localization CUB TokenCut Top-1 Localization Accuracy 72.9 Self-Supervised Transformers for Unsupervised Object... YangtaoWANG95/TokenCut 1 Compare
Zero-Shot Learning CUB-200 - 0-Shot Learning zsl_ADA Average Per-Class Accuracy 70.9 A Generative Framework for Zero-Shot Learning with... vkkhare/ZSL-ADA 1 Compare

Papers archive 2025-07-28

30 shown of 211 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 2,235. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Structural feature enhanced transformer for fine-grained image recognition 0 1 14 Jun 2025 not harvested
Multi-scale Activation, Refinement, and Aggregation: Exploring Diverse Cues for Fine-Grained Bird Recognition 0 1 12 Apr 2025 not harvested
Enhancing Cognition and Explainability of Multimodal Foundation Models with Self-Synthesized Data 1 1 19 Feb 2025 ran 2 of 2 samples (0 unverified; 1 pointer-only for licence)
Interweaving Insights: High-Order Feature Interaction for Fine-Grained Visual Recognition 1 1 20 Oct 2024 not harvested
Data Extrapolation for Text-to-image Generation on Small Datasets 1 1 2 Oct 2024 not harvested
EQ-CBM: A Probabilistic Concept Bottleneck with Energy-based Models and Quantized Vectors 0 1 22 Sep 2024 not harvested
Enhancing Few-Shot Image Classification through Learnable Multi-Scale Embedding and Attention Mechanisms 1 1 12 Sep 2024 not harvested
Concept Graph Embedding Models for Enhanced Accuracy and Interpretability 1 1 13 Aug 2024 not harvested
The Balanced-Pairwise-Affinities Feature Transform 1 2 25 Jun 2024 not harvested
Exploring Data Efficiency in Zero-Shot Learning with Diffusion Models 0 2 5 Jun 2024 not harvested
Few-shot Tuning of Foundation Models for Class-incremental Learning 1 1 26 May 2024 ran 6 of 10 samples (4 unverified)
Label Propagation for Zero-shot Classification with Vision-Language Models 1 3 5 Apr 2024 ran 1 of 2 samples (1 unverified)
Sparse Concept Bottleneck Models: Gumbel Tricks in Contrastive Learning 2 1 4 Apr 2024 not harvested
Learning Latent Partial Matchings with Gumbel-IPF Networks 1 1 3 Apr 2024 not harvested
Pre-trained Vision and Language Transformers Are Few-Shot Incremental Learners 1 2 2 Apr 2024 ran 12 of 14 samples (2 unverified)
A Bag of Tricks for Few-Shot Class-Incremental Learning 0 1 21 Mar 2024 not harvested
Less is More: Fewer Interpretable Region via Submodular Subset Selection 1 2 14 Feb 2024 ran 3 of 3 samples (0 unverified; 3 pointer-only for licence)
Learning Semantic Proxies from Visual Prompts for Parameter-Efficient Fine-Tuning in Deep Metric Learning 1 1 4 Feb 2024 ran 3 of 5 samples (2 unverified)
Zero-shot Classification using Hyperdimensional Computing 0 2 30 Jan 2024 not harvested
Q-SENN: Quantized Self-Explaining Neural Networks 1 1 21 Dec 2023 not harvested
Exploring Foveation and Saccade for Improved Weakly-Supervised Localization 1 1 16 Dec 2023 not harvested
Context-Aware Meta-Learning 1 2 17 Oct 2023 ran 5 of 16 samples (11 unverified)
PCNN: Probable-Class Nearest-Neighbor Explanations Improve Fine-Grained Image Classification Accuracy for AIs and Humans 2 2 25 Aug 2023 not harvested
MES-Loss: Mutually equidistant separation metric learning loss function 0 2 1 Aug 2023 not harvested
Center Contrastive Loss for Metric Learning 0 1 1 Aug 2023 not harvested
Generative Prompt Model for Weakly Supervised Object Localization 1 2 19 Jul 2023 ran 1 of 1 samples (0 unverified)
Unsupervised Semantic Correspondence Using Stable Diffusion 1 1 24 May 2023 ran 4 of 12 samples (8 unverified)
Feature Channel Adaptive Enhancement for Fine-Grained Visual Classification 0 1 11 May 2023 not harvested
ESPT: A Self-Supervised Episodic Spatial Pretext Task for Improving Few-Shot Learning 1 2 26 Apr 2023 ran 0 of 1 samples (1 unverified; 1 pointer-only for licence)
Unicom: Universal and Compact Representation Learning for Image Retrieval 3 1 12 Apr 2023 ran 3 of 6 samples (3 unverified; 6 pointer-only for licence)

The full list of 211 is in the JSON twin.

Dataset loaders archive 2025-07-28

5 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Unknown

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • CUB-200-2011
  • CUB-200-2011, 10 samples per class
  • CUB-200-2011 5-way (5-shot)
  • CUB-200-2011 5-way (1-shot)
  • Imbalanced CUB-200-2011
  • CUB-200-2011, 5 samples per class
  • CUB-200-2011, 30 samples per class
  • CUB-LT
  • CUB-200-2011 - 0-Shot
  • CUB-200 - 0-Shot Learning
  • CUB Birds
  • CUB 200 50-way (0-shot)
  • CUB 200 5-way 5-shot
  • CUB 200 5-way 1-shot
  • CUB 128 x 128
  • CUB
  • CUB-200-2011

17 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections