Datasets › PASCAL VOC

PASCAL VOC (PASCAL Visual Object Classes Challenge)

Introduced by Zheng Dong et al. in Location-aware Single Image Reflection Removal1 Jan 2010 archive 2025-07-28

The PASCAL Visual Object Classes (VOC) 2012 dataset contains 20 object categories including vehicles, household, animals, and other: aeroplane, bicycle, boat, bus, car, motorbike, train, bottle, chair, dining table, potted plant, sofa, TV/monitor, bird, cat, cow, dog, horse, sheep, and person. Each image in this dataset has pixel-level segmentation annotations, bounding box annotations, and object class annotations. This dataset has been widely used as a benchmark for object detection, semantic segmentation, and classification tasks. The PASCAL VOC dataset is split into three subsets: 1,464 images for training, 1,449 images for validation and a private testing set.

Source: Self-supervised Visual Feature Learning with Deep Neural Networks: A Survey Image Source: http://host.robots.ox.ac.uk/pascal/VOC/voc2012/examples/images/sheep_06.jpg

Benchmarks archive 2025-07-28

All 18 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Graph Matching PASCAL VOC URL F1 score 0.717±0.005 Universe Points Representation Learning for Partial... — 31 Compare
Node Classification PascalVOC-SP NeuralWalker macro F1 0.4912 ± 0.0042 Learning Long Range Dependencies on Graphs via Random Walks borgwardtlab/neuralwalker 21 Compare
Open Vocabulary Semantic Segmentation PascalVOC-20 UMG-CLIP-L/14 mIoU 97.9 UMG-CLIP: A Unified Multi-Granularity Vision Generalist... lygsbw/umg-clip 20 Compare
Zero-Shot Semantic Segmentation PASCAL VOC OTSeg+ Transductive Setting hIoU 94.4 OTSeg: Multi-prompt Sinkhorn Attention for Zero-Shot... cubeyoung/OTSeg 13 Compare
Unsupervised Semantic Segmentation with Language-image Pre-training PASCAL VOC CorrCLIP mIoU 76.7 CorrCLIP: Reconstructing Correlations in CLIP with... zdk258/CorrCLIP 10 Compare
Unsupervised Semantic Segmentation with Language-image Pre-training PascalVOC-20 CorrCLIP mIoU 91.8 CorrCLIP: Reconstructing Correlations in CLIP with... zdk258/CorrCLIP 10 Compare
Open Vocabulary Semantic Segmentation PascalVOC-20b UMG-CLIP-E/14 mIoU 85.4 UMG-CLIP: A Unified Multi-Granularity Vision Generalist... lygsbw/umg-clip 4 Compare
Image Segmentation PASCAL VOC OneNete,4-C mIoU 63.6 OneNet: A Channel-Wise 1D Convolutional U-Net shbyun080/onenet 3 Compare
Object Counting PASCAL VOC TFOC mRMSE 0.0084 Training-free Object Counting with Prompts shizenglin/training-free-object-counter 3 Compare
Single-object discovery VOC_all Large-scale rOSD CorLoc 49.4 Toward unsupervised, multi-object discovery in... huyvvo/rOSD 3 Compare
Interactive Segmentation PASCAL VOC ICL CFR-1 (ViT-H, C+L) NoC@95 2.45 CFR-ICL: Cascade-Forward Refinement with Iterative Click... TitorX/CFR-ICL-Interactive-Segmentation 2 Compare
Knowledge Distillation PASCAL VOC LSHFM (T: ResNet101 S: ResNet50) mAP 93.17 Distilling Knowledge by Mimicking Features DoctorKey/LSHFM.singleclassification +2 2 Compare
Multi-object discovery VOC_all Large-scale rOSD Detection Rate 38.3 Toward unsupervised, multi-object discovery in... huyvvo/rOSD 2 Compare
Object Detection PASCAL VOC 10% DETReg (MDef-DETR) AP 58.78 Class-agnostic Object Detection with Multi-modal Transformer mmaaz60/mvits_for_class_agnostic_od 2 Compare
Single-object discovery VOC_6x2 rOSD CorLoc 72.5 Toward unsupervised, multi-object discovery in... huyvvo/rOSD 2 Compare
Multi-object colocalization VOC_all rOSD Detection Rate 49.4 Toward unsupervised, multi-object discovery in... huyvvo/rOSD 1 Compare
Object Detection PASCAL VOC TinyissimoYOLO-v8 Parameters(K) 839 Ultra-Efficient On-Device Object Detection on... eth-pbl/tinyissimoyolo 1 Compare
Semantic Segmentation PASCAL VOC SegCLIP mIoU 52.6 SegCLIP: Patch Aggregation with Learnable Centers for... arrowluo/segclip 1 Compare

Papers archive 2025-07-28

30 shown of 83 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 198. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
TextRegion: Text-Aligned Region Tokens from Frozen Image-Text Models 1 2 29 May 2025 not harvested
Improving the Effective Receptive Field of Message-Passing Neural Networks 1 2 29 May 2025 ran 1 of 1 samples (0 unverified)
Unlocking the Potential of Classic GNNs for Graph-level Tasks: Simple Architectures Meet Excellence 1 1 13 Feb 2025 ran 4 of 12 samples (8 unverified)
MaskCLIP++: A Mask-Based CLIP Fine-tuning Framework for Open-Vocabulary Image Segmentation 1 1 16 Dec 2024 not harvested
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training 1 1 2 Dec 2024 not harvested
HyperSeg: Towards Universal Visual Segmentation with Large Language Model 1 1 26 Nov 2024 ran 7 of 17 samples (10 unverified)
CorrCLIP: Reconstructing Correlations in CLIP with Off-the-Shelf Foundation Models for Open-Vocabulary Semantic Segmentation 1 2 15 Nov 2024 ran 0 of 9 samples (9 unverified; 9 pointer-only for licence)
OneNet: A Channel-Wise 1D Convolutional U-Net 1 3 14 Nov 2024 not harvested
Harnessing Vision Foundation Models for High-Performance, Training-Free Open Vocabulary Segmentation 1 2 14 Nov 2024 not harvested
ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary Segmentation 1 2 9 Aug 2024 ran 3 of 9 samples (6 unverified; 9 pointer-only for licence)
In Defense of Lazy Visual Grounding for Open-Vocabulary Semantic Segmentation 1 1 9 Aug 2024 not harvested
Collaborative Vision-Text Representation Optimizing for Open-Vocabulary Segmentation 1 1 1 Aug 2024 not harvested
Next Level Message-Passing with Hierarchical Support Graphs 1 1 22 Jun 2024 ran 4 of 10 samples (6 unverified)
Open-Vocabulary Semantic Segmentation with Image Embedding Balancing 1 1 14 Jun 2024 ran 3 of 12 samples (9 unverified)
Learning Long Range Dependencies on Graphs via Random Walks 1 1 5 Jun 2024 ran 13 of 13 samples (0 unverified)
Learning Latent Partial Matchings with Gumbel-IPF Networks 1 1 3 Apr 2024 not harvested
TTD: Text-Tag Self-Distillation Enhancing Image-Text Alignment in CLIP to Alleviate Single Tag Bias 1 2 30 Mar 2024 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
Cross-modal Retrieval with Noisy Correspondence via Consistency Refining and Mining 1 1 25 Mar 2024 not harvested
OTSeg: Multi-prompt Sinkhorn Attention for Zero-Shot Semantic Segmentation 1 2 21 Mar 2024 ran 9 of 10 samples (1 unverified; 10 pointer-only for licence)
UMG-CLIP: A Unified Multi-Granularity Vision Generalist for Open-World Understanding 1 2 12 Jan 2024 not harvested
Exploring Regional Clues in CLIP for Zero-Shot Semantic Segmentation 1 1 1 Jan 2024 not harvested
TagAlign: Improving Vision-Language Alignment with Multi-Tag Classification 1 3 21 Dec 2023 not harvested
TagCLIP: A Local-to-Global Framework to Enhance Open-Vocabulary Multi-Label Classification of CLIP Without Training 1 1 20 Dec 2023 not harvested
Open-Vocabulary Segmentation with Semantic-Assisted Calibration 2 1 7 Dec 2023 ran 5 of 10 samples (5 unverified; 10 pointer-only for licence)
GMTR: Graph Matching Transformers 1 2 14 Nov 2023 not harvested
Ultra-Efficient On-Device Object Detection on AI-Integrated Smart Glasses with TinyissimoYOLO 1 1 2 Nov 2023 not harvested
SILC: Improving Vision Language Pretraining with Self-Distillation 0 2 20 Oct 2023 not harvested
Learning Mask-aware CLIP Representations for Zero-Shot Segmentation 2 1 30 Sep 2023 ran 1 of 7 samples (6 unverified; 2 pointer-only for licence)
Where Did the Gap Go? Reassessing the Long-Range Graph Benchmark 2 4 1 Sep 2023 ran 5 of 12 samples (7 unverified)
Convolutions Die Hard: Open-Vocabulary Segmentation with Single Frozen Convolutional CLIP 1 2 4 Aug 2023 ran 1 of 2 samples (1 unverified)

The full list of 83 is in the JSON twin.

Dataset loaders archive 2025-07-28

5 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Custom

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • Pascal VOC 2007 count-test
  • VOCASET
  • VOC_all
  • VOC_6x2
  • PascalVOC-SP
  • PascalVOC-59
  • PascalVOC-459
  • PascalVOC-20b
  • PascalVOC-20
  • PASCAL VOC 10%
  • PASCAL VOC

11 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections