Datasets › MSCOCO

MSCOCO

archive 2025-07-28

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

Provide:

  • a high-level explanation of the dataset characteristics
  • explain motivations and summary of its content
  • potential use cases of the dataset

Benchmarks archive 2025-07-28

All 13 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Open Vocabulary Object Detection MSCOCO Cooperative Foundational Models AP 0.5 50.3 Enhancing Novel Object Detection via Cooperative... rohit901/cooperative-foundational-models 32 Compare
Real-time Instance Segmentation MSCOCO RTMDet-Ins-x mask AP 44.6 RTMDet: An Empirical Study of Designing Real-Time Object... open-mmlab/mmdetection +13 22 Compare
Object Detection MSCOCO PP-PicoDet-L mAP @0.5:0.95 40.9 PP-PicoDet: A Better Real-Time Object Detector on Mobile Devices PaddlePaddle/PaddleDetection +3 7 Compare
Zero-Shot Object Detection MSCOCO Grounding DINO 1.6 Pro (without COCO data) AP 55.4 Grounding DINO 1.5: Advance the "Edge" of Open-Set... mit-han-lab/efficientvit +2 7 Compare
Multi-Label Image Classification MSCOCO IDA-R101(H) mAP 84.8 Causality Compensated Attention for Contextual Biased... yu-gi-oh-leilei/IDA_2023ICLR 4 Compare
Image Retrieval MSCOCO HADA Recall@1 58.46 HADA: A Graph-based Amalgamation Framework in Image-text... m2man/hada +1 3 Compare
Cross-Modal Retrieval MSCOCO 3SHNet Image-to-text R@1 85.8 3SHNet: Boosting Image-Sentence Retrieval via Visual... xurige1995/3shnet 1 Compare
Few Shot Open Set Object Detection MSCOCO FOODv2 AR_U 16.52 HSIC-based Moving WeightAveraging for Few-Shot Open-Set... binyisu/food 1 Compare
Image Captioning MSCOCO CapDec BLEU-4 26.4 Text-Only Training for Image Captioning using Noise-Injected CLIP davidhuji/capdec +3 1 Compare
Image Outpainting MSCOCO NUWA-3D CLIP Similarity 32.26 Learning 3D Photography Videos via Self-supervised... — 1 Compare
Object Detection MSCOCO DAS Average mAP 39.7 DAS: A Deformable Attention to Capture Salient... — 1 Compare
Paraphrase Generation MSCOCO HRQ-VAE BLEU 27.90 Hierarchical Sketch Induction for Paraphrase Generation tomhosking/hrq-vae 1 Compare
Weakly Supervised Object Detection MSCOCO CASD(ResNet50) mAP 13.9 Comprehensive Attention Self-Distillation for... DeLightCMU/CASD 1 Compare

Papers archive 2025-07-28

30 shown of 52 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 90. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
CP-DETR: Concept Prompt Guide DETR Toward Stronger Universal Object Detection 0 1 13 Dec 2024 not harvested
SIA-OVD: Shape-Invariant Adapter for Bridging the Image-Region Gap in Open-Vocabulary Detection 1 2 8 Oct 2024 not harvested
OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective Fusion 1 1 10 Jul 2024 not harvested
OV-DQUO: Open-Vocabulary DETR with Denoising Text Query Training and Open-World Unknown Objects Supervision 1 2 28 May 2024 ran 9 of 11 samples (2 unverified)
Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection 3 2 16 May 2024 ran 1 of 2 samples (1 unverified)
3SHNet: Boosting Image-Sentence Retrieval via Visual Semantic-Spatial Self-Highlighting 1 1 26 Apr 2024 ran 10 of 10 samples (0 unverified)
Retrieval-Augmented Open-Vocabulary Object Detection 1 1 8 Apr 2024 not harvested
YOLOv8-AM: YOLOv8 Based on Effective Attention Mechanisms for Pediatric Wrist Fracture Detection 1 1 14 Feb 2024 not harvested
YOLO-World: Real-Time Open-Vocabulary Object Detection 3 1 30 Jan 2024 ran 0 of 3 samples (3 unverified)
CLIM: Contrastive Language-Image Mosaic for Region Representation 1 1 18 Dec 2023 ran 1 of 3 samples (2 unverified; 3 pointer-only for licence)
DAS: A Deformable Attention to Capture Salient Information in CNNs 0 1 20 Nov 2023 not harvested
Enhancing Novel Object Detection via Cooperative Foundational Models 1 1 19 Nov 2023 not harvested
YOLOv8-Based Visual Detection of Road Hazards: Potholes, Sewer Covers, and Manholes 0 1 31 Oct 2023 not harvested
HSIC-based Moving WeightAveraging for Few-Shot Open-Set Object Detection 1 1 27 Oct 2023 not harvested
LP-OVOD: Open-Vocabulary Object Detection by Linear Probing 1 2 26 Oct 2023 ran 2 of 2 samples (0 unverified; 2 pointer-only for licence)
CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction 1 1 2 Oct 2023 ran 2 of 6 samples (4 unverified; 6 pointer-only for licence)
Region-centric Image-Language Pretraining for Open-Vocabulary Detection 2 1 29 Sep 2023 not harvested
Detect Everything with Few Examples 1 1 22 Sep 2023 ran 15 of 20 samples (5 unverified)
Contrastive Feature Masking Open-Vocabulary Vision Transformer 0 1 2 Sep 2023 not harvested
ScaleDet: A Scalable Multi-Dataset Object Detector 0 1 8 Jun 2023 not harvested
CORA: Adapting CLIP for Open-Vocabulary Detection with Region Prompting and Anchor Pre-Matching 1 2 23 Mar 2023 ran 2 of 3 samples (1 unverified)
Object-Aware Distillation Pyramid for Open-Vocabulary Object Detection 1 2 10 Mar 2023 not harvested
Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection 10 1 9 Mar 2023 ran 2 of 5 samples (3 unverified)
Aligning Bag of Regions for Open-Vocabulary Object Detection 1 1 27 Feb 2023 not harvested
Causality Compensated Attention for Contextual Biased Visual Recognition 1 3 25 Feb 2023 not harvested
Learning 3D Photography Videos via Self-supervised Diffusion on Single Images 0 1 21 Feb 2023 not harvested
HADA: A Graph-based Amalgamation Framework in Image-text Retrieval 2 3 11 Jan 2023 not harvested
RTMDet: An Empirical Study of Designing Real-Time Object Detectors 14 4 14 Dec 2022 ran 3 of 20 samples (17 unverified)
Open-vocabulary Attribute Detection 1 1 23 Nov 2022 ran 3 of 10 samples (7 unverified)
Text-Only Training for Image Captioning using Noise-Injected CLIP 4 1 1 Nov 2022 ran 4 of 5 samples (1 unverified; 1 pointer-only for licence)

The full list of 52 is in the JSON twin.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

No variants listed.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections