Datasets › COCO (Common Objects in Context) › Papers, page 2

COCO (Common Objects in Context)

Papers archive 2025-07-28

papers with a benchmark row: 579 · with a code link: 504 · where Syntology ran a sample: 256 (226 with a run with no instrument failure, 30 where every run was a failure of Syntology's instrument) Syntology

Show: all papers with a benchmark rowonly where code ran (256 of 579 with a benchmark row: 226 with a run with no instrument failure, 30 where every run was a failure of Syntology's instrument)

Page 2 of 6: papers 101 to 200 of 579 with a leaderboard row on this dataset's benchmarks, newest first by the archive's date (ties by slug; undated papers last).

The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset, not that list; the archive's count for this dataset is 11,922. The Syntology column is from Syntology's graph, stated per sample; it is not part of any archive number. A line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified” (C is a part of N, never taken away from it); the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. When the archive marks a repository official for the paper, the cell starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record.

PaperCodeResultsDateSamples run Syntology
Position-guided Text Prompt for Vision-Language Pre-training 1 2 19 Dec 2022 official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified
RTMDet: An Empirical Study of Designing Real-Time Object Detectors 14 2 14 Dec 2022 community repositories only · 16 ran (of which 0 constructed an object rather than computing a result; 16 with no instrument failure: 0 honoured, 0 violated, 16 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (3 pointer-only for licence)
DeepCut: Unsupervised Segmentation using Graph Neural Networks Clustering 1 1 12 Dec 2022 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (4 pointer-only for licence)
NMS Strikes Back 1 1 12 Dec 2022 not harvested
Resolving Semantic Confusions for Improved Zero-Shot Detection 1 1 12 Dec 2022 not harvested
X-Paste: Revisiting Scalable Copy-Paste for Instance Segmentation using CLIP and StableDiffusion 2 1 7 Dec 2022 official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (3 pointer-only for licence)
DiffusionInst: Diffusion Model for Instance Segmentation 2 4 6 Dec 2022 official: harvested, nothing ran · 12 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 3 honoured, 0 violated, 4 with no contract checked; 5 where Syntology's instrument failed) · 4 unverified (7 pointer-only for licence)
Box2Mask: Box-supervised Instance Segmentation via Level-set Evolution 2 1 3 Dec 2022 not harvested
GRiT: A Generative Region-to-text Transformer for Object Understanding 1 1 1 Dec 2022 official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (8 pointer-only for licence)
SegCLIP: Patch Aggregation with Learnable Centers for Open-Vocabulary Semantic Segmentation 1 1 27 Nov 2022 not harvested
Rethinking Alignment and Uniformity in Unsupervised Semantic Segmentation 0 1 26 Nov 2022 not harvested
Shifted Diffusion for Text-to-image Generation 1 2 24 Nov 2022 official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 2 honoured, 1 violated, 7 with no contract checked; 3 where Syntology's instrument failed) · 4 unverified (6 pointer-only for licence)
DAMO-YOLO : A Report on Real-Time Object Detection Design 3 4 23 Nov 2022 official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 7 unverified (12 pointer-only for licence)
DETRs with Collaborative Hybrid Assignments Training 6 6 22 Nov 2022 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified
Retrieval-Augmented Multimodal Language Modeling 0 13 22 Nov 2022 not harvested
X²-VLM: All-In-One Pre-trained Model For Vision-Language Tasks 2 2 22 Nov 2022 official (archive's flag): 2 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (6 pointer-only for licence)
Towards All-in-one Pre-training via Maximizing Multi-modal Mutual Information 1 2 17 Nov 2022 not harvested
EVA: Exploring the Limits of Masked Visual Representation Learning at Scale 6 4 14 Nov 2022 official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified
InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions 3 11 10 Nov 2022 official: harvested, nothing ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified
OneFormer: One Transformer to Rule Universal Image Segmentation 4 6 10 Nov 2022 official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified
MogaNet: Multi-order Gated Aggregation Network 7 10 7 Nov 2022 official (archive's flag): 12 ran · 12 ran (of which 7 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified
Group DETR v2: Strong Object Detector with Encoder-Decoder Pretraining 0 1 7 Nov 2022 not harvested
Could Giant Pretrained Image Models Extract Universal Representations? 0 2 3 Nov 2022 not harvested
eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers 2 1 2 Nov 2022 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 0 violated, 9 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified
ERNIE-ViLG 2.0: Improving Text-to-Image Diffusion Model with Knowledge-Enhanced Mixture-of-Denoising-Experts 2 1 27 Oct 2022 not harvested
Dissecting Deep Metric Learning Losses for Image-Text Retrieval 2 1 21 Oct 2022 not harvested
Unsupervised Image Semantic Segmentation through Superpixels and Graph Neural Networks 0 1 21 Oct 2022 not harvested
Towards Sustainable Self-supervised Learning 1 1 20 Oct 2022 official (archive's flag): 6 ran · 6 ran (of which 6 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified; every one of the 6 samples that ran constructed an object rather than computing a result (8 pointer-only for licence)
A Tri-Layer Plugin to Improve Occluded Detection 1 1 18 Oct 2022 not harvested
Perceptual Grouping in Contrastive Vision-Language Models 2 1 18 Oct 2022 not harvested
Swinv2-Imagen: Hierarchical Vision Transformer Diffusion Models for Text-to-Image Generation 0 1 18 Oct 2022 not harvested
MOVE: Unsupervised Movable Object Segmentation and Detection 1 2 14 Oct 2022 official (archive's flag): 3 ran · 3 ran (of which 1 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 1 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (6 pointer-only for licence)
BoxTeacher: Exploring High-Quality Pseudo Labels for Weakly Supervised Instance Segmentation 1 1 11 Oct 2022 not harvested
Towards Discriminative and Transferable One-Stage Few-Shot Object Detectors 0 1 11 Oct 2022 not harvested
K-means for unsupervised instance segmentation using a self-supervised transformer 0 1 4 Oct 2022 not harvested
MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision Models 2 18 4 Oct 2022 not harvested
ERNIE-ViL 2.0: Multi-view Contrastive Learning for Image-Text Pre-training 1 2 30 Sep 2022 not harvested
Dilated Neighborhood Attention Transformer 7 2 29 Sep 2022 not harvested
Re-Imagen: Retrieval-Augmented Text-to-Image Generator 0 2 29 Sep 2022 not harvested
All are Worth Words: A ViT Backbone for Diffusion Models 3 2 25 Sep 2022 official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (3 pointer-only for licence)
OmniVL:One Foundation Model for Image-Language and Video-Language Tasks 0 1 15 Sep 2022 not harvested
Combining Metric Learning and Attention Heads For Accurate and Efficient Multilabel Image Classification 1 1 14 Sep 2022 not harvested
Exploring Target Representations for Masked Autoencoders 1 8 8 Sep 2022 official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 0 violated, 9 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (2 pointer-only for licence)
YOLOv6: A Single-Stage Object Detection Framework for Industrial Applications 7 1 7 Sep 2022 community repositories only · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (2 pointer-only for licence)
DPIT: Dual-Pipeline Integrated Transformer for Human Pose Estimation 0 1 2 Sep 2022 not harvested
gSwin: Gated MLP Vision Model with Hierarchical Structure of Shifted Window 0 3 24 Aug 2022 not harvested
Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks 2 3 22 Aug 2022 not harvested
Open Vocabulary Multi-Label Classification with Dual-Modal Decoder on Aligned Visual-Textual Features 0 3 19 Aug 2022 not harvested
Hierarchical Attention Network for Few-Shot Object Detection via Meta-Contrastive Learning 1 1 15 Aug 2022 not harvested
Exploiting Multiple Sequence Lengths in Fast End to End Training for Image Captioning 1 1 13 Aug 2022 not harvested
Analog Bits: Generating Discrete Data using Diffusion Models with Self-Conditioning 7 1 8 Aug 2022 community repositories only · 13 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 3 violated, 0 with no contract checked; 9 where Syntology's instrument failed) · 7 unverified (4 pointer-only for licence)
ALADIN: Distilling Fine-grained Alignment Scores for Efficient Image-Text Matching and Retrieval 1 1 29 Jul 2022 not harvested
HorNet: Efficient High-Order Spatial Interactions with Recursive Gated Convolutions 8 1 28 Jul 2022 official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified
kMaX-DeepLab: k-means Mask Transformer 3 4 8 Jul 2022 not harvested
Self-Constrained Inference Optimization on Structural Groups for Human Pose Estimation 0 1 6 Jul 2022 not harvested
YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors 21 10 6 Jul 2022 community repositories only · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (5 pointer-only for licence)
Boosting R-CNN: Reweighting R-CNN Samples by RPN's Error for Underwater Object Detection 2 1 28 Jun 2022 not harvested
I^2R-Net: Intra- and Inter-Human Relation Network for Multi-Person Pose Estimation 1 1 22 Jun 2022 not harvested
0/1 Deep Neural Networks via Block Coordinate Descent 0 1 19 Jun 2022 not harvested
CMT-DeepLab: Clustering Mask Transformers for Panoptic Segmentation 2 2 17 Jun 2022 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified
Deep Multi-Task Networks For Occluded Pedestrian Pose Estimation 0 1 15 Jun 2022 not harvested
GLIPv2: Unifying Localization and Vision-Language Understanding 1 1 12 Jun 2022 official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified
Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and Segmentation 10 6 6 Jun 2022 community repositories only · 11 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 8 where Syntology's instrument failed) · 2 unverified (13 pointer-only for licence)
Architecture-Agnostic Masked Image Modeling -- From ViT back to CNN 3 4 27 May 2022 not harvested
Contrastive Learning Rivals Masked Image Modeling in Fine-tuning via Feature Distillation 1 2 27 May 2022 official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 5 unverified (8 pointer-only for licence)
MixMAE: Mixed and Masked Autoencoder for Efficient Pretraining of Hierarchical Vision Transformers 1 2 26 May 2022 not harvested
Revealing the Dark Secrets of Masked Image Modeling 1 2 26 May 2022 official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified
Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding 5 1 23 May 2022 5 ran (of which 4 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified (9 pointer-only for licence)
Uniform Masking: Enabling MAE Pre-training for Pyramid-based Vision Transformers with Locality 1 1 20 May 2022 not harvested
Integrally Migrating Pre-trained Transformer Encoder-decoders for Visual Object Detection 3 2 19 May 2022 not harvested
Vision Transformer Adapter for Dense Predictions 2 11 17 May 2022 not harvested
Simple Open-Vocabulary Object Detection with Vision Transformers 2 1 12 May 2022 not harvested
AggPose: Deep Aggregation Vision Transformer for Infant Pose Estimation 1 1 11 May 2022 official (archive's flag): 5 ran · 5 ran (of which 5 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified; every one of the 5 samples that ran constructed an object rather than computing a result (7 pointer-only for licence)
CoCa: Contrastive Captioners are Image-Text Foundation Models 6 1 4 May 2022 10 ran (of which 5 constructed an object rather than computing a result; 10 with no instrument failure: 2 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 7 unverified (2 pointer-only for licence)
Lite Pose: Efficient Architecture Design for 2D Human Pose Estimation 1 1 3 May 2022 official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified
Flamingo: a Visual Language Model for Few-Shot Learning 5 1 29 Apr 2022 18 ran (of which 6 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 1 violated, 11 with no contract checked; 6 where Syntology's instrument failed) · 6 unverified (8 pointer-only for licence)
CogView2: Faster and Better Text-to-Image Generation via Hierarchical Transformers 1 2 28 Apr 2022 official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified
Understanding The Robustness in Vision Transformers 2 1 26 Apr 2022 not harvested
ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation 6 2 26 Apr 2022 community repositories only · 18 ran (of which 8 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 6 where Syntology's instrument failed) · 13 unverified (6 pointer-only for licence)
Dite-HRNet: Dynamic Lightweight High-Resolution Network for Human Pose Estimation 1 1 22 Apr 2022 not harvested
Recurrent Affine Transformation for Text-to-image Synthesis 2 1 22 Apr 2022 not harvested
CenterNet++ for Object Detection 3 1 18 Apr 2022 not harvested
YOLO-Pose: Enhancing YOLO for Multi Person Pose Estimation Using Object Keypoint Similarity Loss 2 1 14 Apr 2022 not harvested
Hierarchical Text-Conditional Image Generation with CLIP Latents 8 1 13 Apr 2022 33 ran (of which 7 constructed an object rather than computing a result; 26 with no instrument failure: 3 honoured, 5 violated, 18 with no contract checked; 7 where Syntology's instrument failed) · 5 unverified (4 pointer-only for licence)
CFA: Constraint-based Finetuning Approach for Generalized Few-Shot Object Detection 0 1 11 Apr 2022 not harvested
DaViT: Dual Attention Vision Transformers 4 2 7 Apr 2022 official (archive's flag): 8 ran · 8 ran (of which 5 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 2 where Syntology's instrument failed) · 7 unverified
KNN-Diffusion: Image Generation via Large-Scale Retrieval 0 1 6 Apr 2022 not harvested
MaxViT: Multi-Axis Vision Transformer 15 3 4 Apr 2022 official (archive's flag): 4 ran · 37 ran (of which 15 constructed an object rather than computing a result; 23 with no instrument failure: 2 honoured, 0 violated, 21 with no contract checked; 14 where Syntology's instrument failed) · 16 unverified (9 pointer-only for licence)
ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval 0 1 31 Mar 2022 not harvested
Exploring Plain Vision Transformer Backbones for Object Detection 11 4 30 Mar 2022 community repositories only · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (1 pointer-only for licence)
PP-YOLOE: An evolved version of YOLO 8 9 30 Mar 2022 official (archive's flag): 2 ran · 22 ran (of which 0 constructed an object rather than computing a result; 21 with no instrument failure: 2 honoured, 0 violated, 19 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (1 pointer-only for licence)
Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors 1 2 24 Mar 2022 4 ran (of which 2 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified
Focal Modulation Networks 9 5 22 Mar 2022 official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified
Active Token Mixer 2 1 11 Mar 2022 not harvested
E2EC: An End-to-End Contour-based Method for High-Quality High-Speed Instance Segmentation 1 1 8 Mar 2022 official (archive's flag): 8 ran · 8 ran (of which 4 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 1 violated, 4 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (10 pointer-only for licence)
DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection 16 4 7 Mar 2022 community repositories only · 10 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 1 violated, 5 with no contract checked; 4 where Syntology's instrument failed) · 5 unverified (5 pointer-only for licence)
DN-DETR: Accelerate DETR Training by Introducing Query DeNoising 17 1 2 Mar 2022 official (archive's flag): 10 ran · 17 ran (of which 3 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 1 violated, 5 with no contract checked; 11 where Syntology's instrument failed) · 5 unverified (9 pointer-only for licence)
LILE: Look In-Depth before Looking Elsewhere -- A Dual Attention Network using Transformers for Cross-Modal Information Retrieval in Histopathology Archives 0 1 2 Mar 2022 not harvested
ISDA: Position-Aware Instance Segmentation with Deformable Attention 1 2 23 Feb 2022 not harvested
Self-Supervised Transformers for Unsupervised Object Discovery using Normalized Cut 1 2 23 Feb 2022 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified