| Object Detection |
COCO test-dev |
Co-DETR box mAP 66.0 |
DETRs with Collaborative Hybrid Assignments Training |
open-mmlab/mmdetection +5 |
225 |
Compare |
| Object Detection |
COCO minival |
PE_spatial (DETA) box AP 66.0 |
Perception Encoder: The best visual embeddings are not... |
facebookresearch/perception_models |
220 |
Compare |
| Instance Segmentation |
COCO test-dev |
Co-DETR mask AP 57.1 |
DETRs with Collaborative Hybrid Assignments Training |
open-mmlab/mmdetection +5 |
112 |
Compare |
| Instance Segmentation |
COCO minival |
Co-DETR mask AP 56.6 |
DETRs with Collaborative Hybrid Assignments Training |
open-mmlab/mmdetection +5 |
93 |
Compare |
| Real-Time Object Detection |
COCO (Common Objects in Context) |
DEIM-D-FINE-X+ box AP 59.5 |
DEIM: DETR with Improved Matching for Fast Convergence |
shihuahuang95/deim |
82 |
Compare |
| Text-to-Image Generation |
COCO (Common Objects in Context) |
RAT-Diffusion FID 5.00 |
Data Extrapolation for Text-to-image Generation on Small Datasets |
senmaoy/RAT-Diffusion |
69 |
Compare |
| Pose Estimation |
COCO test-dev |
ViTPose (ViTAE-G, ensemble) AP 81.1 |
ViTPose: Simple Vision Transformer Baselines for Human... |
huggingface/transformers +5 |
47 |
Compare |
| Panoptic Segmentation |
COCO test-dev |
Mask DINO (single scale) PQ 59.5 |
Mask DINO: Towards A Unified Transformer-based Framework... |
PaddlePaddle/PaddleDetection +9 |
38 |
Compare |
| Cross-Modal Retrieval |
COCO 2014 |
VAST Text-to-image R@1 68.0 |
VAST: A Vision-Audio-Subtitle-Text Omni-Modality... |
TXH-mercury/VALOR +1 |
36 |
Compare |
| Multi-Label Classification |
MS-COCO |
ADDS(ViT-L-336, resolution 1344) mAP 93.54 |
Open Vocabulary Multi-Label Classification with... |
— |
34 |
Compare |
| Few-Shot Object Detection |
MS-COCO (10-shot) |
Training-free AP 36.6 |
No time to train! Training-Free Reference-Based Instance... |
miquel-espinosa/no-time-to-train +1 |
33 |
Compare |
| Panoptic Segmentation |
COCO minival |
HyperSeg (Swin-B) PQ 61.2 |
HyperSeg: Towards Universal Visual Segmentation with... |
congvvc/HyperSeg |
31 |
Compare |
| Keypoint Detection |
COCO (Common Objects in Context) |
4xRSN-50(384×288) Test AP 78.6 |
Learning Delicate Local Representations for Multi-Person... |
open-mmlab/mmpose +3 |
24 |
Compare |
| Object Detection |
COCO 2017 |
MaxViT-B AP 53.4 |
MaxViT: Multi-Axis Vision Transformer |
huggingface/pytorch-image-models +14 |
24 |
Compare |
| Zero-Shot Cross-Modal Retrieval |
COCO 2014 |
InternVL-G Image-to-text R@1 74.9 |
InternVL: Scaling up Vision Foundation Models and... |
opengvlab/internvl +1 |
18 |
Compare |
| Image Captioning |
COCO (Common Objects in Context) |
ExpansionNet v2 CIDEr 143.7 |
Exploiting Multiple Sequence Lengths in Fast End to End... |
jchenghu/expansionnet_v2 |
17 |
Compare |
| Keypoint Detection |
COCO test-dev |
HRNet* APM 73.4 |
Deep High-Resolution Representation Learning for Human... |
open-mmlab/mmdetection +38 |
16 |
Compare |
| Multi-Person Pose Estimation |
COCO (Common Objects in Context) |
RSN AP 0.792 |
Learning Delicate Local Representations for Multi-Person... |
open-mmlab/mmpose +3 |
15 |
Compare |
| Multi-Person Pose Estimation |
COCO test-dev |
SCIO (HRNet-48) AP 79.2 |
Self-Constrained Inference Optimization on Structural... |
— |
15 |
Compare |
| Visual Question Answering (VQA) |
COCO Visual Question Answering (VQA) real images 1.0 open ended |
MCB 7 att. Percentage correct 66.5 |
Multimodal Compact Bilinear Pooling for Visual Question... |
Cadene/vqa.pytorch +9 |
14 |
Compare |
| Pose Estimation |
COCO (Common Objects in Context) |
OmniPose (WASPv2) AP 79.5 |
OmniPose: A Multi-Scale Framework for Multi-Person Pose... |
bmartacho/OmniPose |
10 |
Compare |
| Single-object discovery |
COCO_20k |
IMST CorLoc 72.2 |
K-means for unsupervised instance segmentation using a... |
— |
10 |
Compare |
| Visual Question Answering (VQA) |
COCO Visual Question Answering (VQA) real images 1.0 multiple choice |
MCB 7 att. Percentage correct 70.1 |
Multimodal Compact Bilinear Pooling for Visual Question... |
Cadene/vqa.pytorch +9 |
10 |
Compare |
| Zero-Shot Composed Image Retrieval (ZS-CIR) |
COCO (Common Objects in Context) |
iSEARLE-XL-OTI (CLIP L/14) Actions Recall@5 32.55 |
iSEARLE: Improving Textual Inversion for Zero-Shot... |
miccunifi/searle +1 |
10 |
Compare |
| Image-to-Text Retrieval |
COCO (Common Objects in Context) |
BLIP-2 (ViT-G, fine-tuned) Recall@1 85.4 |
BLIP-2: Bootstrapping Language-Image Pre-training with... |
huggingface/transformers +16 |
9 |
Compare |
| Semantic Segmentation |
COCO (Common Objects in Context) |
HyperSeg mIoU 77.2 |
HyperSeg: Towards Universal Visual Segmentation with... |
congvvc/HyperSeg |
9 |
Compare |
| Zero-Shot Object Detection |
MS-COCO |
UniFa mAP 26.00 |
UniFa: A unified feature hallucination framework for... |
— |
9 |
Compare |
| Keypoint Detection |
COCO test-challenge |
4×RSN-50 AR 82.6 |
Learning Delicate Local Representations for Multi-Person... |
open-mmlab/mmpose +3 |
8 |
Compare |
| Box-supervised Instance Segmentation |
COCO test-dev |
Box2Mask-T mask AP 42.4 |
Box2Mask: Box-supervised Instance Segmentation via... |
LiWentomng/BoxInstSeg +1 |
7 |
Compare |
| Image-level Supervised Instance Segmentation |
COCO test-dev |
WeakSAM-Mask2Former (with SAM) AP 25.9 |
WeakSAM: Segment Anything Meets Weakly-supervised... |
hustvl/weaksam |
7 |
Compare |
| Object Counting |
COCO count-test |
ens m-reIRMSE 0.18 |
Counting Everyday Objects in Everyday Scenes |
prithv1/cvpr2017_counting |
7 |
Compare |
| Weakly-supervised instance segmentation |
COCO test-dev |
DiscoBox (ResNeXt-101-DCN-FPN) AP 37.9 |
DiscoBox: Weakly Supervised Instance Segmentation and... |
NVlabs/DiscoBox +2 |
7 |
Compare |
| Image Retrieval |
COCO (Common Objects in Context) |
BLIP-2 ViT-G (fine-tuned) recall@1 68.3 |
BLIP-2: Bootstrapping Language-Image Pre-training with... |
huggingface/transformers +16 |
6 |
Compare |
| Unsupervised Semantic Segmentation |
COCO-Stuff-3 |
SAN Pixel Accuracy 80.3 |
Rethinking Alignment and Uniformity in Unsupervised... |
— |
6 |
Compare |
| Layout-to-Image Generation |
COCO-Stuff 256x256 |
LayoutDiffusion (25steps) FID 31.68 |
LayoutDiffusion: Controllable Diffusion Model for... |
zgctroy/layoutdiffusion +1 |
5 |
Compare |
| Weakly Supervised Object Detection |
COCO (Common Objects in Context) |
MSLPD MAP 56.6 |
Few-Example Object Detection with Model Communication |
D-X-Y/DXY-Projects |
5 |
Compare |
| Knowledge Distillation |
COCO (Common Objects in Context) |
ADLIK-Faster (T: Faster R-CNN vit-base S: Faster R-CNN deit-small) box AP 47.6 |
Focal and Global Knowledge Distillation for Detectors |
yzd-v/FGD |
4 |
Compare |
| Multi-Person Pose Estimation |
COCO minival |
HRNet-W48plus AP 79.1 |
AID: Pushing the Performance Boundary of Human Pose... |
open-mmlab/mmpose +1 |
4 |
Compare |
| One-Shot Object Detection |
COCO (Common Objects in Context) |
OWL-ViT (R50+H/32) AP 0.5 41.8 |
Simple Open-Vocabulary Object Detection with Vision Transformers |
google-research/scenic +1 |
4 |
Compare |
| Question Generation |
COCO Visual Question Answering (VQA) real images 1.0 open ended |
MDN BLEU-1 65.1 |
Multimodal Differential Network for Visual Question Generation |
badripatro/MDN-VQG |
4 |
Compare |
| Visual Question Answering (VQA) |
COCO Visual Question Answering (VQA) real images 2.0 open ended |
HDU-USYD-UNCC Percentage correct 68.16 |
VQA: Visual Question Answering |
ramprs/grad-cam +20 |
4 |
Compare |
| Visual Question Answering (VQA) |
COCO Visual Question Answering (VQA) abstract images 1.0 open ended |
Graph VQA Percentage correct 70.42 |
Graph-Structured Representations for Visual Question Answering |
— |
4 |
Compare |
| Visual Question Answering (VQA) |
COCO Visual Question Answering (VQA) abstract 1.0 multiple choice |
Graph VQA Percentage correct 74.37 |
Graph-Structured Representations for Visual Question Answering |
— |
4 |
Compare |
| Weakly Supervised Object Detection |
COCO test-dev |
wetectron(single-model, VGG16) AP50 24.8 |
Instance-aware, Context-focused, and Memory-efficient... |
NVlabs/wetectron +1 |
4 |
Compare |
| Object Detection |
COCO (Common Objects in Context) |
MOAT-3 22K+1K box AP 59.2 |
MOAT: Alternating Mobile Convolution and Attention... |
google-research/deeplab2 +1 |
3 |
Compare |
| Conditional Image Generation |
COCO-Animals |
U-Net GAN FID 13.73 |
A U-Net Based Discriminator for Generative Adversarial Networks |
boschresearch/unetgan +2 |
2 |
Compare |
| Cross-Modal Retrieval |
MSCOCO-1k |
NAPReg Image-to-text R@1 81.9 |
NAPReg: Nouns As Proxies Regularization for Semantically... |
bhavinjawade/NAPReq |
2 |
Compare |
| Open World Object Detection |
COCO 2017 (Outdoor, Accessories, Appliance, Truck) |
ORE (MDef-DETR) Unknown Recall 49.54 |
Class-agnostic Object Detection with Multi-modal Transformer |
mmaaz60/mvits_for_class_agnostic_od |
2 |
Compare |
| Open World Object Detection |
COCO 2017 (Sports, Food) |
ORE (MDef-DETR) Unknown Recall 50.89 |
Class-agnostic Object Detection with Multi-modal Transformer |
mmaaz60/mvits_for_class_agnostic_od |
2 |
Compare |
| Open World Object Detection |
COCO 2017 (Electronic, Indoor, Kitchen, Furniture) |
ORE (MDef-DETR) MAP 31.66 |
Class-agnostic Object Detection with Multi-modal Transformer |
mmaaz60/mvits_for_class_agnostic_od |
2 |
Compare |
| Panoptic Segmentation |
COCO panoptic |
VAN-B6* PQ 58.2 |
Visual Attention Network |
huggingface/transformers +20 |
2 |
Compare |
| Robust Object Detection |
COCO (Common Objects in Context) |
Faster R-CNN with Stylized Training Data mPC [AP] 20.4 |
Benchmarking Robustness in Object Detection: Autonomous... |
bethgelab/imagecorruptions +3 |
2 |
Compare |
| Text-to-Image Generation |
MS-COCO |
AttnGAN Inception score 25.89 |
AttnGAN: Fine-Grained Text to Image Generation with... |
taoxugit/AttnGAN +19 |
2 |
Compare |
| Active Object Detection |
COCO (Common Objects in Context) |
RetinaNet AP (7.3, 13.8, 16.9, 19.1, 20.8) on 2% ~ 10% |
Multiple instance active learning for object detection |
yuantn/MI-AOD |
1 |
Compare |
| Activeness Detection |
COCO test-dev |
Lightweight OpenPose Accuracy (%) 76.67 |
ActiveNet: A computer-vision based approach to determine lethargy |
aaditagarwal/ActiveNet |
1 |
Compare |
| Few-Shot Object Detection |
COCO 2017 |
DETReg (ours) AP 30 |
DETReg: Unsupervised Pretraining with Region Priors for... |
amirbar/detreg |
1 |
Compare |
| Homography Estimation |
COCO 2014 |
PFNet MACE 0.92 |
Rethinking Planar Homography Estimation Using Perspective Fields |
ruizengalways/PFNet |
1 |
Compare |
| Image Captioning |
MS-COCO |
NeuSyRE BLEU-1 79.1 |
NeuSyRE: Neuro-Symbolic Visual Understanding and... |
jaleedkhan/neusire |
1 |
Compare |
| Instance Segmentation |
coco minval |
R3-CNN (ResNet-50-FPN, GC-Net) APL 56 |
Recursively Refined R-CNN: Instance Segmentation with... |
IMPLabUniPr/mmdetection |
1 |
Compare |
| Interactive Segmentation |
COCO (Common Objects in Context) |
IOG Instance Average IoU 85.2 |
Interactive Object Segmentation With Inside-Outside Guidance |
shiyinzhang/Inside-Outside-Guidance +1 |
1 |
Compare |
| Interactive Segmentation |
COCO minival |
ViT-B+MST+CL NoC@85 2.08 |
MST: Adaptive Multi-Scale Tokens Guided Interactive Segmentation |
hahamyt/mst |
1 |
Compare |
| Multi-Label Learning |
COCO 2014 |
SADCL CF1 79.8 |
Semantic-Aware Dual Contrastive Learning for Multi-label... |
yu-gi-oh-leilei/sadcl |
1 |
Compare |
| Multi-object discovery |
COCO_20k |
Large-scale rOSD Detection Rate 12.0 |
Toward unsupervised, multi-object discovery in... |
huyvvo/rOSD |
1 |
Compare |
| Object Detection |
COCO+ |
RepPoints + Self-adaptation mAR (COCO+ XS) 28.4 |
Slender Object Detection: Diagnoses and Improvements |
wanzysky/SlenderObjDet |
1 |
Compare |
| Object Proposal Generation |
COCO (Common Objects in Context) |
MDef-DETR (Off-the-shelf evaluation) Average Recall 0.6503 |
Class-agnostic Object Detection with Multi-modal Transformer |
mmaaz60/mvits_for_class_agnostic_od |
1 |
Compare |
| One-Shot Instance Segmentation |
COCO (Common Objects in Context) |
Siamese Mask R-CNN AP 0.5 14.5 |
One-Shot Instance Segmentation |
bethgelab/siamese-mask-rcnn +2 |
1 |
Compare |
| Point-Supervised Instance Segmentation |
COCO test-dev |
BESTIE (proposal-free) AP 17.8 |
Beyond Semantic to Instance Segmentation:... |
clovaai/BESTIE |
1 |
Compare |
| Pose Estimation |
COCO minival |
MSPN AP 75.9 |
Rethinking on Multi-Stage Networks for Human Pose Estimation |
open-mmlab/mmpose +6 |
1 |
Compare |
| Pose Estimation |
MS-COCO |
UniHCP (finetune) AP 76.5 |
UniHCP: A Unified Model for Human-Centric Perceptions |
opengvlab/unihcp |
1 |
Compare |
| Quantization |
COCO (Common Objects in Context) |
SSD ResNet50 V1 FPN 640x640 MAP 34.3 |
HPTQ: Hardware-Friendly Post Training Quantization |
sony/model_optimization |
1 |
Compare |
| Question Answering |
COCO Visual Question Answering (VQA) real images 1.0 open ended |
MaMMUT (2B) Test 80.8 |
MaMMUT: A Simple Architecture for Joint Learning for... |
lucidrains/mammut-pytorch |
1 |
Compare |
| Real-time Instance Segmentation |
MSCOCO-1k |
RTMDet-Ins-x APM 49.0 |
RTMDet: An Empirical Study of Designing Real-Time Object... |
open-mmlab/mmdetection +13 |
1 |
Compare |
| Scene Graph Generation |
MS-COCO |
NeuSyRE R@100 38.5 |
NeuSyRE: Neuro-Symbolic Visual Understanding and... |
jaleedkhan/neusire |
1 |
Compare |
| Semi Supervised Learning for Image Captioning |
COCO (Common Objects in Context) |
Perturb, Predict & Paraphrase CIDEr 84.5 |
Perturb, Predict & Paraphrase: Semi-Supervised Learning... |
csalt-research/perturb-predict-paraphrase |
1 |
Compare |
| Unsupervised Object Localization |
COCO_20k |
DeepCut CorLoc 61.6 |
DeepCut: Unsupervised Segmentation using Graph Neural... |
sampl-weizmann/deepcut |
1 |
Compare |
| Unsupervised Semantic Segmentation with Language-image Pre-training |
COCO (Common Objects in Context) |
CLIPpy ViT-B Mean IoU (val) 25.5 |
Perceptual Grouping in Contrastive Vision-Language Models |
kahnchana/clippy +1 |
1 |
Compare |
| Visual Question Answering |
COCO Visual Question Answering (VQA) real images 2.0 open ended |
MaMMUT (2B) Percentage correct 80.7 |
MaMMUT: A Simple Architecture for Joint Learning for... |
lucidrains/mammut-pytorch |
1 |
Compare |