| Image Classification |
ImageNet |
CoCa (finetuned) Top 1 Accuracy 91.0% |
CoCa: Contrastive Captioners are Image-Text Foundation Models |
mlfoundations/open_clip +5 |
1,060 |
Compare |
| Self-Supervised Image Classification |
ImageNet |
DINOv2+reg (ViT-g/14) Top 1 Accuracy 87.1 |
Vision Transformers Need Registers |
rwightman/pytorch-image-models +5 |
144 |
Compare |
| Neural Architecture Search |
ImageNet |
DeepMAD-50M Top-1 Error Rate 16.1 |
DeepMAD: Mathematical Architecture Design for Deep... |
alibaba/lightweight-neural-architecture-search |
135 |
Compare |
| Semi-Supervised Image Classification |
ImageNet - 10% labeled data |
DHO (ViT-Large) Top 1 Accuracy 85.9% |
Simple Semi-supervised Knowledge Distillation from... |
erjui/DHO |
75 |
Compare |
| Self-Supervised Image Classification |
ImageNet (finetuned) |
DINOv2 (ViT-g/14, 448) Top 1 Accuracy 88.9% |
DINOv2: Learning Robust Visual Features without Supervision |
huggingface/transformers +25 |
65 |
Compare |
| Semi-Supervised Image Classification |
ImageNet - 1% labeled data |
DHO (ViT-Large) Top 1 Accuracy 84.6% |
Simple Semi-supervised Knowledge Distillation from... |
erjui/DHO |
65 |
Compare |
| Knowledge Distillation |
ImageNet |
ScaleKD (T:BEiT-L S:ViT-B/14) Top-1 accuracy % 86.43 |
ScaleKD: Strong Vision Transformers Could Be Excellent Teachers |
deep-optimization/scalekd |
52 |
Compare |
| Image Classification |
ImageNet V2 |
Model soups (BASIC-L) Top 1 Accuracy 84.63 |
Model soups: averaging weights of multiple fine-tuned... |
mlfoundations/model-soups +5 |
33 |
Compare |
| Quantization |
ImageNet |
FQ-ViT (ViT-L) Top-1 Accuracy (%) 85.03 |
FQ-ViT: Post-Training Quantization for Fully Quantized... |
megvii-research/FQ-ViT |
27 |
Compare |
| Zero-Shot Transfer Image Classification |
ImageNet |
M2-Encoder Param 10B |
M2-Encoder: Advancing Bilingual Image-Text Understanding... |
alipay/Ant-Multi-Modal-Framework |
23 |
Compare |
| Data Augmentation |
ImageNet |
DeiT-B (+MixPro) Accuracy (%) 82.9 |
MixPro: Data Augmentation with MaskMix and Progressive... |
fistyee/mixpro |
17 |
Compare |
| Network Pruning |
ImageNet |
ResNet50-2.3 GFLOPs Accuracy 78.79 |
Pruning Filters for Efficient ConvNets |
PaddlePaddle/PaddleOCR +20 |
16 |
Compare |
| Image Reconstruction |
ImageNet |
MGVQ (16x16x8) FID 0.49 |
MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer... |
MKJia/MGVQ |
15 |
Compare |
| Prompt Engineering |
ImageNet |
PromptKD Harmonic mean 77.62 |
PromptKD: Unsupervised Prompt Distillation for... |
zhengli97/promptkd |
15 |
Compare |
| Contrastive Learning |
imagenet-1k |
ResNet50 ImageNet Top-1 Accuracy 73.6 |
Matrix Information Theory for Self-Supervised Learning |
yifanzhang-pro/matrix-ssl +2 |
14 |
Compare |
| Zero-Shot Transfer Image Classification |
ImageNet V2 |
BASIC (Lion) Accuracy (Private) 81.2 |
— |
— |
13 |
Compare |
| Image Clustering |
ImageNet |
TURTLE (CLIP + DINOv2) Accuracy 72.9 |
Let Go of Your Labels with Unsupervised Transfer |
mlbio-epfl/turtle |
12 |
Compare |
| Model Compression |
ImageNet |
ADLIK-MO-ResNet50+W4A4 Top-1 77.878 |
Learned Step Size Quantization |
zhutmost/lsq-net +8 |
12 |
Compare |
| Zero-Shot Composed Image Retrieval (ZS-CIR) |
ImageNet |
iSEARLE-XL (CLIP L/14) Average Recall 24.46 |
iSEARLE: Improving Textual Inversion for Zero-Shot... |
miccunifi/searle +1 |
11 |
Compare |
| Sparse Learning |
ImageNet |
Resnet-50: 80% Sparse Top-1 Accuracy 77.1 |
Rigging the Lottery: Making All Tickets Winners |
google-research/rigl +10 |
9 |
Compare |
| Unsupervised Image Classification |
ImageNet |
TURTLE (CLIP + DINOv2) Accuracy (%) 72.9 |
Let Go of Your Labels with Unsupervised Transfer |
mlbio-epfl/turtle |
9 |
Compare |
| Feature Upsampling |
ImageNet |
JAFAR ADCC 73.3 |
JAFAR: Jack up Any Feature at Any Resolution |
PaulCouairon/JAFAR |
8 |
Compare |
| Few-Shot Image Classification |
ImageNet - 1-shot |
ViT-MoE-15B (Every-2) Top 1 Accuracy 68.66 |
Scaling Vision with Sparse Mixture of Experts |
google-research/vmoe |
8 |
Compare |
| Few-Shot Image Classification |
ImageNet - 5-shot |
ViT-MoE-15B (Every-2) Top 1 Accuracy 82.78 |
Scaling Vision with Sparse Mixture of Experts |
google-research/vmoe |
8 |
Compare |
| Prompt Engineering |
ImageNet V2 |
HPT++ Top-1 accuracy % 65.31 |
HPT++: Hierarchically Prompting Vision-Language Models... |
vill-lab/2024-aaai-hpt +1 |
8 |
Compare |
| Few-Shot Image Classification |
ImageNet - 10-shot |
MAWS (ViT-6.5B) Top 1 Accuracy 84.6 |
The effectiveness of MAE pre-pretraining for... |
facebookresearch/maws |
7 |
Compare |
| Image Super-Resolution |
ImageNet |
DAVI FID 36.27 |
Diffusion Prior-Based Amortized Variational Inference... |
mlvlab/davi +1 |
6 |
Compare |
| JPEG Decompression |
ImageNet |
Palette (QF: 20) FID-5K 4.3 |
Palette: Image-to-Image Diffusion Models |
Janspiry/Palette-Image-to-Image-Diffusion-Models +4 |
6 |
Compare |
| Weakly-Supervised Object Localization |
ImageNet |
Stable diffusion GT-known localization accuracy 75.0 |
Generative Prompt Model for Weakly Supervised Object Localization |
callsys/genpromp |
6 |
Compare |
| Few-Shot Image Classification |
ImageNet - 0-Shot |
DebiasPL (ResNet50) Accuracy 68.3% |
Debiased Learning from Naturally Imbalanced Pseudo-Labels |
frank-xwang/debiased-pseudo-labeling |
5 |
Compare |
| Image Inpainting |
ImageNet |
WavePaint FID 3.21 |
WavePaint: Resource-efficient Token-mixer for... |
pranavphoenix/WavePaint |
5 |
Compare |
| Adversarial Robustness |
ImageNet |
ResNet-50 (SGD, Cosine) Accuracy 77.4 |
Are Transformers More Robust Than CNNs? |
ytongbai/ViTs-vs-CNNs |
4 |
Compare |
| Image Classification with Differential Privacy |
ImageNet |
NFResnet-50 Top 1 Accuracy 39.2 |
TAN Without a Burn: Scaling Laws of DP-SGD |
facebookresearch/tan |
4 |
Compare |
| Image Colorization |
ImageNet |
DDRM Consistency 260.4 |
Zero-Shot Image Restoration Using Denoising Diffusion... |
wyhuai/ddnm +3 |
4 |
Compare |
| Weakly Supervised Object Detection |
ImageNet |
PCL-OB-G-Ens + FRCNN MAP 19.6 |
PCL: Proposal Cluster Learning for Weakly Supervised... |
ppengtang/pcl.pytorch +3 |
4 |
Compare |
| Adversarial Defense |
ImageNet |
ResNet101 Accuracy 99.8% |
NOMARO: Defending against Adversarial Attacks by... |
as791/NOMARO_defense |
3 |
Compare |
| Image Deblurring |
ImageNet |
DDNM FID 1.15 |
Zero-Shot Image Restoration Using Denoising Diffusion... |
wyhuai/ddnm +3 |
3 |
Compare |
| Semi-Supervised Image Classification |
ImageNet - 0.2% labeled data |
DebiasPL (ResNet-50) ImageNet Top-1 Accuracy 69.6% |
Debiased Learning from Naturally Imbalanced Pseudo-Labels |
frank-xwang/debiased-pseudo-labeling |
3 |
Compare |
| Color Image Denoising |
ImageNet sigma100 |
DMID-p LPIPS 0.156 |
Stimulating Diffusion Model for Image Denoising via... |
li-tong-621/dmid |
2 |
Compare |
| Color Image Denoising |
ImageNet sigma150 |
DMID-p LPIPS 0.259 |
Stimulating Diffusion Model for Image Denoising via... |
li-tong-621/dmid |
2 |
Compare |
| Color Image Denoising |
ImageNet sigma200 |
DMID-p LPIPS 0.259 |
Stimulating Diffusion Model for Image Denoising via... |
li-tong-621/dmid |
2 |
Compare |
| Color Image Denoising |
ImageNet sigma250 |
DMID-p LPIPS 0.289 |
Stimulating Diffusion Model for Image Denoising via... |
li-tong-621/dmid |
2 |
Compare |
| Color Image Denoising |
ImageNet sigma50 |
DMID-p LPIPS 0.087 |
Stimulating Diffusion Model for Image Denoising via... |
li-tong-621/dmid |
2 |
Compare |
| Medical Image Classification |
ImageNet |
DaViT-T GFLOPs 4.5 |
DaViT: Dual Attention Vision Transformers |
rwightman/pytorch-image-models +3 |
2 |
Compare |
| Zero-Shot Learning |
ImageNet |
ZLaP Top 1 Accuracy 72.1 |
Label Propagation for Zero-shot Classification with... |
vladan-stojnic/zlap |
2 |
Compare |
| Image Classification |
imagenet-1k |
BinaryViT Top 1 Accuracy 70.6 |
BinaryViT: Pushing Binary Vision Transformers Towards... |
phuoc-hoan-le/binaryvit |
1 |
Compare |
| Image Clustering |
imagenet-1k |
TAC ARI 0.435 |
Image Clustering with External Guidance |
xlearning-scu/2024-icml-tac |
1 |
Compare |
| Image Segmentation |
ImageNet |
MobileOne-S0 GFLOPs 0.275 |
MobileOne: An Improved One millisecond Mobile Backbone |
rwightman/pytorch-image-models +9 |
1 |
Compare |
| Transductive Zero-Shot Classification |
ImageNet |
ZLaP Top 1 Accuracy 72.7 |
Label Propagation for Zero-shot Classification with... |
vladan-stojnic/zlap |
1 |
Compare |
| Visual Question Answering (VQA) |
ImageNet |
BLIP-2 OPT ClipMatch@1 57.10 |
Open-ended VQA benchmarking of Vision-Language models by... |
lmb-freiburg/ovqa |
1 |
Compare |
| Classification |
imagenet-1k |
no rows |
— |
— |
0 |
Compare |
| Semi-Supervised Image Classification |
ImageNet |
no rows |
— |
— |
0 |
Compare |