Browse State-of-the-Art › Object Detection

Object Detection

4,657 papers with code · 123 benchmarks · 332 datasets archive 2025-07-28

Computer Vision

Benchmarks archive 2025-07-28

124 leaderboard tables shown for this task (3 more in the archive withheld as spam; see /not-shown), 123 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 124 until expanded.

DatasetBest model (first row in archive order)PaperCodeSyntologyCompare
COCO test-dev (225 rows) Co-DETR DETRs with Collaborative Hybrid Assignments Training code Syntology ran 0 of 5 samples · 5 unverified Compare
COCO minival (220 rows) PE_spatial (DETA) Perception Encoder: The best visual embeddings are not at the... code Syntology ran 12 of 18 samples · 6 unverified Compare
COCO-O (45 rows) EVA EVA: Exploring the Limits of Masked Visual Representation Learning at Scale code Syntology ran 1 of 3 samples · 2 unverified Compare
COCO 2017 val (33 rows) Mr. DETR (Swin-L, 1x, 5cale) Mr. DETR: Instructive Multi-Route Training for Detection Transformers code — Compare
PASCAL VOC 2007 (30 rows) Cascade Eff-B7 NAS-FPN (Copy Paste pre-training, single-scale) Simple Copy-Paste is a Strong Data Augmentation Method for... code Syntology ran 2 of 2 samples · 0 unverified Compare
COCO 2017 (24 rows) MaxViT-B MaxViT: Multi-Axis Vision Transformer code Syntology ran 33 of 53 samples · 20 unverified Compare
CrowdHuman (full body) (19 rows) InternImage-H InternImage: Exploring Large-Scale Vision Foundation Models with... code Syntology ran 2 of 4 samples · 2 unverified Compare
CPPE-5 (16 rows) TridentNet CPPE-5: Medical Personal Protective Equipment Dataset code — Compare
LVIS v1.0 val (15 rows) Co-DETR (single-scale) DETRs with Collaborative Hybrid Assignments Training code Syntology ran 0 of 5 samples · 5 unverified Compare
Manga109-s 15test (15 rows) YOLOX-L USB: Universal-Scale Object Detection Benchmark code — Compare
Waymo 2D detection all_ns f0val (15 rows) YOLOX-L USB: Universal-Scale Object Detection Benchmark code — Compare
PKU-DDD17-Car (14 rows) CAFR Embracing Events and Frames with Hierarchical Feature Refinement... code Syntology ran 1 of 1 samples · 0 unverified Compare
USB (Standard USB 1.0 protocol) (13 rows) UniverseNet-20.08 USB: Universal-Scale Object Detection Benchmark code — Compare
DSEC (12 rows) CAFR Embracing Events and Frames with Hierarchical Feature Refinement... code Syntology ran 1 of 1 samples · 0 unverified Compare
SFCHD (12 rows) YOLOv8+SCALE Large, Complex, and Realistic Safety Clothing and Helmet... code — Compare
GEN1 Detection (11 rows) ERGO-12 From Chaos Comes Order: Ordering Event Representations for Object... code — Compare
SeaDronesSee (10 rows) Synth Pretrained Faster R-CNN ResNeXt-101-FPN Leveraging Synthetic Data in Object Detection on Unmanned Aerial Vehicles code — Compare
UA-DETRAC (9 rows) VSTAM Video Sparse Transformer With Attention-Guided Memory for Video... code — Compare
ODinW Full-Shot 13 Tasks (8 rows) CP-DETR-L(only optimize prompt) CP-DETR: Concept Prompt Guide DETR Toward Stronger Universal... — — Compare
UAVDT (8 rows) PRB-FPN Parallel Residual Bi-Fusion Feature Pyramid Network for Accurate... code — Compare
AI-TOD (7 rows) DNTR A DeNoising FPN With Transformer R-CNN for Tiny Object Detection code — Compare
MSCOCO (7 rows) PP-PicoDet-L PP-PicoDet: A Better Real-Time Object Detector on Mobile Devices code — Compare
NAO (7 rows) Mask RCNN R50 Natural Adversarial Objects — — Compare
PASCAL VOC 2012 (7 rows) InternImage-H InternImage: Exploring Large-Scale Vision Foundation Models with... code Syntology ran 2 of 4 samples · 2 unverified Compare
TBBR (7 rows) Swin-T (ImageNet-1k pretrain) Deep learning approaches to building rooftop thermal bridge... code — Compare
BigDetection val (6 rows) Cascade R-CNN (R50-FPN) BigDetection: A Large-scale Benchmark for Improved Object Detector... code — Compare
EventPed (6 rows) MMPedestron When Pedestrian Detection Meets Multi-Modal Learning: Generalist... code — Compare
GRAZPEDWRI-DX (6 rows) YOLOv8x Enhancing Wrist Fracture Detection with YOLO code — Compare
InOutDoor (6 rows) MMPedestron When Pedestrian Detection Meets Multi-Modal Learning: Generalist... code — Compare
LVIS v1.0 minival (6 rows) Co-DETR (single-scale) DETRs with Collaborative Hybrid Assignments Training code Syntology ran 0 of 5 samples · 5 unverified Compare
PeopleArt (6 rows) PVT (Pyramid Vision Transformer; trained on PeopleArt and PopArt) Poses of People in Art: A Data Set for Human Pose Estimation in... — — Compare
STCrowd (6 rows) MMPedestron When Pedestrian Detection Meets Multi-Modal Learning: Generalist... code — Compare
iSAID (5 rows) PANet++ iSAID: A Large-scale Dataset for Instance Segmentation in Aerial Images code Syntology ran 1 of 1 samples · 0 unverified Compare
KITTI Cars Easy (5 rows) Patches Patch Refinement -- Localized 3D Object Detection — — Compare
KITTI Cars Hard (5 rows) Patches Patch Refinement -- Localized 3D Object Detection — — Compare
VisDrone-DET2019 (5 rows) CZ Det. Cascaded Zoom-in Detector for High Resolution Aerial Images code — Compare
Waymo Open Dataset (5 rows) LeapMotor_Det 2nd Place Solution for Waymo Open Dataset Challenge - Real-time 2D... code — Compare
India Driving Dataset (4 rows) YOLOv5x YOLO-Drone:Airborne real-time detection of dense small objects... — — Compare
KITTI Cars Moderate (4 rows) Patches Patch Refinement -- Localized 3D Object Detection — — Compare
Visual Genome (4 rows) KnowZRel KnowZRel: Common Sense Knowledge-based Zero-Shot Relationship... code — Compare
WaterScenes (4 rows) YOLOv8-M WaterScenes: A Multi-Task 4D Radar-Camera Fusion Dataset and... code — Compare
WiderPerson (4 rows) IterDet (Faster RCNN, ResNet50, 2 iterations) IterDet: Iterative Scheme for Object Detection in Crowded Environments — — Compare
COCO (Common Objects in Context) (3 rows) MOAT-3 22K+1K MOAT: Alternating Mobile Convolution and Attention Brings Strong... code — Compare
FlickrLogos-32 (3 rows) Logo-Yolo LogoDet-3K: A Large-Scale Image Dataset for Logo Detection code — Compare
OoDIS (3 rows) UGainS UGainS: Uncertainty Guided Anomaly Instance Segmentation code — Compare
Pascal VOC to Clipart1K (3 rows) DDT Diffusion Domain Teacher: Diffusion Guided Domain Adaptive Object Detector code — Compare
VEDAI (3 rows) GHOST Guided Hybrid Quantization for Object detection in Multimodal... code — Compare
(unnamed in the archive) (2 rows) Fine tuned Yolov5xu Crossing Language Borders: A Pipeline for Indonesian Manhwa Translation code — Compare
Drone vs Bird (2 rows) OBSS YOLOv5+Track Boosting (Including Synthetic Data) Track Boosting and Synthetic Data Aided Drone Detection — — Compare
IndustReal (2 rows) YoloV8 IndustReal: A Dataset for Procedure Step Recognition Handling... code Syntology ran 4 of 5 samples · 1 unverified Compare
Manga109 (2 rows) DASS-Detector (YOLOX XL) Domain-Adaptive Self-Supervised Pre-Training for Face & Body... code — Compare
nuScenes (2 rows) BIRANet(RGB+Radar) Radar+RGB Attentive Fusion for Robust Object Detection in... code — Compare
ODinW Full-shot 35 Tasks (2 rows) Grounding DINO 1.5 Pro Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection code Syntology ran 1 of 2 samples · 1 unverified Compare
OpenImages-v6 (2 rows) ScaleDet ScaleDet: A Scalable Multi-Dataset Object Detector — — Compare
PASCAL VOC 10% (2 rows) DETReg (MDef-DETR) Class-agnostic Object Detection with Multi-modal Transformer code — Compare
PASCAL VOC to Watercolor2k (2 rows) DDT Diffusion Domain Teacher: Diffusion Guided Domain Adaptive Object Detector code — Compare
PASCAL VOC to Comic2k (2 rows) DDT Diffusion Domain Teacher: Diffusion Guided Domain Adaptive Object Detector code — Compare
SA-Det-100k (2 rows) Relation-DETR (ResNet50 1x) Relation DETR: Exploring Explicit Position Relation Prior for... code Syntology ran 1 of 1 samples · 0 unverified Compare
SIXray (2 rows) LRPz Towards Best Practice in Explaining Neural Network Decisions with LRP code — Compare
SpaceNet 2 (2 rows) YOLT SpaceNet: A Remote Sensing Dataset and Challenge Series code — Compare
A Dataset of Multispectral Potato Plants Images (1 row) Retina-UNet-Ag Potato Crop Stress Identification in Aerial Images using Deep... — — Compare
A2D (1 row) RL [10] Lpixel Paint Transformer: Feed Forward Neural Painting with Stroke Prediction code — Compare
AODRaw (1 row) Cascade RCNN (ConvNext-T, RAW pre-training) Towards RAW Object Detection in Diverse Conditions code Syntology ran 1 of 2 samples · 1 unverified Compare
AquaTrash (1 row) Aquavision AquaVision: Automating the detection of waste in water bodies... code — Compare
BDD100K (1 row) CDDMSL Semi-Supervised Domain Generalization for Object Detection via... code — Compare
BDD100K val (1 row) hybrid incremental net On Generalizing Detection Models for Unconstrained Environments code — Compare
C2A: Human Detection in Disaster Scenarios (1 row) B2BDet From Blurry to Brilliant Detection: YOLOv5-Based Aerial Object... — — Compare
CISOL - Track A - TD-TSR (1 row) YOLO v8.1m CISOL: An Open and Extensible Dataset for Table Structure... — — Compare
CISOL - Track B - TSR-only (1 row) YOLO v8.1m CISOL: An Open and Extensible Dataset for Table Structure... — — Compare
CityPersons (1 row) V2F-Net V2F-Net: Explicit Decomposition of Occluded Pedestrian Detection — — Compare
Cityscapes to Foggy Cityscapes (1 row) CDDMSL Semi-Supervised Domain Generalization for Object Detection via... code — Compare
Clipart1k (1 row) CDDMSL Semi-Supervised Domain Generalization for Object Detection via... code — Compare
COCO (1 row) ColorMAE-Green-ViTB-1600 ColorMAE: Exploring data-independent masking strategies in Masked... code — Compare
COCO+ (1 row) RepPoints + Self-adaptation Slender Object Detection: Diagnoses and Improvements code — Compare
COCO val2017 (1 row) SynCo (ResNet-50) 200ep SynCo: Synthetic Hard Negatives in Contrastive Learning for Better... code — Compare
Comic2k (1 row) CDDMSL Semi-Supervised Domain Generalization for Object Detection via... code — Compare
CrowdHuman (1 row) S-RCNN+Ours Progressive End-to-End Object Detection in Crowded Scenes code Syntology ran 1 of 1 samples · 0 unverified Compare
DeepTrash (1 row) YOLOv5 A Robotic Approach towards Quantifying Epipelagic Bound Plastic... code — Compare
Drinking Waste Classification (1 row) EfficientDet-D2 Waste detection in Pomerania: non-profit project for detecting... code — Compare
ELEVATER (1 row) GLIP-T ELEVATER: A Benchmark and Toolkit for Evaluating... code Syntology ran 4 of 20 samples · 16 unverified Compare
EVD4UAV (1 row) yolov8x-seg EVD4UAV: An Altitude-Sensitive Benchmark to Evade Vehicle Detection in UAV code — Compare
ExDark (1 row) EMV-YOLO Toward Highly Efficient Semantic-Guided Machine Vision for... code — Compare
Extended TACO-1 (1 row) EfficientDet-D2 Waste detection in Pomerania: non-profit project for detecting... code — Compare
Extended TACO-7 (1 row) EfficientDet-D2 Waste detection in Pomerania: non-profit project for detecting... code — Compare
Extragalactic Planetary Nebulae (1 row) PNe within NGC1380 & NGC1404 Fornax 3D project: automated detection of planetary nebulae in the... code — Compare
FLIR (1 row) MiPa MiPa: Mixed Patch Infrared-Visible Modality Agnostic Object Detection code — Compare
GMOT-40 (1 row) iGDINO MAC-SORT TP-GMOT: Tracking Generic Multiple Object by Textual Prompt with... code — Compare
GQA (1 row) KnowZRel KnowZRel: Common Sense Knowledge-based Zero-Shot Relationship... code — Compare
KITTI Cyclists Easy (1 row) Vote3Deep Vote3Deep: Fast Object Detection in 3D Point Clouds Using... — — Compare
KITTI Cyclists Hard (1 row) Vote3Deep Vote3Deep: Fast Object Detection in 3D Point Clouds Using... — — Compare
KITTI Cyclists Moderate (1 row) Vote3Deep Vote3Deep: Fast Object Detection in 3D Point Clouds Using... — — Compare
KITTI Pedestrians Moderate (1 row) Vote3Deep Vote3Deep: Fast Object Detection in 3D Point Clouds Using... — — Compare
KITTI Pedestrians Easy (1 row) Vote3Deep Vote3Deep: Fast Object Detection in 3D Point Clouds Using... — — Compare
KITTI Pedestrians Hard (1 row) Vote3Deep Vote3Deep: Fast Object Detection in 3D Point Clouds Using... — — Compare
LDD (1 row) R^3-CNN LDD: A Dataset for Grape Diseases Object Detection and Instance... — — Compare
LeukemiaAttri (1 row) AttriDet A Large-scale Multi Domain Leukemia Dataset for the White Blood... code — Compare
LLVIP (1 row) MiPa MiPa: Mixed Patch Infrared-Visible Modality Agnostic Object Detection code — Compare
LVIS v1.0 (1 row) ScaleDet ScaleDet: A Scalable Multi-Dataset Object Detector — — Compare
MJU-Waste (1 row) EfficientDet-D2 Waste detection in Pomerania: non-profit project for detecting... code — Compare
MSCOCO (1 row) DAS DAS: A Deformable Attention to Capture Salient Information in CNNs — — Compare
Multispectral Dataset (1 row) TarDAL Target-aware Dual Adversarial Learning and a Multi-scenario... code — Compare
MUSES: MUlti-SEnsor Semantic perception dataset (1 row) Mask2Former (R50) MUSES: The Multi-Sensor Semantic Perception Dataset for Driving... code Syntology ran 12 of 13 samples · 1 unverified Compare
NII-CU MAPD (1 row) YOLOv3 Deep learning with RGB and thermal images onboard a drone for... — — Compare
Objects365 (1 row) ScaleDet ScaleDet: A Scalable Multi-Dataset Object Detector — — Compare
PASCAL Part 2010 - Animals (1 row) Attention-based Joint Detection of Object and Semantic Part Attention-based Joint Detection of Object and Semantic Part code Syntology ran 3 of 12 samples · 9 unverified Compare
PASCAL VOC (1 row) TinyissimoYOLO-v8 Ultra-Efficient On-Device Object Detection on AI-Integrated Smart... code — Compare
PASCAL VOC 2007 (15+5) (1 row) Faster R-CNN Faster R-CNN: Towards Real-Time Object Detection with Region... code Syntology ran 59 of 124 samples · 65 unverified Compare
PASCAL VOC 2012 test (1 row) SynCo (ResNet-50) 200ep SynCo: Synthetic Hard Negatives in Contrastive Learning for Better... code — Compare
PKU-DDD17-Car (1 row) Faster-RCNN Faster R-CNN: Towards Real-Time Object Detection with Region... code Syntology ran 59 of 124 samples · 65 unverified Compare
STN PLAD (1 row) MS-PAD STN PLAD: A Dataset for Multi-Size Power Line Assets Detection in... code — Compare
SAR-AIRcraft-1.0 (1 row) PGD-YOLOv8 Physics-Guided Detector for SAR Airplanes code — Compare
SHEL5K (1 row) YOLO SHEL5K: An Extended Dataset and Benchmarking for Safety Helmet Detection code — Compare
Songdo Vision (1 row) Geo-trax Advanced computer vision for extracting georeferenced vehicle... code — Compare
SpaceNet 1 (1 row) YOLT SpaceNet: A Remote Sensing Dataset and Challenge Series code — Compare
SUN-RGBD val (1 row) CDSSD How To Extract Fashion Trends From Social Media? A Robust Object... code — Compare
TexBiG 2022 test (1 row) VSR (Vison, Semantics and Relation Model) A Dataset for Analysing Complex Document Layouts in the Digital... code — Compare
TexBiG 2023 test (1 row) DetectoRS + LAEM Drawing the Same Bounding Box Twice? Coping Noisy Annotations in... code — Compare
UAVVaste (1 row) EfficientDet-D2 Waste detection in Pomerania: non-profit project for detecting... code — Compare
VisDrone- 1% labeled data (1 row) SSOD + Crop (L + U) Density Crop-guided Semi-supervised Object Detection in Aerial Images code — Compare
VisDrone - 10% labeled data (1 row) SSOD + Crop (L + U) Density Crop-guided Semi-supervised Object Detection in Aerial Images code — Compare
VisDrone - 5% labeled data (1 row) SSOD + Crop (L + U) Density Crop-guided Semi-supervised Object Detection in Aerial Images code — Compare
Watercolor2k (1 row) CDDMSL Semi-Supervised Domain Generalization for Object Detection via... code — Compare
Waymo 2D detection all_ns test (1 row) UniverseNet USB: Universal-Scale Object Detection Benchmark code — Compare
Fire and Smoke Dataset (0 rows) no rows in the archive — —

Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.

Libraries

Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.

Datasets archive 2025-07-28

332 datasets whose archive record lists this task, ordered by the archive's paper count. 30 shown of 332 until expanded.

COCO (Common Objects in Context)KITTInuScenesVisual GenomeGQALVISWaymo Open DatasetSUN RGB-DBDD100KManga10910,000 People - Human Pose Recognition DataFoggy CityscapesPASCAL3D+Kvasir-SEGPASCAL VOCLabelMeNoCapsCrowdHumanObjects365fMoWCityPersonsPASCAL VOC 2007PubLayNetKvasirLLVIPPASCAL VOC 2012 testABC DatasetUAVDTxViewMSCOCOUIEB11k HandsiSAIDApolloScapeVisDroneSPair-71kFSODPASCAL-PartScanRefer DatasetETHExDarkWoodScapeDAQUARUA-DETRACDVQAClipart1kSynscapesCOCO-OA2DH3DSIXrayA*3DOpen Images V4Watercolor2kJTAClear WeatherEgoHandsInteriorNetAVDComic2kVT5000UVOFlickrLogos-32ELEVATERTinyPersonCCPDINRIA PersonRADIATERecipeQACADCDSECOpenImages-v6COWCIP102MINOSSatlasSKU110KMSegDense FogEORSSDPlantDocArgoverse-HDModaNetAI-TODCoIRSeaDronesSeeRPCSceneNetMinneAppleMALFWashington RGB-DCADPGEN1 DetectionPreSILTTPLAUFDDWaterScenesGMOT-40Hyper-Kvasir DatasetPeopleArtSODA10MTJU-DHDTrashCanWiderPerson2024 AI City ChallengeCropAndWeedDeepScoresDescription Detection DatasetGRAZPEDWRI-DXSpaceNet 2Gun Detection DatasetPhenoBenchReDWeb-SUIISAPRICOTCops-RefDTTD-MobileDuke Breast Cancer MRIDUOHS-SODRIT-18SKU110K-RSODAGARBambooBigDetectionFreiburg GroceriesMobilityAidsPFN-PICPIDrayPS-BattlesYT-BBBAAI-VANJEECOCO-TasksEuroCity PersonsFATIIIT-AR-13KIndustRealLytro IllumMOR-UAVOpen Images V7Prophesee GEN4 DatasetVEDAICBCHJDatasetInsPLADOoDISRF100SoccerDBSpaceNet 1SYNTHIA-ALZeroWasteDeepSpaceYoloDatasetHeavy SnowfallKitchen ScenesMJU-WasteMUADMUSES: MUlti-SEnsor Semantic perception datasetS2TLDSeparated COCOSOD4SBTexBiGTimberSeg 1.0Aircraft Context DatasetCURE-TSDEAGLEFOD-ALOGO-NetNODOpenTTGamesParcel2D RealPoPArtPTLRailEye3D DatasetSpaceNet MVOIStream-51UIIS10KUP-COUNT360-SOD5,011 Images – Human Frontal face Data (Male)BdSLImsetC2A: Human Detection in Disaster ScenariosCOCO Object Detection VIPriors subsetCOMPASS-XPCPPE-5Deep PCBDGTA-SeaDronesSeeDGTA-VisDroneEndotect Polyp Segmentation Challenge DatasetEVD4UAVFDDB-360Forward-Looking Sonar Marine Debris DatasetsFracAtlasFSVOD-500Human-PartsINRIA-HorseKvasir-CapsuleLeukemiaAttriLight SnowfallLindenthal Camera TrapsMSDAOccluded COCOOktoberfest Food DatasetParasitic Egg Detection and Classification in Microscopic ImagesParcel3DPESMODSA-Det-100kSalient Object Subitizing DatasetSmartCityStreaksYoloDataset: labeled raw astronomical images for streaks detectionTBBRTNCR DatasetUnderwater Trash DetectionxView3-SARA Dataset of Multispectral Potato Plants ImagesActive TerahertzAODRawApron DatasetAquaTrashBBBC041Boreal Forest FireCattleCISOLCLADcranfield-synthetic-drone-detectionD2CityDARaiDeepTrashDGTA-CattleDrIFTDrinking Waste ClassificationDrone vs BirdEGO-CH-GazeEmbrapa ADD 256EXPO-HDFashionFailFGVDFire and Smoke DatasetGDITGLIB: image datasetHA-ViDHGPIndraEyeIRBFDLDDLiDAR-CSLocountLVVOMarine Microalgae Detection in Microscopy ImagesMD4KMETU-ALETMIS-Check DamMLGESTURE DATASETMuCeDmultiRAWNAONII-CU MAPDObject-Centric Stylized COCOPOGPoTATO: A Dataset for Analyzing Polarimetric Traces of Afloat Trash ObjectsPothole DatasetRadioGalaxyNETRASMDRetail50KS-ODv2SAR-AIRcraft-1.0SemanticSugarBeetsSFCHDSFU-HW-Objects-v1SIDODSOMPT22Songdo VisionSpiideo SoccerNet SynLocSSBI DatasetSSLSTN PLADStreet DatasetSunspotsYoloDataset: annotated solar images captured with smart telescopes (January 2023 - May 2024)TAMPARTea sickness - object detectionTimberVisionTiRODUAVBillboardsUAVVasteUDA-CHUSBVETRAVisual 3D shape matching datasetVizWiz-FewShotVME & CDSIVQA-OVYouTube-GDDAppleBBCH76AppleBBCH81Autorickshaw Image Dataset | Niche Vehicle DatasetBottles and Cups Dataset | Household ObjectsCherryBBCH72CherryBBCH81CodeSCANCODroneCornHubCracked Mobile Screen DataseteAppleScabElectronics Object Image Dataset | Computer PartsFacesInThingsHindi Text Image Dataset | Hindi in the wildHuman Palm and Gloves Dataset | Human Body Parts DatasetHumans in 3DIITKGP_Fence DatasetIndian Food Image DatasetInfinity Spills Basic DatasetIntersection Markings DatasetMasks Dataset | Unattended Mask ImagesMobile Phone Dataset | Smartphone & Feature PhoneMoroccan Monay datasetMS-EVS DatasetOCT5kOximeter Image Dataset | Medical Device ReadingPALMSPear640PFruitlet640RF100-VLSESYD DatasetShipRSImageNetStairs Image Dataset | Parts of House | IndoorSuitcase/Luggage Dataset Indoor Object ImageToloka Business ID RecognitionToloka WaterMeterstomato detectiontomato fruits detectionVolleyVisionWalnutData

Subtasks archive 2025-07-28

39 subtasks in the archive's task tree. 30 shown of 39 until expanded.

Parent tasks archive 2025-07-28

Most implemented papers archive 2025-07-28

30 shown of 4,657 papers with code (10,957 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.

  • 10 Dec 2015 484 repositories listed Syntology ran 230 of 377 samples · 147 unverified · 187 pointer-only (licence)
    Deep residual nets are foundations of our submissions to ILSVRC & COCO 2015 competitions, where we also won the 1st places on the tasks of ImageNet detection, ImageNet localization, COCO detection, and COCO segmentation.
  • 8 Apr 2018 311 repositories listed Syntology ran 18 of 124 samples · 106 unverified · 19 pointer-only (licence)
    At 320x320 YOLOv3 runs in 22 ms at 28.
  • 7 Aug 2017 234 repositories listed Syntology ran 11 of 11 samples · 0 unverified · 6 pointer-only (licence)
    Our novel Focal Loss focuses training on a sparse set of hard examples and prevents the vast number of easy negatives from overwhelming the detector during training.
  • 25 Dec 2016 231 repositories listed Syntology ran 16 of 60 samples · 44 unverified · 22 pointer-only (licence)
    On the 156 classes not in COCO, YOLO9000 gets 16.
  • 23 Apr 2020 223 repositories listed Syntology ran 24 of 184 samples · 160 unverified · 8 pointer-only (licence)
    There are a huge number of features which are said to improve Convolutional Neural Network (CNN) accuracy.
  • 8 Dec 2015 221 repositories listed Syntology ran 19 of 131 samples · 112 unverified · 5 pointer-only (licence)
    Experimental results on the PASCAL VOC, MS COCO, and ILSVRC datasets confirm that SSD has comparable accuracy to methods that utilize an additional object proposal step and is much faster, while providing a unified…
  • 4 Jun 2015 196 repositories listed Syntology ran 59 of 124 samples · 65 unverified · 42 pointer-only (licence)
    In this work, we introduce a Region Proposal Network (RPN) that shares full-image convolutional features with the detection network, thus enabling nearly cost-free region proposals.
  • 20 Mar 2017 179 repositories listed Syntology ran 42 of 140 samples · 98 unverified · 23 pointer-only (licence)
    Our approach efficiently detects objects in an image while simultaneously generating a high-quality segmentation mask for each instance.
  • 13 Jan 2018 159 repositories listed Syntology ran 85 of 111 samples · 26 unverified · 64 pointer-only (licence)
    In this paper we describe a new mobile architecture, MobileNetV2, that improves the state of the art performance of mobile models on multiple tasks and benchmarks as well as across a spectrum of different model sizes.
  • 17 Apr 2017 159 repositories listed Syntology ran 52 of 83 samples · 31 unverified · 48 pointer-only (licence)
    We present a class of efficient models called MobileNets for mobile and embedded vision applications.
  • 8 Jun 2015 144 repositories listed Syntology ran 80 of 148 samples · 68 unverified · 98 pointer-only (licence)
    A single neural network predicts bounding boxes and class probabilities directly from full images in one evaluation.
  • 17 Jun 2019 142 repositories listed Syntology ran 14 of 82 samples · 68 unverified
    In this paper, we introduce the various features of this toolbox.
  • 27 Nov 2019 123 repositories listed
    Neural networks have enabled state-of-the-art approaches to achieve incredible results on computer vision tasks such as object detection.
  • 2 Apr 2019 87 repositories listed Syntology ran 13 of 40 samples · 27 unverified · 18 pointer-only (licence)
    By eliminating the predefined set of anchor boxes, FCOS completely avoids the complicated computation related to anchor boxes such as calculating overlapping during training.
  • 5 Sep 2017 85 repositories listed Syntology ran 2 of 9 samples · 7 unverified · 6 pointer-only (licence)
    Squeeze-and-Excitation Networks formed the foundation of our ILSVRC 2017 classification submission which won first place and reduced the top-5 error to 2.
  • 9 Dec 2016 85 repositories listed Syntology ran 16 of 51 samples · 35 unverified · 11 pointer-only (licence)
    Feature pyramids are a basic component in recognition systems for detecting objects at different scales.
  • 17 Sep 2014 83 repositories listed Syntology ran 27 of 42 samples · 15 unverified · 20 pointer-only (licence)
    We propose a deep convolutional neural network architecture codenamed "Inception", which was responsible for setting the new state of the art for classification and detection in the ImageNet Large-Scale Visual…
  • 25 Mar 2021 80 repositories listed Syntology ran 108 of 207 samples · 99 unverified · 43 pointer-only (licence)
    This paper presents a new vision Transformer, called Swin Transformer, that capably serves as a general-purpose backbone for computer vision.
  • 16 Apr 2019 76 repositories listed Syntology ran 10 of 130 samples · 120 unverified
    We model an object as a single point --- the center point of its bounding box.
  • 22 Nov 2017 68 repositories listed Syntology ran 2 of 4 samples · 2 unverified · 2 pointer-only (licence)
    In this work, we study 3D object detection from RGB-D data in both indoor and outdoor scenes.
  • 6 May 2019 67 repositories listed Syntology ran 58 of 105 samples · 47 unverified · 46 pointer-only (licence)
    We achieve new state of the art results for mobile classification, detection and segmentation.
  • 20 Nov 2019 64 repositories listed Syntology ran 11 of 70 samples · 59 unverified · 3 pointer-only (licence)
    Model efficiency has become increasingly important in computer vision.
  • 24 Nov 2016 61 repositories listed Syntology ran 4 of 23 samples · 19 unverified · 4 pointer-only (licence)
    We present an approach to efficiently detect the 2D pose of multiple people in an image.
  • 19 Jun 2017 59 repositories listed Syntology ran 9 of 17 samples · 8 unverified · 13 pointer-only (licence)
    Its principled nature also enables us to identify methods for both training and attacking neural networks that are reliable and, in a certain sense, universal.
  • 11 Nov 2021 58 repositories listed Syntology ran 71 of 137 samples · 66 unverified · 73 pointer-only (licence)
    Our MAE approach is simple: we mask random patches of the input image and reconstruct the missing pixels.
  • 10 Jan 2022 54 repositories listed Syntology ran 54 of 80 samples · 26 unverified · 11 pointer-only (licence)
    The "Roaring 20s" of visual recognition began with the introduction of Vision Transformers (ViTs), which quickly superseded ConvNets as the state-of-the-art image classification model.
  • 10 Jan 2019 53 repositories listed
    COCO test-dev results are up to 41.
  • 20 May 2016 48 repositories listed Syntology ran 1 of 11 samples · 10 unverified
    In contrast to previous region-based detectors such as Fast/Faster R-CNN that apply a costly per-region subnetwork hundreds of times, our region-based detector is fully convolutional with almost all computation shared…
  • 7 Feb 2017 47 repositories listed
    We adapted the join-training scheme of Faster RCNN framework from Caffe to TensorFlow as a baseline implementation for object detection.
  • 17 Nov 2017 44 repositories listed Syntology ran 2 of 4 samples · 2 unverified · 4 pointer-only (licence)
    Accurate detection of objects in 3D point clouds is a central problem in many applications, such as autonomous navigation, housekeeping robots, and augmented/virtual reality.

Syntology lines on 27 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections