Papers › GPUNet: Searching the Deployable Convolution Neural Networks for GPUs

GPUNet: Searching the Deployable Convolution Neural Networks for GPUs

26 Apr 2022arXiv:2205.00841archive 2025-07-28

Linnan Wang, Chenhan Yu, Satish Salian, Slawomir Kierat, Szymon Migacz, Alex Fit Florea

Customizing Convolution Neural Networks (CNN) for production use has been a challenging task for DL practitioners. This paper intends to expedite the model customization with a model hub that contains the optimized models tiered by their inference latency using Neural Architecture Search (NAS). To achieve this goal, we build a distributed NAS system to search on a novel search space that consists of prominent factors to impact latency and accuracy. Since we target GPU, we name the NAS optimized models as GPUNet, which establishes a new SOTA Pareto frontier in inference latency and accuracy. Within 1$ms$, GPUNet is 2x faster than EfficientNet-X and FBNetV3 with even better accuracy. We also validate GPUNet on detection tasks, and GPUNet consistently outperforms EfficientNet-X and FBNetV3 on COCO detection tasks in both latency and accuracy. All of these data validate that our NAS system is effective and generic to handle different design tasks. With this NAS system, we expand GPUNet to cover a wide range of latency targets such that DL practitioners can deploy our models directly in different scenarios.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Neural Architecture Search

1 archive task tag without a task page not shown.

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Neural Architecture Search ImageNet GPUNet-D3 FLOPs 15.6G #3 of 135 Archive leaderboard report
Neural Architecture Search ImageNet GPUNet-D3 Params 19M #3 of 135 Archive leaderboard report
Neural Architecture Search ImageNet GPUNet-D3 Top-1 Error Rate 16.4 #3 of 135 Archive leaderboard report
Neural Architecture Search ImageNet GPUNet-D1 FLOPs 3.66G #4 of 135 Archive leaderboard report
Neural Architecture Search ImageNet GPUNet-D1 Params 10.6M #4 of 135 Archive leaderboard report
Neural Architecture Search ImageNet GPUNet-D1 Top-1 Error Rate 17.5 #4 of 135 Archive leaderboard report
Neural Architecture Search ImageNet GPUNet-D0 FLOPs 720M #28 of 135 Archive leaderboard report
Neural Architecture Search ImageNet GPUNet-D0 Params 6.2M #28 of 135 Archive leaderboard report
Neural Architecture Search ImageNet GPUNet-D0 Top-1 Error Rate 20.3 #28 of 135 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Convolution

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections