Papers › TResNet: High Performance GPU-Dedicated Architecture

TResNet: High Performance GPU-Dedicated Architecture

30 Mar 2020arXiv:2003.13630archive 2025-07-28

Tal Ridnik, Hussam Lawen, Asaf Noy, Emanuel Ben Baruch, Gilad Sharir, Itamar Friedman

Many deep learning models, developed in recent years, reach higher ImageNet accuracy than ResNet50, with fewer or comparable FLOPS count. While FLOPs are often seen as a proxy for network efficiency, when measuring actual GPU training and inference throughput, vanilla ResNet50 is usually significantly faster than its recent competitors, offering better throughput-accuracy trade-off. In this work, we introduce a series of architecture modifications that aim to boost neural networks' accuracy, while retaining their GPU training and inference efficiency. We first demonstrate and discuss the bottlenecks induced by FLOPs-optimizations. We then suggest alternative designs that better utilize GPU structure and assets. Finally, we introduce a new family of GPU-dedicated models, called TResNet, which achieve better accuracy and efficiency than previous ConvNets. Using a TResNet model, with similar GPU throughput to ResNet50, we reach 80.8 top-1 accuracy on ImageNet. Our TResNet models also transfer well and achieve state-of-the-art accuracy on competitive single-label classification datasets such as Stanford cars (96.0%), CIFAR-10 (99.0%), CIFAR-100 (91.5%) and Oxford-Flowers (99.1%). They also perform well on multi-label classification and object detection tasks. Implementation is available at: https://github.com/mrT23/TResNet.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

mrT23/TResNet officialmentioned in papermentioned on GitHubpytorchApache-2.0 report
rwightman/pytorch-image-models officialmentioned in papermentioned on GitHubpytorch report
Alibaba-MIIL/TResNet mentioned on GitHubpytorchApache-2.0 report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Fine-Grained Image ClassificationGeneral ClassificationImage ClassificationMUlTI-LABEL-ClASSIFICATIONMulti-Label ClassificationObject DetectionVocal Bursts Intensity Predictionobject-detection

1 archive task tag without a task page not shown.

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Fine-Grained Image Classification Oxford 102 Flowers TResNet-L Accuracy 99.1% #6 of 25 Archive leaderboard report
Image Classification CIFAR-10 TResNet-XL Percentage correct 99 #22 of 265 Archive leaderboard report
Image Classification CIFAR-100 TResNet-L-V2 Percentage correct 92.6 #12 of 211 Archive leaderboard report
Image Classification Flowers-102 TResNet-L Accuracy 99.1% #14 of 52 Archive leaderboard report
Image Classification ImageNet TResNet-XL Number of params 77M #329 of 1060 Archive leaderboard report
Image Classification ImageNet TResNet-XL Top 1 Accuracy 84.3% #329 of 1060 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Introduced by this paper: TResNet

1x1 ConvolutionAnti-Alias DownsamplingAutoAugmentAverage PoolingBatch NormalizationColorJitterConvolutionCutoutDense ConnectionsGlobal Average PoolingInPlace-ABNLSTMLabel SmoothingReLUResidual ConnectionSigmoid ActivationSqueeze-and-Excitation BlockTResNetTanh ActivationWeight Decay

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections