Papers › Rescaling Large Datasets Based on Validation Outcomes of a Pre-trained Network

Rescaling Large Datasets Based on Validation Outcomes of a Pre-trained Network

1 Jul 2024Pattern Recognition Letters 2024 7archive 2025-07-28

Thanh Tuan NGUYEN, Thanh Phuong Nguyen

In fact, several categories in a large dataset are not difficult for recent advanced deep neural networks to recognize. Eliminating them for a challenging smaller subset will assist the early network proposals in taking a quick trial of verification. To this end, we propose an efficient rescaling method based on the validation outcomes of a pre-trained model. Firstly, we will take out the sensitive images of the lowest-accuracy classes of the validation outcomes. Each of such images is then considered to identify which label it was confused with. Gathering the lowest-accuracy classes along with the most confused ones can produce a smaller subset with a higher challenge for quick validation of an early network draft. Finally, a rescaling application is introduced to rescale two popular large datasets (ImageNet and Places365) for different tiny subsets (i.e., ReINΩ and RePLΩ respectively). Experiments for image classification have proved that neural networks obtaining good performance on the original datasets also achieve good results on their rescaled subsets. For instance, MobileNetV1 and MobileNetV2 with 70.6% and 72% on ImageNet respectively obtained 46.53% and 47.47% on its small subset ReIN30, which only contains about 39000 images. It can be observed that the better performance of MobileNetV2 on ImageNet correspondingly leads to the better rate on its rescaled subset. Appropriately, utilizing these rescaled sets would help researchers save time and computational costs in the way of designing deep neural architectures.

PaperPDFCode

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Image Classificationimage-classification

Datasets

Introduced by this paper, per the archive.

ReINs and RePLs: Challenging, small datasets for quick validations of designing deep neural networks for image classification

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

1x1 ConvolutionAverage PoolingBatch NormalizationConvolutionDense ConnectionsDepthwise ConvolutionDepthwise Separable ConvolutionGlobal Average PoolingInverted Residual BlockMobileNetV1Pointwise ConvolutionReLUSoftmax

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections