Papers › Mining limited data for more robust and generalized ML models

Mining limited data for more robust and generalized ML models

1 Dec 2021Association for the Advancement of Artificial Intelligence 2021 12archive 2025-07-28

Jun Yu, Hao Chang, Keda Lu

For the last ten years, the success of deep learning comes from big data and large models. In recent years, the architecture of neural networks has been very mature. It’s more efficient to look for ways improving the data based a fixed neural network architecture. Similarly, in terms of robust machine learning, defense methods based on deep learning models have been proposed to mitigate potential threats about adversarial samples, but most of them pursue high-performance models under fixed constraints and dataset. Therefore, how to construct universal and effective dataset to train robust models is still a problem to be explored. In this paper, we consider to generate a more robust and efficient machine learning model by mining limited data. In detail, we proposed Robust TrivialAugment(RTA) and Iterative Search to obtain better dataset to make the trained model achieve better performance with the same amount of data. Moreover, we have applied the proposed methods to competition AAAI2022 DataCentric Robust Learning on ML Models that is organized by Alibaba on the Tianchi platform and won top 10 in 3691 teams. Code is available at https://github.com/wujiekd/RTAIterative-Search-AAAI2022.

PaperPDFCode

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections