{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/using-feature-grouping-as-a-stochastic","title":"Feature Grouping as a Stochastic Regularizer for High-Dimensional Structured Data","arxiv_id":"1807.11718","date":"2018-07-31","proceeding":null,"authors":["Sergul Aydore","Bertrand Thirion","Gael Varoquaux"],"abstract":"In many applications where collecting data is expensive, for example\nneuroscience or medical imaging, the sample size is typically small compared to\nthe feature dimension. It is challenging in this setting to train expressive,\nnon-linear models without overfitting. These datasets call for intelligent\nregularization that exploits known structure, such as correlations between the\nfeatures arising from the measurement device. However, existing structured\nregularizers need specially crafted solvers, which are difficult to apply to\ncomplex models. We propose a new regularizer specifically designed to leverage\nstructure in the data in a way that can be applied efficiently to complex\nmodels. Our approach relies on feature grouping, using a fast clustering\nalgorithm inside a stochastic gradient descent loop: given a family of feature\ngroupings that capture feature covariations, we randomly select these groups at\neach iteration. We show that this approach amounts to enforcing a denoising\nregularizer on the solution. The method is easy to implement in many model\narchitectures, such as fully connected neural networks, and has a linear\ncomputational cost. We apply this regularizer to a real-world fMRI dataset and\nthe Olivetti Faces datasets. Experiments on both datasets demonstrate that the\nproposed approach produces models that generalize better than those trained\nwith conventional regularizers, and also improves convergence speed.","url_abs":"http://arxiv.org/abs/1807.11718v2","url_pdf":"http://arxiv.org/pdf/1807.11718v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"using-feature-grouping-as-a-stochastic","repo_url":"https://github.com/sergulaydore/Feature-Grouping-Regularizer","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"clustering","task_name":"Clustering"},{"task_slug":"denoising","task_name":"Denoising"},{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"high","task_name":"Vocal Bursts Intensity Prediction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}