{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/recursive-nearest-agglomeration-rena-fast","title":"Recursive nearest agglomeration (ReNA): fast clustering for approximation of structured signals","arxiv_id":"1609.04608","date":"2016-09-15","proceeding":null,"authors":["Andrés Hoyos-Idrobo","Gaël Varoquaux","Jonas Kahn","Bertrand Thirion"],"abstract":"In this work, we revisit fast dimension reduction approaches, as with random\nprojections and random sampling. Our goal is to summarize the data to decrease\ncomputational costs and memory footprint of subsequent analysis. Such dimension\nreduction can be very efficient when the signals of interest have a strong\nstructure, such as with images. We focus on this setting and investigate\nfeature clustering schemes for data reductions that capture this structure. An\nimpediment to fast dimension reduction is that good clustering comes with large\nalgorithmic costs. We address it by contributing a linear-time agglomerative\nclustering scheme, Recursive Nearest Agglomeration (ReNA). Unlike existing fast\nagglomerative schemes, it avoids the creation of giant clusters. We empirically\nvalidate that it approximates the data as well as traditional\nvariance-minimizing clustering schemes that have a quadratic complexity. In\naddition, we analyze signal approximation with feature clustering and show that\nit can remove noise, improving subsequent analysis steps. As a consequence,\ndata reduction by clustering features with ReNA yields very fast and accurate\nmodels, enabling to process large datasets on budget. Our theoretical analysis\nis backed by extensive experiments on publicly-available data that illustrate\nthe computation efficiency and the denoising properties of the resulting\ndimension reduction scheme.","url_abs":"http://arxiv.org/abs/1609.04608v2","url_pdf":"http://arxiv.org/pdf/1609.04608v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"recursive-nearest-agglomeration-rena-fast","repo_url":"https://github.com/ahoyosid/ReNA","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"clustering","task_name":"Clustering"},{"task_slug":"denoising","task_name":"Denoising"},{"task_slug":"dimensionality-reduction","task_name":"Dimensionality Reduction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}