{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/ultra-scalable-spectral-clustering-and","title":"Ultra-Scalable Spectral Clustering and Ensemble Clustering","arxiv_id":"1903.01057","date":"2019-03-04","proceeding":null,"authors":["Dong Huang","Chang-Dong Wang","Jian-Sheng Wu","Jian-Huang Lai","Chee-Keong Kwoh"],"abstract":"This paper focuses on scalability and robustness of spectral clustering for\nextremely large-scale datasets with limited resources. Two novel algorithms are\nproposed, namely, ultra-scalable spectral clustering (U-SPEC) and\nultra-scalable ensemble clustering (U-SENC). In U-SPEC, a hybrid representative\nselection strategy and a fast approximation method for K-nearest\nrepresentatives are proposed for the construction of a sparse affinity\nsub-matrix. By interpreting the sparse sub-matrix as a bipartite graph, the\ntransfer cut is then utilized to efficiently partition the graph and obtain the\nclustering result. In U-SENC, multiple U-SPEC clusterers are further integrated\ninto an ensemble clustering framework to enhance the robustness of U-SPEC while\nmaintaining high efficiency. Based on the ensemble generation via multiple\nU-SEPC's, a new bipartite graph is constructed between objects and base\nclusters and then efficiently partitioned to achieve the consensus clustering\nresult. It is noteworthy that both U-SPEC and U-SENC have nearly linear time\nand space complexity, and are capable of robustly and efficiently partitioning\nten-million-level nonlinearly-separable datasets on a PC with 64GB memory.\nExperiments on various large-scale datasets have demonstrated the scalability\nand robustness of our algorithms. The MATLAB code and experimental data are\navailable at https://www.researchgate.net/publication/330760669.","url_abs":"http://arxiv.org/abs/1903.01057v2","url_pdf":"http://arxiv.org/pdf/1903.01057v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"clustering","task_name":"Clustering"},{"task_slug":"imagedocument-clustering","task_name":"Image/Document Clustering"}],"methods":[{"method_slug":"large-scale-spectral-clustering","method_name":"Large-scale spectral clustering"},{"method_slug":"spectral-clustering","method_name":"Spectral Clustering"},{"method_slug":"pc","method_name":"pc"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/image-document-clustering-on-pendigits","task":"Image/Document Clustering","dataset":"pendigits","model":"U-SPEC","rank_in_archive_order":3,"of":7,"metrics":{"NMI":"0.803","runtime (s)":"1.01"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1903.01057","atlas_url":"https://app.syntology.ai/?focus=1903.01057","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}