{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/scalable-mutual-information-estimation-using","title":"Scalable Mutual Information Estimation using Dependence Graphs","arxiv_id":"1801.09125","date":"2018-01-27","proceeding":null,"authors":["Morteza Noshad","Yu Zeng","Alfred O. Hero III"],"abstract":"The Mutual Information (MI) is an often used measure of dependency between\ntwo random variables utilized in information theory, statistics and machine\nlearning. Recently several MI estimators have been proposed that can achieve\nparametric MSE convergence rate. However, most of the previously proposed\nestimators have the high computational complexity of at least $O(N^2)$. We\npropose a unified method for empirical non-parametric estimation of general MI\nfunction between random vectors in $\\mathbb{R}^d$ based on $N$ i.i.d. samples.\nThe reduced complexity MI estimator, called the ensemble dependency graph\nestimator (EDGE), combines randomized locality sensitive hashing (LSH),\ndependency graphs, and ensemble bias-reduction methods. We prove that EDGE\nachieves optimal computational complexity $O(N)$, and can achieve the optimal\nparametric MSE rate of $O(1/N)$ if the density is $d$ times differentiable. To\nthe best of our knowledge EDGE is the first non-parametric MI estimator that\ncan achieve parametric MSE rates with linear time complexity. We illustrate the\nutility of EDGE for the analysis of the information plane (IP) in deep\nlearning. Using EDGE we shed light on a controversy on whether or not the\ncompression property of information bottleneck (IB) in fact holds for ReLu and\nother rectification functions in deep neural networks (DNN).","url_abs":"http://arxiv.org/abs/1801.09125v2","url_pdf":"http://arxiv.org/pdf/1801.09125v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"scalable-mutual-information-estimation-using","repo_url":"https://github.com/mrtnoshad/EDGE","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"information-plane","task_name":"Information Plane"},{"task_slug":"mutual-information-estimation","task_name":"Mutual Information Estimation"}],"methods":[{"method_slug":"relu","method_name":"ReLU"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1801.09125","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}