{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/crosscat-a-fully-bayesian-nonparametric","title":"CrossCat: A Fully Bayesian Nonparametric Method for Analyzing Heterogeneous, High Dimensional Data","arxiv_id":"1512.01272","date":"2015-12-03","proceeding":null,"authors":["Vikash Mansinghka","Patrick Shafto","Eric Jonas","Cap Petschulat","Max Gasner","Joshua B. Tenenbaum"],"abstract":"There is a widespread need for statistical methods that can analyze\nhigh-dimensional datasets with- out imposing restrictive or opaque modeling\nassumptions. This paper describes a domain-general data analysis method called\nCrossCat. CrossCat infers multiple non-overlapping views of the data, each\nconsisting of a subset of the variables, and uses a separate nonparametric\nmixture to model each view. CrossCat is based on approximately Bayesian\ninference in a hierarchical, nonparamet- ric model for data tables. This model\nconsists of a Dirichlet process mixture over the columns of a data table in\nwhich each mixture component is itself an independent Dirichlet process mixture\nover the rows; the inner mixture components are simple parametric models whose\nform depends on the types of data in the table. CrossCat combines strengths of\nmixture modeling and Bayesian net- work structure learning. Like mixture\nmodeling, CrossCat can model a broad class of distributions by positing latent\nvariables, and produces representations that can be efficiently conditioned and\nsampled from for prediction. Like Bayesian networks, CrossCat represents the\ndependencies and independencies between variables, and thus remains accurate\nwhen there are multiple statistical signals. Inference is done via a scalable\nGibbs sampling scheme; this paper shows that it works well in practice. This\npaper also includes empirical results on heterogeneous tabular data of up to 10\nmillion cells, such as hospital cost and quality measures, voting records,\nunemployment rates, gene expression measurements, and images of handwritten\ndigits. CrossCat infers structure that is consistent with accepted findings and\ncommon-sense knowledge in multiple domains and yields predictive accuracy\ncompetitive with generative, discriminative, and model-free alternatives.","url_abs":"http://arxiv.org/abs/1512.01272v1","url_pdf":"http://arxiv.org/pdf/1512.01272v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"crosscat-a-fully-bayesian-nonparametric","repo_url":"https://github.com/ririw/data-talk-2016-06","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"bayesian-inference","task_name":"Bayesian Inference"},{"task_slug":"common-sense-reasoning","task_name":"Common Sense Reasoning"},{"task_slug":"high","task_name":"Vocal Bursts Intensity Prediction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1512.01272","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}