{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/data-smashing","title":"Data Smashing","arxiv_id":"1401.0742","date":"2014-01-03","proceeding":null,"authors":["Ishanu Chattopadhyay","Hod Lipson"],"abstract":"Investigation of the underlying physics or biology from empirical data\nrequires a quantifiable notion of similarity - when do two observed data sets\nindicate nearly identical generating processes, and when they do not. The\ndiscriminating characteristics to look for in data is often determined by\nheuristics designed by experts, $e.g.$, distinct shapes of \"folded\" lightcurves\nmay be used as \"features\" to classify variable stars, while determination of\npathological brain states might require a Fourier analysis of brainwave\nactivity. Finding good features is non-trivial. Here, we propose a universal\nsolution to this problem: we delineate a principle for quantifying similarity\nbetween sources of arbitrary data streams, without a priori knowledge, features\nor training. We uncover an algebraic structure on a space of symbolic models\nfor quantized data, and show that such stochastic generators may be added and\nuniquely inverted; and that a model and its inverse always sum to the generator\nof flat white noise. Therefore, every data stream has an anti-stream: data\ngenerated by the inverse model. Similarity between two streams, then, is the\ndegree to which one, when summed to the other's anti-stream, mutually\nannihilates all statistical structure to noise. We call this data smashing. We\npresent diverse applications, including disambiguation of brainwaves pertaining\nto epileptic seizures, detection of anomalous cardiac rhythms, and\nclassification of astronomical objects from raw photometry. In our examples,\nthe data smashing principle, without access to any domain knowledge, meets or\nexceeds the performance of specialized algorithms tuned by domain experts.","url_abs":"http://arxiv.org/abs/1401.0742v1","url_pdf":"http://arxiv.org/pdf/1401.0742v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"data-smashing","repo_url":"https://github.com/pslii/datasmash","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}