{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/universal-indexes-for-highly-repetitive","title":"Universal Indexes for Highly Repetitive Document Collections","arxiv_id":"1604.08897","date":"2016-05-24","proceeding":null,"authors":["Claude Francisco","Fariña Antonio","Martínez-Prieto Miguel A.","Navarro Gonzalo"],"abstract":"Indexing highly repetitive collections has become a relevant problem with the\nemergence of large repositories of versioned documents, among other\napplications. These collections may reach huge sizes, but are formed mostly of\ndocuments that are near-copies of others. Traditional techniques for indexing\nthese collections fail to properly exploit their regularities in order to\nreduce space.\n  We introduce new techniques for compressing inverted indexes that exploit\nthis near-copy regularity. They are based on run-length, Lempel-Ziv, or grammar\ncompression of the differential inverted lists, instead of the usual practice\nof gap-encoding them. We show that, in this highly repetitive setting, our\ncompression methods significantly reduce the space obtained with classical\ntechniques, at the price of moderate slowdowns. Moreover, our best methods are\nuniversal, that is, they do not need to know the versioning structure of the\ncollection, nor that a clear versioning structure even exists.\n  We also introduce compressed self-indexes in the comparison. These are\ndesigned for general strings (not only natural language texts) and represent\nthe text collection plus the index structure (not an inverted index) in\nintegrated form. We show that these techniques can compress much further, using\na small fraction of the space required by our new inverted indexes. Yet, they\nare orders of magnitude slower.","url_abs":"http://arxiv.org/abs/1604.08897v2","url_pdf":"http://arxiv.org/pdf/1604.08897v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"universal-indexes-for-highly-repetitive","repo_url":"https://github.com/migumar2/uiHRDC","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}