{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/chimp-efficient-lossless-floating-point","title":"Chimp: Efficient Lossless Floating Point Compression for Time Series Databases","arxiv_id":null,"date":"2022-07-01","proceeding":"Proceedings of the VLDB Endowment 2022 7","authors":["Panagiotis Liakos","Katia Papakonstantinopoulou","Yannis Kotidis"],"abstract":"Applications in diverse domains such as astronomy, economics\r\nand industrial monitoring, increasingly press the need for analyz\u0002ing massive collections of time series data. The sheer size of the\r\nlatter hinders our ability to efficiently store them and also yields\r\nsignificant storage costs. Applying general purpose compression al\u0002gorithms would effectively reduce the size of the data, at the expense\r\nof introducing significant computational overhead. Time Series\r\nManagement Systems that have emerged to address the challenge\r\nof handling this overwhelming amount of information, cannot suf\u0002fer the ingestion rate restrictions that such compression algorithms\r\nwould cause. Data points are usually encoded using faster, stream\u0002ing compression approaches. However, the techniques that contem\u0002porary systems use do not fully utilize the compression potential\r\nof time series data, with implications in both storage requirements\r\nand access times. In this work, we propose a novel streaming com\u0002pression algorithm, suitable for floating point time series data. We\r\nempirically establish properties exhibited by a diverse set of time\r\nseries and harness these features in our proposed encodings. Our\r\nexperimental evaluation demonstrates that our approach readily\r\noutperforms competing techniques, attaining compression ratios\r\nthat are competitive with slower general purpose algorithms, and\r\non average around 50% of the space required by state-of-the-art\r\nstreaming approaches. Moreover, our algorithm outperforms all\r\nearlier techniques with regards to both compression and access time,\r\noffering a significantly improved trade-off between space and speed.\r\nThe aforementioned benefits of our approach –in terms of all space\r\nrequirements, compression time and read access– significantly im\u0002prove the efficiency in which we can store and analyze time series\r\ndata.","url_abs":"https://dl.acm.org/doi/10.14778/3551793.3551852","url_pdf":"https://dl.acm.org/doi/pdf/10.14778/3551793.3551852","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"chimp-efficient-lossless-floating-point","repo_url":"https://github.com/xolvio/chimp","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"astronomy","task_name":"Astronomy"},{"task_slug":"time-series-1","task_name":"Time Series"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}