Browse State-of-the-Art › Data Compression
Data Compression
131 papers with code · 0 benchmarks · 0 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
No dataset record in the archive lists this task.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 131 papers with code (459 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
9 Mar 2016 28 repositories listed Syntology ran 0 of 17 samples · 17 unverifiedIn this paper, we describe a scalable end-to-end tree boosting system called XGBoost, which is used widely by data scientists to achieve state-of-the-art results on many machine learning challenges.
-
26 Jun 2023 6 repositories listed Syntology ran 8 of 23 samples · 15 unverifiedDecoding the linguistic intricacies of the genome is a crucial problem in biology, and pre-trained foundational models such as DNABERT and Nucleotide Transformer have made significant strides in this area.
-
9 Jul 2024 3 repositories listed Syntology ran 4 of 5 samples · 1 unverifiedInspired by the information compression nature of LLMs, we uncover an ``entropy law'' that connects LLM performance with data compression ratio and first-epoch training loss, which reflect the information redundancy of…
-
14 Feb 2022 3 repositories listed Syntology ran 0 of 16 samples · 16 unverifiedNeural compression is the application of neural networks and other machine learning methods to data compression.
-
29 Sep 2021 3 repositories listedNeural data compression based on nonlinear transform coding has made great progress over the last few years, mainly due to improvements in prior models, quantization methods and nonlinear transforms.
-
26 Jun 2017 3 repositories listedThere is a rich literature on approximating the unknown manifold, and on exploiting such approximations in clustering, data compression, and prediction.
-
24 Mar 2025 2 repositories listedFinally, through the dynamic combination of UELC and VRCM, we achieve lossy compression, lossless compression, variable rate and complexity within a unified framework.
-
28 Feb 2024 2 repositories listed Syntology ran 3 of 6 samples · 3 unverified · 6 pointer-only (licence)Tokenization is a foundational step in natural language processing (NLP) tasks, bridging raw text and language models.
-
30 Jan 2024 2 repositories listed Syntology ran 5 of 5 samples · 0 unverifiedAs LLMs continue to advance, it is crucial to develop diverse and appropriate metrics for their evaluation.
-
19 Dec 2023 2 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)In this paper, we introduce a compact scene representation organizing the parameters of 3D Gaussian Splatting (3DGS) into a 2D grid with local homogeneity, ensuring a drastic reduction in storage requirements without…
-
28 Jun 2023 2 repositories listedModern Bayesian inference involves a mixture of computational techniques for estimating, validating, and drawing conclusions from probabilistic models as part of principled workflows for data analysis.
-
7 Mar 2023 2 repositories listedThanks to its concise but effective summarization, CubeScope can also detect the sudden appearance of anomalies and identify the types of anomalies that occur in practice.
-
24 Jan 2023 2 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedMulti-view image compression plays a critical role in 3D-related applications.
-
9 Nov 2022 2 repositories listedWe present a new convolution layer for deep learning architectures which we call QuadConv -- an approximation to continuous convolution via quadrature.
-
20 Jul 2022 2 repositories listed Syntology ran 4 of 7 samples · 3 unverifiedDataset Condensation is a newly emerging technique aiming at learning a tiny dataset that captures the rich information encoded in the original dataset.
-
7 Jan 2022 2 repositories listedWe show that BottleFit decreases power consumption and latency respectively by up to 49% and 89% with respect to (w.
-
5 Jan 2022 2 repositories listedEntropy coding is the backbone data compression.
-
23 Nov 2021 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)By contrast, this paper makes the first attempt at an algorithm for sandwiching the R-D function of a general (not necessarily discrete) source requiring only i.
-
21 Aug 2021 2 repositories listed Syntology ran 1 of 17 samples · 16 unverified · 1 pointer-only (licence)There has been much interest in deploying deep learning algorithms on low-powered devices, including smartphones, drones, and medical sensors.
-
21 May 2021 2 repositories listed Syntology ran 1 of 2 samples · 1 unverified · 2 pointer-only (licence)This work attempts to provide a plausible theoretical framework that aims to interpret modern deep (convolutional) networks from the principles of data compression and discriminative representation.
-
12 Nov 2019 2 repositories listed Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)In this paper, we present a new angle to analyze the quantization error, which decomposes the quantization error into norm error and direction error.
-
10 Jun 2025 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedThe site conditions that make astronomical observatories in space and on the ground so desirable -- cold and dark -- demand a physical remoteness that leads to limited data transmission capabilities.
-
27 May 2025 1 repository listedApproximation of a target probability distribution using a finite set of points is a problem of fundamental importance, arising in cubature, data compression, and optimisation.
-
26 May 2025 1 repository listedLong Context Understanding (LCU) is a critical area for exploration in current large language models (LLMs).
-
23 May 2025 1 repository listedWe also created a synthetic teacher-student setup to investigate compression in a controlled continuous setting.
-
19 May 2025 1 repository listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)In recent years, dataset distillation has provided a reliable solution for data compression, where models trained on the resulting smaller synthetic datasets achieve performance comparable to those trained on the…
-
10 May 2025 1 repository listedSpecifically, we compress historical data by searching for a class semantic embedding in the conditional space of the pre-trained diffusion model, which can guide the model to replay data with fine-grained pixel…
-
3 May 2025 1 repository listedMost existing 3D Gaussian Splatting (3DGS) compression schemes focus on producing compact 3DGS representation via implicit data embedding.
-
31 Mar 2025 1 repository listedNonlinear matrix decomposition (NMD) with the ReLU function, denoted ReLU-NMD, is the following problem: given a sparse, nonnegative matrix X and a factorization rank r, identify a rank-r matrix Θ such that X≈max(0,Θ).
-
5 Jan 2025 1 repository listedWhile deep learning facilitates joint design of the compression mapping along with encoding and inference rules, existing learned compression mechanisms are static, and struggle in adapting their resolution to changes…
Syntology lines on 15 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections