Browse State-of-the-Art › Imputation
Imputation
479 papers with code · 4 benchmarks · 12 datasets archive 2025-07-28
Substituting missing data with values according to some criteria.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
4 leaderboard tables shown for this task, 4 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| Adult (1 row) | ANN | Missing Data Imputation for Supervised Learning | code | — | Compare |
| HMNIST (1 row) | GP-VAE (B-NLST) | Seq2Tens: An Efficient Representation of Sequences by Low-Rank... | code | — | Compare |
| PhysioNet Challenge 2012 (1 row) | GP-VAE (B-NLST) | Seq2Tens: An Efficient Representation of Sequences by Low-Rank... | code | — | Compare |
| Sprites (1 row) | GP-VAE (B-NLST) | Seq2Tens: An Efficient Representation of Sequences by Low-Rank... | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
12 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 479 papers with code (1,237 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
7 Jun 2018 8 repositories listed Syntology ran 2 of 8 samples · 6 unverified · 1 pointer-only (licence)Accordingly, we call our method Generative Adversarial Imputation Nets (GAIN).
-
6 Jun 2016 7 repositories listed Syntology ran 1 of 3 samples · 2 unverified · 2 pointer-only (licence)Multivariate time series data in practical applications, such as health care, geoscience, and biology, are characterized by a variety of missing values.
-
22 Oct 2022 6 repositories listed Syntology ran 1 of 4 samples · 3 unverified · 1 pointer-only (licence)Under each task, we describe the most recent developments in classical and deep learning methods and discuss their advantages and disadvantages.
-
6 Oct 2020 6 repositories listed Syntology ran 5 of 32 samples · 27 unverified · 2 pointer-only (licence)In this work we propose for the first time a transformer-based framework for unsupervised representation learning of multivariate time series.
-
8 Mar 2019 6 repositories listed Syntology ran 0 of 4 samples · 4 unverifiedIn this work, we introduce a general probabilistic model that describes sparse high dimensional imaging data as being generated by a deep non-linear embedding.
-
6 Feb 2024 5 repositories listedThis survey aims to serve as a valuable resource for researchers and practitioners in the field of time series analysis and missing data imputation tasks.
-
18 Sep 2023 5 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedThrough empirical evaluation across the benchmark, we demonstrate that our approach outperforms deep-learning generation methods in data generation tasks and remains competitive in data imputation.
-
30 May 2023 5 repositories listed Syntology ran 0 of 21 samples · 21 unverifiedPyPOTS is an open-source Python library dedicated to data mining and analysis on multivariate partially-observed time series, i.
-
7 Jul 2021 5 repositories listed Syntology ran 10 of 19 samples · 9 unverified · 1 pointer-only (licence)In this paper, we propose Conditional Score-based Diffusion models for Imputation (CSDI), a novel time series imputation method that utilizes score-based diffusion models conditioned on observed data.
-
27 May 2018 5 repositories listed Syntology ran 2 of 4 samples · 2 unverifiedIt is ubiquitous that time series contains many missing values.
-
18 Jun 2024 4 repositories listedDespite the development of numerous deep learning algorithms for time series imputation, the community lacks standardized and comprehensive benchmark platforms to effectively evaluate imputation performance across…
-
8 Mar 2024 4 repositories listed Syntology ran 8 of 9 samples · 1 unverified · 4 pointer-only (licence)Such architectures impose hard constraints on the model.
-
9 Jul 2019 4 repositories listed Syntology ran 0 of 12 samples · 12 unverifiedMultivariate time series with missing values are common in areas such as healthcare and finance, and have grown in number and complexity over the years.
-
1 Jun 2015 4 repositories listedWe used Tiled Convolutional Neural Networks (tiled CNNs) on 20 standard datasets to learn high-level features from the individual and compound GASF-GADF-MTF images.
-
5 Oct 2022 3 repositories listedTimesBlock can discover the multi-periodicity adaptively and extract the complex temporal variations from transformed 2D tensors by a parameter-efficient inception block.
-
17 Feb 2022 3 repositories listed Syntology ran 3 of 10 samples · 7 unverifiedMissing data in time series is a pervasive problem that puts obstacles in the way of advanced analysis.
-
29 Jan 2022 3 repositories listedRandom forests are considered one of the best out-of-the-box classification and regression algorithms due to their high level of predictive performance with relatively little tuning.
-
23 Oct 2020 3 repositories listedDeveloping deep learning methods on EHRs data is critical for personalized treatment, precise diagnosis and medical management.
-
20 May 2020 3 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedPulmonary opacification is the inflammation in the lungs caused by many respiratory ailments, including the novel corona virus disease 2019 (COVID-19).
-
14 Oct 2019 3 repositories listedIn this paper, we propose a Bayesian temporal factorization (BTF) framework for modeling multidimensional time series -- in particular spatiotemporal data -- in the presence of missing values.
-
17 May 2019 3 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedIn order to integrate uncertainty estimates into deep time-series modelling, Kalman Filters (KFs) (Kalman et al., 1960) have been integrated with deep learning models, however, such approaches typically rely on…
-
19 Feb 2019 3 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)A striking result is that the widely-used method of imputing with a constant, such as the mean prior to learning is consistent when missing values are not informative.
-
10 Jul 2018 3 repositories listed Syntology ran 0 of 11 samples · 11 unverifiedVariational autoencoders (VAEs), as well as other generative models, have been shown to be efficient and accurate for capturing the latent structure of vast amounts of complex high-dimensional data.
-
6 Jun 2018 3 repositories listed Syntology ran 0 of 7 samples · 7 unverifiedWe propose a single neural probabilistic model based on variational autoencoder that can be conditioned on an arbitrary subset of observed features and then sample the remaining features in "one shot".
-
8 May 2017 3 repositories listedMissing data is a significant problem impacting all domains.
-
22 Sep 2016 3 repositories listed Syntology ran 4 of 4 samples · 0 unverifiedWe show that many existing neural network architectures can be made input-convex with a minor modification, and develop specialized optimization algorithms tailored to this setting.
-
16 Jun 2025 2 repositories listedAccurate weather forecasts are essential for supporting a wide range of activities and decision-making processes, as well as mitigating the impacts of adverse weather events.
-
17 Feb 2025 2 repositories listed Syntology ran 0 of 19 samples · 19 unverifiedWe introduce a simple method for probabilistic predictions on tabular data based on Large Language Models (LLMs) called JoLT (Joint LLM Process for Tabular data).
-
21 Jan 2025 2 repositories listed Syntology ran 1 of 5 samples · 4 unverifiedSynthetic data generation for tabular datasets must balance fidelity, efficiency, and versatility to meet the demands of real-world applications.
-
20 Dec 2024 2 repositories listedHere, we develop and apply an MPS-based algorithm, MPSTime, for learning a joint probability distribution underlying an observed time-series dataset, and show how it can be used to tackle important time-series ML…
Syntology lines on 20 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections