Browse State-of-the-Art › Data Integration
Data Integration
130 papers with code · 0 benchmarks · 7 datasets archive 2025-07-28
Data integration (also called information integration) is the process of consolidating data from a set of heterogeneous data sources into a single uniform data set (materialized integration) or view on the data (virtual integration). Data integration pipelines involve subtasks such as schema matching, table annotation, entity resolution, value normalization, data cleansing, and data fusion. Application domains of data integration include data warehousing, data lakes, and knowledge base consolidation. Surveys on Data integration:
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
7 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
3 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 130 papers with code (431 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
3 Jan 2025 3 repositories listedImage segmentation is a vital task for providing human assistance and enhancing autonomy in our daily lives.
-
10 Jun 2024 2 repositories listedWe introduce Semantic Parsing in Contextual Environments (SPICE), a task designed to enhance artificial agents' contextual awareness by integrating multimodal inputs with prior contexts.
-
22 May 2024 2 repositories listed Syntology ran 5 of 6 samples · 1 unverifiedTraditional time series forecasting methods predominantly rely on historical data patterns, neglecting external interventions that significantly shape future dynamics.
-
15 Dec 2022 2 repositories listed Syntology ran 3 of 6 samples · 3 unverified · 6 pointer-only (licence)Many, if not most, systems of interest in science are naturally described as nonlinear dynamical systems.
-
4 Nov 2020 2 repositories listedIdentifying potential drug targets using metabolic modeling requires integrating multiple modeling methods and heterogenous biological datasets, which can be challenging without sophisticated tools.
-
17 Apr 2017 2 repositories listedWe introduce a novel Bayesian hybrid matrix factorisation model (HMF) for data integration, based on combining multiple matrix factorisation methods, that can be used for in- and out-of-matrix prediction of missing…
-
24 May 2016 2 repositories listedMIMIC-III (‘Medical Information Mart for Intensive Care’) is a large, single-center database comprising information relating to patients admitted to critical care units at a large tertiary care hospital.
-
11 Jul 2025 1 repository listedCancer research is increasingly driven by the integration of diverse data modalities, spanning from genomics and proteomics to imaging and clinical factors.
-
10 Jun 2025 1 repository listedOur analysis reveals task-specific trade-offs: the specialized single-cell frameworks, scVI, CLAIRE, and the finetuned scGPT excel at uni-modal batch correction, while generic SSL methods, such as VICReg and SimCLR,…
-
10 Jun 2025 1 repository listedSingle-cell omics technologies have transformed our understanding of cellular diversity by enabling high-resolution profiling of individual cells.
-
6 Jun 2025 1 repository listedOur results on KRAMABENCH show that, although the models are sufficiently capable of solving well-specified data science code generation tasks, when extensive data processing and domain knowledge are required to…
-
29 May 2025 1 repository listedLeveraging a labeled dataset of thousands of videos, we evaluate a variety of LLMs in a zero-shot setting and compare their performance to a fine-tuned RoBERTa baseline.
-
25 May 2025 1 repository listedIn this work, we propose a data fusion approach for precipitation nowcasting by integrating data from meteorological and rain gauge stations in Rio de Janeiro metropolitan area with ERA5 reanalysis data and GFS…
-
15 Apr 2025 1 repository listedGPCCA addresses key challenges in multi-modal data analysis by handling missing values within the model, enabling the integration of more than two modalities, and identifying informative features while accounting for…
-
7 Apr 2025 1 repository listedRisk Analysis: Assesses market volatility and systemic risk using network analysis.
-
26 Mar 2025 1 repository listedWe evaluate the efficacy of our methodology in three settings: single-dataset learning, multi-dataset learning, and transfer learning.
-
24 Mar 2025 1 repository listedA key challenge today lies in the ability to efficiently handle multi-omics data since such multimodal data may provide a more comprehensive overview of the underlying processes in a system.
-
18 Mar 2025 1 repository listedWorld models significantly enhance hierarchical understanding, improving data integration and learning efficiency.
-
24 Feb 2025 1 repository listedWe employ star-shaped query processing and extend this technique to mapping candidate selection.
-
22 Feb 2025 1 repository listedThis highlights the potential of incorporating discrete weather station data to enhance the performance of deep learning-based weather nowcasting models.
-
29 Jan 2025 1 repository listedWe present a novel, non-invasive approach using multimodal predictive models that integrate RGB or thermal video data with patient metadata.
-
7 Jan 2025 1 repository listedKnowledge graphs have proven successful in integrating heterogeneous data across various domains.
-
1 Jan 2025 1 repository listedWe further explore the scenario where training data for the CPA task is available and can be used for selecting demonstrations or fine-tuning the model.
-
22 Dec 2024 1 repository listedEver since the vision was formulated, the Semantic Web has inspired many generations of innovations.
-
21 Dec 2024 1 repository listedAccelerating model convergence in resource-constrained environments is essential for fast and efficient neural network training.
-
20 Dec 2024 1 repository listedThe integration of multi-omic data is pivotal for understanding complex diseases, but its high dimensionality and noise present significant challenges.
-
13 Dec 2024 1 repository listedResting-state fMRI captures spontaneous neural activity characterized by complex spatiotemporal dynamics.
-
1 Dec 2024 1 repository listedRecent advancements in multimodal pre-training models have significantly advanced computational pathology.
-
23 Nov 2024 1 repository listedFederated Learning (FL) is a decentralized machine learning approach that has gained attention for its potential to enable collaborative model training across clients while protecting data privacy, making it an…
-
18 Nov 2024 1 repository listedAccurate mapping of the built asset information to established data classification systems and taxonomies is crucial for effective asset management, whether for compliance at project handover or ad-hoc data integration…
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections