{"url":"/task/data-integration","name":"Data Integration","slug":"data-integration","description_markdown":"**Data integration** (also called information integration) is the process of consolidating data from a set of heterogeneous data sources into a single uniform data set (materialized integration) or view on the data (virtual integration). Data integration pipelines involve subtasks such as schema matching, table annotation, entity resolution, value normalization, data cleansing, and data fusion. Application domains of data integration include data warehousing, data lakes, and knowledge base consolidation.\r\nSurveys on Data integration: \r\n\r\n-  [Dong, Srivastava: Big data integration](https://ieeexplore.ieee.org/abstract/document/6544914),  2013.\r\n\r\n-  [Doan, Halevy, Ives: Principles of Data Integration](https://research.cs.wisc.edu/dibook/),  2012.","categories":[{"name":"Knowledge Base","url":"/area/knowledge-base"},{"name":"Natural Language Processing","url":"/area/natural-language-processing"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":431,"papers_with_code":130,"benchmarks":0,"benchmark_tables_in_archive":0,"benchmark_tables_shown":0,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":7,"subtasks":3,"parent_tasks":0},"benchmarks":[],"datasets":[{"url":"/dataset/amazon-google","name":"Amazon-Google","full_name":"","num_papers_in_archive":20},{"url":"/dataset/abt-buy","name":"Abt-Buy","full_name":"","num_papers_in_archive":19},{"url":"/dataset/wdc-sotab-v2","name":"WDC SOTAB V2","full_name":"","num_papers_in_archive":7},{"url":"/dataset/wikitables-turl","name":"WikiTables-TURL","full_name":"","num_papers_in_archive":7},{"url":"/dataset/wdc-products-1","name":"WDC Products","full_name":"","num_papers_in_archive":6},{"url":"/dataset/wdc-sotab","name":"WDC SOTAB","full_name":"","num_papers_in_archive":2},{"url":"/dataset/wdc-block","name":"WDC Block","full_name":"WDC Block: A Blocking Benchmark","num_papers_in_archive":1}],"subtasks":[{"url":"/task/entity-alignment","name":"Entity Alignment"},{"url":"/task/entity-resolution","name":"Entity Resolution"},{"url":"/task/table-annotation","name":"Table annotation"}],"parent_tasks":[],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":30,"of":130,"tagged_in_all":431,"items":[{"url":"/paper/iam-enhancing-rgb-d-instance-segmentation","title":"IAM: Enhancing RGB-D Instance Segmentation with New Benchmarks","date":"2025-01-03","arxiv_id":"2501.01685","repositories_listed":3,"syntology":null},{"url":"/paper/multimodal-contextualized-semantic-parsing","title":"Multimodal Contextualized Semantic Parsing from Speech","date":"2024-06-10","arxiv_id":"2406.06438","repositories_listed":2,"syntology":null},{"url":"/paper/beyond-trend-and-periodicity-guiding-time","title":"Intervention-Aware Forecasting: Breaking Historical Limits from a System Perspective","date":"2024-05-22","arxiv_id":"2405.13522","repositories_listed":2,"syntology":{"n":6,"n_ran":5,"n_unverified":1,"n_pointer_only":0}},{"url":"/paper/multimodal-teacher-forcing-for-reconstructing","title":"Integrating Multimodal Data for Joint Generative Modeling of Complex Dynamics","date":"2022-12-15","arxiv_id":"2212.07892","repositories_listed":2,"syntology":{"n":6,"n_ran":3,"n_unverified":3,"n_pointer_only":6}},{"url":"/paper/madrid-a-pipeline-for-metabolic-drug","title":"COMO: A Pipeline for Multi-Omics Data Integration in Metabolic Modeling and Drug Discovery","date":"2020-11-04","arxiv_id":"2011.02103","repositories_listed":2,"syntology":null},{"url":"/paper/bayesian-hybrid-matrix-factorisation-for-data","title":"Bayesian Hybrid Matrix Factorisation for Data Integration","date":"2017-04-17","arxiv_id":"1704.04962","repositories_listed":2,"syntology":null},{"url":"/paper/mimic-iii-a-freely-accessible-critical-care","title":"MIMIC-III, a freely accessible critical care database","date":"2016-05-24","arxiv_id":null,"repositories_listed":2,"syntology":null},{"url":"/paper/from-classical-machine-learning-to-emerging","title":"From Classical Machine Learning to Emerging Foundation Models: Review on Multimodal Data Integration for Cancer Research","date":"2025-07-11","arxiv_id":"2507.09028","repositories_listed":1,"syntology":null},{"url":"/paper/2506-10031","title":"scSSL-Bench: Benchmarking Self-Supervised Learning for Single-Cell Data","date":"2025-06-10","arxiv_id":"2506.10031","repositories_listed":1,"syntology":null},{"url":"/paper/the-cell-ontology-in-the-age-of-single-cell","title":"The Cell Ontology in the age of single-cell omics","date":"2025-06-10","arxiv_id":"2506.10037","repositories_listed":1,"syntology":null},{"url":"/paper/kramabench-a-benchmark-for-ai-systems-on-data","title":"KramaBench: A Benchmark for AI Systems on Data-to-Insight Pipelines over Data Lakes","date":"2025-06-06","arxiv_id":"2506.06541","repositories_listed":1,"syntology":null},{"url":"/paper/evaluating-ai-capabilities-in-detecting","title":"Evaluating AI capabilities in detecting conspiracy theories on YouTube","date":"2025-05-29","arxiv_id":"2505.23570","repositories_listed":1,"syntology":null},{"url":"/paper/towards-a-spatiotemporal-fusion-approach-to","title":"Towards a Spatiotemporal Fusion Approach to Precipitation Nowcasting","date":"2025-05-25","arxiv_id":"2505.19258","repositories_listed":1,"syntology":null},{"url":"/paper/generalized-probabilistic-canonical","title":"Generalized probabilistic canonical correlation analysis for multi-modal data integration with full or partial observations","date":"2025-04-15","arxiv_id":"2504.11610","repositories_listed":1,"syntology":null},{"url":"/paper/omniecon-nexus-global-microeconomic","title":"OmniEcon Nexus: Global Microeconomic Simulation Engine","date":"2025-04-07","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/multi-dataset-and-transfer-learning-using","title":"Multi-dataset and Transfer Learning Using Gene Expression Knowledge Graphs","date":"2025-03-26","arxiv_id":"2503.20400","repositories_listed":1,"syntology":null},{"url":"/paper/modis-multi-omics-data-integration-for-small","title":"MODIS: Multi-Omics Data Integration for Small and Unpaired Datasets","date":"2025-03-24","arxiv_id":"2503.18856","repositories_listed":1,"syntology":null},{"url":"/paper/fusdreamer-label-efficient-remote-sensing","title":"FusDreamer: Label-efficient Remote Sensing World Model for Multimodal Data Classification","date":"2025-03-18","arxiv_id":"2503.13814","repositories_listed":1,"syntology":null},{"url":"/paper/intermediate-triple-table-a-general","title":"Intermediate triple table: A general architecture for virtual knowledge graphs","date":"2025-02-24","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/integrating-weather-station-data-and-radar","title":"Integrating Weather Station Data and Radar for Precipitation Nowcasting: SmaAt-fUsion and SmaAt-Krige-GNet","date":"2025-02-22","arxiv_id":"2502.16116","repositories_listed":1,"syntology":null},{"url":"/paper/pulmofusion-advancing-pulmonary-health-with","title":"PulmoFusion: Advancing Pulmonary Health with Efficient Multi-Modal Fusion","date":"2025-01-29","arxiv_id":"2501.17699","repositories_listed":1,"syntology":null},{"url":"/paper/reckg-knowledge-graph-for-recommender-systems","title":"RecKG: Knowledge Graph for Recommender Systems","date":"2025-01-07","arxiv_id":"2501.03598","repositories_listed":1,"syntology":null},{"url":"/paper/column-property-annotation-using-large","title":"Column Property Annotation using Large Language Models","date":"2025-01-01","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/semantic-web-past-present-and-future","title":"Semantic Web: Past, Present, and Future","date":"2024-12-22","arxiv_id":"2412.17159","repositories_listed":1,"syntology":null},{"url":"/paper/learn2mix-training-neural-networks-using","title":"Learn2Mix: Training Neural Networks Using Adaptive Data Integration","date":"2024-12-21","arxiv_id":"2412.16482","repositories_listed":1,"syntology":null},{"url":"/paper/graphseqlm-a-unified-graph-language-framework","title":"GraphSeqLM: A Unified Graph Language Framework for Omic Graph Learning","date":"2024-12-20","arxiv_id":"2412.15790","repositories_listed":1,"syntology":null},{"url":"/paper/data-integration-with-fusion-searchlight","title":"Data Integration with Fusion Searchlight: Classifying Brain States from Resting-state fMRI","date":"2024-12-13","arxiv_id":"2412.10161","repositories_listed":1,"syntology":null},{"url":"/paper/towards-unified-molecule-enhanced-pathology","title":"Towards Unified Molecule-Enhanced Pathology Image Representation Learning via Integrating Spatial Transcriptomics","date":"2024-12-01","arxiv_id":"2412.00651","repositories_listed":1,"syntology":null},{"url":"/paper/federated-learning-in-chemical-engineering-a","title":"Federated Learning in Chemical Engineering: A Tutorial on a Framework for Privacy-Preserving Collaboration Across Distributed Data Sources","date":"2024-11-23","arxiv_id":"2411.16737","repositories_listed":1,"syntology":null},{"url":"/paper/benchmarking-pre-trained-text-embedding","title":"Benchmarking pre-trained text embedding models in aligning built asset information","date":"2024-11-18","arxiv_id":"2411.12056","repositories_listed":1,"syntology":null}],"syntology_records":2,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}