{"url":"/task/data-summarization","name":"Data Summarization","slug":"data-summarization","description_markdown":"**Data Summarization** is a central problem in the area of machine learning, where we want to compute a small summary of the data.\n\n\n<span class=\"description-source\">Source: [How to Solve Fair k-Center in Massive Data Models ](https://arxiv.org/abs/2002.07682)</span>","categories":[{"name":"Miscellaneous","url":"/area/miscellaneous"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":97,"papers_with_code":34,"benchmarks":0,"benchmark_tables_in_archive":0,"benchmark_tables_shown":0,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":3,"subtasks":0,"parent_tasks":0},"benchmarks":[],"datasets":[{"url":"/dataset/mentsum","name":"MentSum","full_name":"Mental Health Summarization Dataset","num_papers_in_archive":2},{"url":"/dataset/m3ls-multi-lingual-multi-modal-summarization","name":"M3LS","full_name":"Multi-Lingual Multi-Modal Summarization Dataset","num_papers_in_archive":1},{"url":"/dataset/pjm-aep","name":"PJM(AEP)","full_name":"","num_papers_in_archive":0}],"subtasks":[],"parent_tasks":[],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":30,"of":34,"tagged_in_all":97,"items":[{"url":"/paper/improving-dataset-distillation","title":"Soft-Label Dataset Distillation and Text Dataset Distillation","date":"2019-10-06","arxiv_id":"1910.02551","repositories_listed":3,"syntology":{"n":13,"n_ran":0,"n_unverified":13,"n_pointer_only":0}},{"url":"/paper/sequential-estimation-of-nonparametric","title":"Sequential estimation of Spearman rank correlation using Hermite series estimators","date":"2020-12-11","arxiv_id":"2012.06287","repositories_listed":2,"syntology":null},{"url":"/paper/flexible-dataset-distillation-learn-labels","title":"Flexible Dataset Distillation: Learn Labels Instead of Images","date":"2020-06-15","arxiv_id":"2006.08572","repositories_listed":2,"syntology":{"n":14,"n_ran":12,"n_unverified":2,"n_pointer_only":0}},{"url":"/paper/iterative-projection-and-matching-finding","title":"Iterative Projection and Matching: Finding Structure-preserving Representatives and Its Application to Computer Vision","date":"2018-11-29","arxiv_id":"1811.12326","repositories_listed":2,"syntology":null},{"url":"/paper/get-rid-of-task-isolation-a-continuous-multi","title":"Get Rid of Isolation: A Continuous Multi-task Spatio-Temporal Learning Framework","date":"2024-10-14","arxiv_id":"2410.10524","repositories_listed":1,"syntology":{"n":10,"n_ran":9,"n_unverified":1,"n_pointer_only":0}},{"url":"/paper/diffred-dimensionality-reduction-guided-by","title":"DiffRed: Dimensionality Reduction guided by stable rank","date":"2024-03-09","arxiv_id":"2403.05882","repositories_listed":1,"syntology":null},{"url":"/paper/time-to-pattern-information-theoretic","title":"Time-to-Pattern: Information-Theoretic Unsupervised Learning for Scalable Time Series Summarization","date":"2023-08-26","arxiv_id":"2308.13722","repositories_listed":1,"syntology":null},{"url":"/paper/chartsumm-a-comprehensive-benchmark-for","title":"ChartSumm: A Comprehensive Benchmark for Automatic Chart Summarization of Long and Short Summaries","date":"2023-04-26","arxiv_id":"2304.13620","repositories_listed":1,"syntology":null},{"url":"/paper/matcha-enhancing-visual-language-pretraining","title":"MatCha: Enhancing Visual Language Pretraining with Math Reasoning and Chart Derendering","date":"2022-12-19","arxiv_id":"2212.09662","repositories_listed":1,"syntology":null},{"url":"/paper/black-box-coreset-variational-inference","title":"Black-box Coreset Variational Inference","date":"2022-11-04","arxiv_id":"2211.02377","repositories_listed":1,"syntology":null},{"url":"/paper/balancing-utility-and-fairness-in-submodular","title":"Balancing Utility and Fairness in Submodular Maximization (Technical Report)","date":"2022-11-02","arxiv_id":"2211.00980","repositories_listed":1,"syntology":null},{"url":"/paper/streaming-algorithms-for-diversity","title":"Streaming Algorithms for Diversity Maximization with Fairness Constraints","date":"2022-07-30","arxiv_id":"2208.00194","repositories_listed":1,"syntology":null},{"url":"/paper/towards-neural-numeric-to-text-generation","title":"Towards Neural Numeric-To-Text Generation From Temporal Personal Health Data","date":"2022-07-11","arxiv_id":"2207.05194","repositories_listed":1,"syntology":null},{"url":"/paper/group-fairness-in-adaptive-submodular","title":"Group Equality in Adaptive Submodular Maximization","date":"2022-07-07","arxiv_id":"2207.03364","repositories_listed":1,"syntology":null},{"url":"/paper/submodlib-a-submodular-optimization-library","title":"Submodlib: A Submodular Optimization Library","date":"2022-02-22","arxiv_id":"2202.10680","repositories_listed":1,"syntology":{"n":3,"n_ran":0,"n_unverified":3,"n_pointer_only":0}},{"url":"/paper/synthetic-dataset-generation-of-driver","title":"Synthetic Dataset Generation of Driver Telematics","date":"2021-01-30","arxiv_id":"2102.00252","repositories_listed":1,"syntology":null},{"url":"/paper/very-fast-streaming-submodular-function","title":"Very Fast Streaming Submodular Function Maximization","date":"2020-10-20","arxiv_id":"2010.10059","repositories_listed":1,"syntology":null},{"url":"/paper/semi-supervised-batch-active-learning-via","title":"Semi-supervised Batch Active Learning via Bilevel Optimization","date":"2020-10-19","arxiv_id":"2010.09654","repositories_listed":1,"syntology":{"n":2,"n_ran":1,"n_unverified":1,"n_pointer_only":0}},{"url":"/paper/streaming-submodular-maximization-with","title":"Fair and Representative Subset Selection from Data Streams","date":"2020-10-09","arxiv_id":"2010.04412","repositories_listed":1,"syntology":null},{"url":"/paper/b-cores-robust-large-scale-bayesian-data","title":"$β$-Cores: Robust Large-Scale Bayesian Data Summarization in the Presence of Outliers","date":"2020-08-31","arxiv_id":"2008.13600","repositories_listed":1,"syntology":null},{"url":"/paper/anova-exemplars-for-understanding-data-drift","title":"Understanding collections of related datasets using dependent MMD coresets","date":"2020-06-24","arxiv_id":"2006.14621","repositories_listed":1,"syntology":null},{"url":"/paper/deuteros-2-0-peptide-level-significance","title":"Deuteros 2.0: Peptide-level significance testing of data from hydrogen deuterium exchange mass spectrometry","date":"2020-05-17","arxiv_id":"2005.08380","repositories_listed":1,"syntology":null},{"url":"/paper/co-optimal-transport","title":"CO-Optimal Transport","date":"2020-02-10","arxiv_id":"2002.03731","repositories_listed":1,"syntology":{"n":5,"n_ran":0,"n_unverified":5,"n_pointer_only":0}},{"url":"/paper/streaming-submodular-maximization-under-a-k","title":"Streaming Submodular Maximization under a $k$-Set System Constraint","date":"2020-02-09","arxiv_id":"2002.03352","repositories_listed":1,"syntology":{"n":3,"n_ran":2,"n_unverified":1,"n_pointer_only":3}},{"url":"/paper/an-empirical-and-comparative-analysis-of-data-1","title":"Scalability vs. Utility: Do We Have to Sacrifice One for the Other in Data Importance Quantification?","date":"2019-11-17","arxiv_id":"1911.07128","repositories_listed":1,"syntology":null},{"url":"/paper/fast-and-accurate-least-mean-squares-solvers","title":"Fast and Accurate Least-Mean-Squares Solvers","date":"2019-06-11","arxiv_id":"1906.04705","repositories_listed":1,"syntology":null},{"url":"/paper/apricot-submodular-selection-for-data","title":"apricot: Submodular selection for data summarization in Python","date":"2019-06-08","arxiv_id":"1906.03543","repositories_listed":1,"syntology":{"n":11,"n_ran":0,"n_unverified":11,"n_pointer_only":0}},{"url":"/paper/fair-k-center-clustering-for-data","title":"Fair k-Center Clustering for Data Summarization","date":"2019-01-24","arxiv_id":"1901.08628","repositories_listed":1,"syntology":null},{"url":"/paper/controlled-random-search-improves-hyper","title":"Coverage-Based Designs Improve Sample Mining and Hyper-Parameter Optimization","date":"2018-09-05","arxiv_id":"1809.01712","repositories_listed":1,"syntology":null},{"url":"/paper/a-mixed-hierarchical-attention-based-encoder","title":"A Mixed Hierarchical Attention based Encoder-Decoder Approach for Standard Table Summarization","date":"2018-04-20","arxiv_id":"1804.07790","repositories_listed":1,"syntology":null}],"syntology_records":8,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}