Browse State-of-the-Art › Data Summarization
Data Summarization
34 papers with code · 0 benchmarks · 3 datasets archive 2025-07-28
Data Summarization is a central problem in the area of machine learning, where we want to compute a small summary of the data.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
3 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 34 papers with code (97 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
6 Oct 2019 3 repositories listed Syntology ran 0 of 13 samples · 13 unverifiedWe propose to simultaneously distill both images and their labels, thus assigning each synthetic sample a `soft' label (a distribution of labels).
-
11 Dec 2020 2 repositories listedTo treat the non-stationary setting, we introduce a novel, exponentially weighted estimator for the Spearman rank correlation, which allows the local nonparametric correlation of a bivariate data stream to be tracked.
-
15 Jun 2020 2 repositories listed Syntology ran 12 of 14 samples · 2 unverifiedIn particular, we study the problem of label distillation - creating synthetic labels for a small set of real images, and show it to be more effective than the prior image-based approach to dataset distillation.
-
29 Nov 2018 2 repositories listedIn our algorithm, at each iteration, the maximum information from the structure of the data is captured by one selected sample, and the captured information is neglected in the next iterations by projection on the…
-
14 Oct 2024 1 repository listed Syntology ran 9 of 10 samples · 1 unverifiedTo this end, we argue that there is an essential to propose a Continuous Multi-task Spatio-Temporal learning framework (CMuST) to empower collective urban intelligence, which reforms the urban spatiotemporal learning…
-
9 Mar 2024 1 repository listedWe rigorously prove that DiffRed achieves a general upper bound of O(√((1-p)/k₂)) on Stress and O((1-p)/(√(k₂*ρ(A^*)))) on M1 where p is the fraction of variance explained by the first k₁ principal components and ρ(A^*)…
-
26 Aug 2023 1 repository listedData summarization is the process of generating interpretable and representative subsets from a dataset.
-
26 Apr 2023 1 repository listedAutomatic chart to text summarization is an effective tool for the visually impaired people along with providing precise insights of tabular data in natural language to the user.
-
19 Dec 2022 1 repository listedVisual language data such as plots, charts, and infographics are ubiquitous in the human world.
-
4 Nov 2022 1 repository listedRecent advances in coreset methods have shown that a selection of representative datapoints can replace massive volumes of data for Bayesian inference, preserving the relevant statistical information and significantly…
-
2 Nov 2022 1 repository listedSubmodular function maximization is a fundamental combinatorial optimization problem with plenty of applications -- including data summarization, influence maximization, and recommendation.
-
30 Jul 2022 1 repository listedGiven a set X of n elements, it asks to select a subset S of k ≪n elements with maximum \emph{diversity}, as quantified by the dissimilarities among the elements in S.
-
11 Jul 2022 1 repository listedWe examine recurrent, convolutional, and Transformer-based encoder-decoder models to automatically generate natural language summaries from numeric temporal personal health data.
-
7 Jul 2022 1 repository listedIn this paper, we study the classic submodular maximization problem subject to a group equality constraint under both non-adaptive and adaptive settings.
-
22 Feb 2022 1 repository listed Syntology ran 0 of 3 samples · 3 unverifiedA recent work has also leveraged submodular functions to propose submodular information measures which have been found to be very useful in solving the problems of guided subset selection and guided summarization.
-
30 Jan 2021 1 repository listedThis article describes techniques employed in the production of a synthetic dataset of driver telematics emulated from a similar real insurance dataset.
-
20 Oct 2020 1 repository listedData summarization has become a valuable tool in understanding even terabytes of data.
-
19 Oct 2020 1 repository listed Syntology ran 1 of 2 samples · 1 unverifiedActive learning is an effective technique for reducing the labeling cost by improving data efficiency.
-
9 Oct 2020 1 repository listedWe study the problem of extracting a small subset of representative items from a large data stream.
-
31 Aug 2020 1 repository listedModern machine learning applications should be able to address the intrinsic challenges arising over inference on massive real-world datasets, including scalability and robustness to outliers.
-
24 Jun 2020 1 repository listedUnderstanding how two datasets differ can help us determine whether one dataset under-represents certain sub-populations, and provides insights into how well models will generalize across datasets.
-
17 May 2020 1 repository listedThere are currently very few software packages available that offer quick and informative comparison of HDX-MS datasets and even few-er which offer statistical analysis and advanced visualization.
-
10 Feb 2020 1 repository listed Syntology ran 0 of 5 samples · 5 unverifiedOptimal transport (OT) is a powerful geometric and probabilistic tool for finding correspondences and measuring similarity between two distributions.
-
9 Feb 2020 1 repository listed Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)In this paper, we propose a novel framework that converts streaming algorithms for monotone submodular maximization into streaming algorithms for non-monotone submodular maximization.
-
17 Nov 2019 1 repository listedQuantifying the importance of each training point to a learning task is a fundamental problem in machine learning and the estimated importance scores have been leveraged to guide a range of data workflows such as data…
-
11 Jun 2019 1 repository listedLeast-mean squares (LMS) solvers such as Linear / Ridge / Lasso-Regression, SVD and Elastic-Net not only solve fundamental machine learning problems, but are also the building blocks in a variety of other methods, such…
-
8 Jun 2019 1 repository listed Syntology ran 0 of 11 samples · 11 unverifiedThis paper presents an explanation of submodular selection, an overview of the features in apricot, and an application to several data sets.
-
24 Jan 2019 1 repository listedIn data summarization we want to choose k prototypes in order to summarize a data set.
-
5 Sep 2018 1 repository listedSampling one or more effective solutions from large search spaces is a recurring idea in machine learning, and sequential optimization has become a popular solution.
-
20 Apr 2018 1 repository listedStructured data summarization involves generation of natural language summaries from structured input data.
Syntology lines on 8 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections