{"url":"/task/data-valuation","name":"Data Valuation","slug":"data-valuation","description_markdown":"Data valuation in machine learning tries to determine the worth of data, or data sets, for downstream tasks. Some methods are task-agnostic and consider datasets as a whole, mostly for decision making in data markets. These look at distributional distances between samples. More often, methods look at how individual points affect performance of specific machine learning models. They assign a scalar to each element of a training set which reflects its contribution to the final performance of some model trained on it. Some concepts of value depend on a specific model of interest, others are model-agnostic.\r\n\r\nConcepts of the usefulness of a datum or its influence on the outcome of a prediction have a long history in statistics and ML, in particular through the notion of the influence function. However, it has only been recently that rigorous and practical notions of value for data, and in particular data-sets, have appeared in the ML literature, often based on concepts from collaborative game theory, but also from generalization estimates of neural networks, or optimal transport theory, among others.","categories":[{"name":"Methodology","url":"/area/methodology"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"derived"},"counts":{"papers_tagged":119,"papers_with_code":53,"benchmarks":0,"benchmark_tables_in_archive":0,"benchmark_tables_shown":0,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":0,"subtasks":1,"parent_tasks":0},"benchmarks":[],"datasets":[],"subtasks":[{"url":"/task/data-interaction","name":"Data Interaction"}],"parent_tasks":[],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":30,"of":53,"tagged_in_all":119,"items":[{"url":"/paper/data-shapley-equitable-valuation-of-data-for","title":"Data Shapley: Equitable Valuation of Data for Machine Learning","date":"2019-04-05","arxiv_id":"1904.02868","repositories_listed":6,"syntology":{"n":2,"n_ran":0,"n_unverified":2,"n_pointer_only":0}},{"url":"/paper/stochastic-amortization-a-unified-approach-to","title":"Stochastic Amortization: A Unified Approach to Accelerate Feature and Data Attribution","date":"2024-01-29","arxiv_id":"2401.15866","repositories_listed":3,"syntology":{"n":7,"n_ran":6,"n_unverified":1,"n_pointer_only":0}},{"url":"/paper/the-shapley-value-in-machine-learning","title":"The Shapley Value in Machine Learning","date":"2022-02-11","arxiv_id":"2202.05594","repositories_listed":3,"syntology":{"n":2,"n_ran":0,"n_unverified":2,"n_pointer_only":0}},{"url":"/paper/efficient-task-specific-data-valuation-for","title":"Efficient Task-Specific Data Valuation for Nearest Neighbor Algorithms","date":"2019-08-22","arxiv_id":"1908.08619","repositories_listed":3,"syntology":{"n":5,"n_ran":3,"n_unverified":2,"n_pointer_only":5}},{"url":"/paper/2d-oob-attributing-data-contribution-through","title":"2D-OOB: Attributing Data Contribution Through Joint Valuation Framework","date":"2024-08-07","arxiv_id":"2408.03572","repositories_listed":2,"syntology":null},{"url":"/paper/chg-shapley-efficient-data-valuation-and","title":"CHG Shapley: Efficient Data Valuation and Selection towards Trustworthy Machine Learning","date":"2024-06-17","arxiv_id":"2406.11730","repositories_listed":2,"syntology":{"n":8,"n_ran":6,"n_unverified":2,"n_pointer_only":0}},{"url":"/paper/opendataval-a-unified-benchmark-for-data-1","title":"OpenDataVal: a Unified Benchmark for Data Valuation","date":"2023-06-18","arxiv_id":"2306.10577","repositories_listed":2,"syntology":{"n":7,"n_ran":6,"n_unverified":1,"n_pointer_only":0}},{"url":"/paper/data-oob-out-of-bag-estimate-as-a-simple-and","title":"Data-OOB: Out-of-bag Estimate as a Simple and Efficient Data Value","date":"2023-04-16","arxiv_id":"2304.07718","repositories_listed":2,"syntology":null},{"url":"/paper/cs-shapley-class-wise-shapley-values-for-data","title":"CS-Shapley: Class-wise Shapley Values for Data Valuation in Classification","date":"2022-11-13","arxiv_id":"2211.06800","repositories_listed":2,"syntology":null},{"url":"/paper/data-banzhaf-a-data-valuation-framework-with","title":"Data Banzhaf: A Robust Data Valuation Framework for Machine Learning","date":"2022-05-30","arxiv_id":"2205.15466","repositories_listed":2,"syntology":null},{"url":"/paper/beta-shapley-a-unified-and-noise-reduced-data","title":"Beta Shapley: a Unified and Noise-reduced Data Valuation Framework for Machine Learning","date":"2021-10-26","arxiv_id":"2110.14049","repositories_listed":2,"syntology":null},{"url":"/paper/data-valuation-using-reinforcement-learning","title":"Data Valuation using Reinforcement Learning","date":"2019-09-25","arxiv_id":"1909.11671","repositories_listed":2,"syntology":{"n":2,"n_ran":2,"n_unverified":0,"n_pointer_only":2}},{"url":"/paper/faithful-group-shapley-value","title":"Faithful Group Shapley Value","date":"2025-05-25","arxiv_id":"2505.19013","repositories_listed":1,"syntology":{"n":1,"n_ran":1,"n_unverified":0,"n_pointer_only":0}},{"url":"/paper/shapley-guided-utility-learning-for-effective","title":"Shapley-Guided Utility Learning for Effective Graph Inference Data Valuation","date":"2025-03-23","arxiv_id":"2503.18195","repositories_listed":1,"syntology":null},{"url":"/paper/fw-shapley-real-time-estimation-of-weighted","title":"FW-Shapley: Real-time Estimation of Weighted Shapley Values","date":"2025-03-09","arxiv_id":"2503.06602","repositories_listed":1,"syntology":null},{"url":"/paper/alinfik-learning-to-approximate-linearized","title":"ALinFiK: Learning to Approximate Linearized Future Influence Kernel for Scalable Third-Party LLM Data Valuation","date":"2025-03-02","arxiv_id":"2503.01052","repositories_listed":1,"syntology":null},{"url":"/paper/dupre-data-utility-prediction-for-efficient","title":"DUPRE: Data Utility Prediction for Efficient Data Valuation","date":"2025-02-22","arxiv_id":"2502.16152","repositories_listed":1,"syntology":null},{"url":"/paper/data-valuation-using-neural-networks-for","title":"Data Valuation using Neural Networks for Efficient Instruction Fine-Tuning","date":"2025-02-14","arxiv_id":"2502.09969","repositories_listed":1,"syntology":null},{"url":"/paper/beyond-models-explainable-data-valuation-and","title":"Beyond Models! Explainable Data Valuation and Metric Adaption for Recommendation","date":"2025-02-12","arxiv_id":"2502.08685","repositories_listed":1,"syntology":null},{"url":"/paper/qless-a-quantized-approach-for-data-valuation","title":"QLESS: A Quantized Approach for Data Valuation and Selection in Large Language Model Fine-Tuning","date":"2025-02-03","arxiv_id":"2502.01703","repositories_listed":1,"syntology":null},{"url":"/paper/lossval-efficient-data-valuation-for-neural","title":"LossVal: Efficient Data Valuation for Neural Networks","date":"2024-12-05","arxiv_id":"2412.04158","repositories_listed":1,"syntology":null},{"url":"/paper/towards-data-valuation-via-asymmetric-data","title":"Towards Data Valuation via Asymmetric Data Shapley","date":"2024-11-01","arxiv_id":"2411.00388","repositories_listed":1,"syntology":null},{"url":"/paper/one-sample-fits-all-approximating-all","title":"One Sample Fits All: Approximating All Probabilistic Values Simultaneously and Efficiently","date":"2024-10-31","arxiv_id":"2410.23808","repositories_listed":1,"syntology":{"n":3,"n_ran":0,"n_unverified":3,"n_pointer_only":3}},{"url":"/paper/data-distribution-valuation","title":"Data Distribution Valuation","date":"2024-10-06","arxiv_id":"2410.04386","repositories_listed":1,"syntology":{"n":13,"n_ran":12,"n_unverified":1,"n_pointer_only":0}},{"url":"/paper/shapiq-shapley-interactions-for-machine","title":"shapiq: Shapley Interactions for Machine Learning","date":"2024-10-02","arxiv_id":"2410.01649","repositories_listed":1,"syntology":{"n":1,"n_ran":1,"n_unverified":0,"n_pointer_only":0}},{"url":"/paper/targeted-synthetic-data-generation-for","title":"Targeted synthetic data generation for tabular data via hardness characterization","date":"2024-10-01","arxiv_id":"2410.00759","repositories_listed":1,"syntology":null},{"url":"/paper/influence-based-attributions-can-be","title":"Influence-based Attributions can be Manipulated","date":"2024-09-08","arxiv_id":"2409.05208","repositories_listed":1,"syntology":{"n":5,"n_ran":4,"n_unverified":1,"n_pointer_only":5}},{"url":"/paper/in-context-probing-approximates-influence","title":"In-Context Probing Approximates Influence Function for Data Valuation","date":"2024-07-17","arxiv_id":"2407.12259","repositories_listed":1,"syntology":null},{"url":"/paper/redefining-contributions-shapley-driven","title":"Redefining Contributions: Shapley-Driven Federated Learning","date":"2024-06-01","arxiv_id":"2406.00569","repositories_listed":1,"syntology":{"n":5,"n_ran":5,"n_unverified":0,"n_pointer_only":0}},{"url":"/paper/scaling-laws-for-the-value-of-individual-data","title":"Scaling Laws for the Value of Individual Data Points in Machine Learning","date":"2024-05-30","arxiv_id":"2405.20456","repositories_listed":1,"syntology":{"n":2,"n_ran":2,"n_unverified":0,"n_pointer_only":0}}],"syntology_records":14,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}