{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/efficient-and-accurate-explanation-estimation","title":"Efficient and Accurate Explanation Estimation with Distribution Compression","arxiv_id":"2406.18334","date":"2024-06-26","proceeding":null,"authors":["Hubert Baniecki","Giuseppe Casalicchio","Bernd Bischl","Przemyslaw Biecek"],"abstract":"We discover a theoretical connection between explanation estimation and distribution compression that significantly improves the approximation of feature attributions, importance, and effects. While the exact computation of various machine learning explanations requires numerous model inferences and becomes impractical, the computational cost of approximation increases with an ever-increasing size of data and model parameters. We show that the standard i.i.d. sampling used in a broad spectrum of algorithms for post-hoc explanation leads to an approximation error worthy of improvement. To this end, we introduce Compress Then Explain (CTE), a new paradigm of sample-efficient explainability. It relies on distribution compression through kernel thinning to obtain a data sample that best approximates its marginal distribution. CTE significantly improves the accuracy and stability of explanation estimation with negligible computational overhead. It often achieves an on-par explanation approximation error 2-3x faster by using fewer samples, i.e. requiring 2-3x fewer model evaluations. CTE is a simple, yet powerful, plug-in for any explanation method that now relies on i.i.d. sampling.","url_abs":"https://arxiv.org/abs/2406.18334v2","url_pdf":"https://arxiv.org/pdf/2406.18334v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"efficient-and-accurate-explanation-estimation","repo_url":"https://github.com/hbaniecki/compress-then-explain","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2406.18334","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2406.18334"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/hbaniecki/compress-then-explain","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran":3},"by_repo_kind":{"official":{"samples":3,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"787be538c707806e","entry":"d_mmd","repo":"hbaniecki/compress-then-explain","repo_kind":"official","path":"experiments/figure_3_8.py","file_url":"https://github.com/hbaniecki/compress-then-explain/blob/HEAD/experiments/figure_3_8.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"787be538c707806e"}},{"code_sha256_prefix":"84df6033caeb3c33","entry":"get_topk_values","repo":"hbaniecki/compress-then-explain","repo_kind":"official","path":"experiments/figure_3_8.py","file_url":"https://github.com/hbaniecki/compress-then-explain/blob/HEAD/experiments/figure_3_8.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"84df6033caeb3c33"}},{"code_sha256_prefix":"0b2e1b743ba30315","entry":"largest_power_of_2_in_sqrt","repo":"hbaniecki/compress-then-explain","repo_kind":"official","path":"experiments/cte/utils.py","file_url":"https://github.com/hbaniecki/compress-then-explain/blob/HEAD/experiments/cte/utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"0b2e1b743ba30315"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}