{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/mandoline-model-evaluation-under-distribution","title":"Mandoline: Model Evaluation under Distribution Shift","arxiv_id":"2107.00643","date":"2021-07-01","proceeding":null,"authors":["Mayee Chen","Karan Goel","Nimit S. Sohoni","Fait Poms","Kayvon Fatahalian","Christopher Ré"],"abstract":"Machine learning models are often deployed in different settings than they were trained and validated on, posing a challenge to practitioners who wish to predict how well the deployed model will perform on a target distribution. If an unlabeled sample from the target distribution is available, along with a labeled sample from a possibly different source distribution, standard approaches such as importance weighting can be applied to estimate performance on the target. However, importance weighting struggles when the source and target distributions have non-overlapping support or are high-dimensional. Taking inspiration from fields such as epidemiology and polling, we develop Mandoline, a new evaluation framework that mitigates these issues. Our key insight is that practitioners may have prior knowledge about the ways in which the distribution shifts, which we can use to better guide the importance weighting procedure. Specifically, users write simple \"slicing functions\" - noisy, potentially correlated binary functions intended to capture possible axes of distribution shift - to compute reweighted performance estimates. We further describe a density ratio estimation framework for the slices and show how its estimation error scales with slice quality and dataset size. Empirical validation on NLP and vision tasks shows that Mandoline can estimate performance on the target distribution up to 3x more accurately compared to standard baselines.","url_abs":"https://arxiv.org/abs/2107.00643v2","url_pdf":"https://arxiv.org/pdf/2107.00643v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"mandoline-model-evaluation-under-distribution","repo_url":"https://github.com/HazyResearch/mandoline","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"density-ratio-estimation","task_name":"Density Ratio Estimation"},{"task_slug":"epidemiology","task_name":"Epidemiology"},{"task_slug":"model","task_name":"model"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2107.00643","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2107.00643"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/HazyResearch/mandoline","reach":null}],"summary":{"ran_honours":1,"ran_violates":1,"ran_fixture":1},"by_repo_kind":{"official":{"samples":3,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"9244f2b360b12414","entry":"Phi","repo":"HazyResearch/mandoline","repo_kind":"official","path":"mandoline.py","file_url":"https://github.com/HazyResearch/mandoline/blob/HEAD/mandoline.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"9244f2b360b12414"}},{"code_sha256_prefix":"e45d15f7a93aca51","entry":"log_partition_ratio","repo":"HazyResearch/mandoline","repo_kind":"official","path":"mandoline.py","file_url":"https://github.com/HazyResearch/mandoline/blob/HEAD/mandoline.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"e45d15f7a93aca51"}},{"code_sha256_prefix":"347405051e7701e2","entry":"mandoline","repo":"HazyResearch/mandoline","repo_kind":"official","path":"mandoline.py","file_url":"https://github.com/HazyResearch/mandoline/blob/HEAD/mandoline.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"347405051e7701e2"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}