{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/mama-mia-a-large-scale-multi-center-breast","title":"A large-scale multicenter breast cancer DCE-MRI benchmark dataset with expert segmentations","arxiv_id":"2406.13844","date":"2024-06-19","proceeding":null,"authors":["Lidia Garrucho","Kaisar Kushibar","Claire-Anne Reidel","Smriti Joshi","Richard Osuala","Apostolia Tsirikoglou","Maciej Bobowicz","Javier del Riego","Alessandro Catanese","Katarzyna Gwoździewicz","Maria-Laura Cosaka","Pasant M. Abo-Elhoda","Sara W. Tantawy","Shorouq S. Sakrana","Norhan O. Shawky-Abdelfatah","Amr Muhammad Abdo-Salem","Androniki Kozana","Eugen Divjak","Gordana Ivanac","Katerina Nikiforaki","Michail E. Klontzas","Rosa García-Dosdá","Meltem Gulsun-Akpinar","Oğuz Lafcı","Ritse Mann","Carlos Martín-Isla","Fred Prior","Kostas Marias","Martijn P. A. Starmans","Fredrik Strand","Oliver Díaz","Laura Igual","Karim Lekadir"],"abstract":"Artificial Intelligence (AI) research in breast cancer Magnetic Resonance Imaging (MRI) faces challenges due to limited expert-labeled segmentations. To address this, we present a multicenter dataset of 1506 pre-treatment T1-weighted dynamic contrast-enhanced MRI cases, including expert annotations of primary tumors and non-mass-enhanced regions. The dataset integrates imaging data from four collections in The Cancer Imaging Archive (TCIA), where only 163 cases with expert segmentations were initially available. To facilitate the annotation process, a deep learning model was trained to produce preliminary segmentations for the remaining cases. These were subsequently corrected and verified by 16 breast cancer experts (averaging 9 years of experience), creating a fully annotated dataset. Additionally, the dataset includes 49 harmonized clinical and demographic variables, as well as pre-trained weights for a baseline nnU-Net model trained on the annotated data. This resource addresses a critical gap in publicly available breast cancer datasets, enabling the development, validation, and benchmarking of advanced deep learning models, thus driving progress in breast cancer diagnostics, treatment response prediction, and personalized care.","url_abs":"https://arxiv.org/abs/2406.13844v3","url_pdf":"https://arxiv.org/pdf/2406.13844v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"mama-mia-a-large-scale-multi-center-breast","repo_url":"https://github.com/lidiagarrucho/mama-mia","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"benchmarking","task_name":"Benchmarking"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2406.13844","atlas_url":"https://app.syntology.ai/?focus=2406.13844","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2406.13844"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/lidiagarrucho/mama-mia","reach":{"status":"ok","spdx":"NOASSERTION"}}],"summary":{"ran":3},"by_repo_kind":{"official":{"samples":3,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"7ea3ef0e6f9fddf4","entry":"filter_clinical_data","repo":"lidiagarrucho/mama-mia","repo_kind":"official","path":"src/clinical_data.py","file_url":"https://github.com/lidiagarrucho/mama-mia/blob/HEAD/src/clinical_data.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"7ea3ef0e6f9fddf4"}},{"code_sha256_prefix":"492419b818aa283d","entry":"filter_patient_info","repo":"lidiagarrucho/mama-mia","repo_kind":"official","path":"src/clinical_data.py","file_url":"https://github.com/lidiagarrucho/mama-mia/blob/HEAD/src/clinical_data.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"492419b818aa283d"}},{"code_sha256_prefix":"be711f3c68679d28","entry":"region_or_label_to_mask","repo":"lidiagarrucho/mama-mia","repo_kind":"official","path":"src/challenge/metrics.py","file_url":"https://github.com/lidiagarrucho/mama-mia/blob/HEAD/src/challenge/metrics.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"be711f3c68679d28"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}