{"url":"/dataset/crackvision12k","name":"CrackVision12K","full_name":null,"description_markdown":"We present the CrackVision12k dataset, a collection of 12,000 crack images derived from 13 publicly available crack datasets. The individual datasets were too small to effectively train a deep learning model. Moreover, the masks in each dataset were annotated using different standards, so unifying the annotations was necessary. To achieve this, we applied various image processing techniques to each dataset to create masks that follow a consistent standard.\r\n\r\nCrack datasets inherently suffer from class imbalance. To mitigate this issue, we selected images containing crack pixels of more than 5000 pixels and applied data augmentation techniques such as Gaussian noise and rotation. Finally, there is a corresponding refined ground truth for each crack image across the dataset to ensure uniformity and reliability.\r\n\r\nThe 13 datasets we combined are as follows: Aigle-RN, ESAR, LCMS, CRACK500, CrackLS315, CRKWH100, CrackTree260, DeepCrack, GAPS384, Masonry, Stone331, CFD, and SDNet2018.","description_withheld":null,"homepage":"https://rdr.ucl.ac.uk/articles/dataset/CrackVision12K/26946472?file=49023628","introduced_date":"2024-09-04","introduced_date_note":null,"introduced_by":{"paper":"/paper/hybrid-segmentor-a-hybrid-approach-to","title":"Hybrid-Segmentor: A Hybrid Approach to Automated Fine-Grained Crack Segmentation in Civil Infrastructure","first_author":"June Moh Goo","url":null},"license":{"name":"CC0 1.0","url":"https://creativecommons.org/publicdomain/zero/1.0/"},"modalities":[{"name":"Images","url":"/datasets/modality/images"}],"tasks":[{"name":"Semantic Segmentation","url":"/task/semantic-segmentation","datasets_with_task":"/datasets/task/semantic-segmentation"},{"name":"2D Semantic Segmentation","url":"/task/2d-semantic-segmentation","datasets_with_task":"/datasets/task/2d-semantic-segmentation"},{"name":"Segmentation","url":"/task/segmentation","datasets_with_task":"/datasets/task/segmentation"},{"name":"Crack Segmentation","url":"/task/crack-segmentation","datasets_with_task":"/datasets/task/crack-segmentation"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["CrackVision12K"],"data_loaders":[],"num_papers_in_archive":4,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/crack-segmentation-on-crackvision12k","task":"Crack Segmentation","dataset_variant":"CrackVision12K","rows":4,"metrics":["mIoU"],"first_row_in_archive_order":{"model":"Hybrid-Segmentor","paper":"/paper/hybrid-segmentor-a-hybrid-approach-to","metrics":{"mIoU":"0.62982"},"code_links":[{"title":"junegoo94/Hybrid-Segmentor","url":"https://github.com/junegoo94/Hybrid-Segmentor"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/hybrid-segmentor-a-hybrid-approach-to","title":"Hybrid-Segmentor: A Hybrid Approach to Automated Fine-Grained Crack Segmentation in Civil Infrastructure","date":"2024-09-04","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/segformer-simple-and-efficient-design-for","title":"SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers","date":"2021-05-31","rows_on_this_dataset":1,"code_links":28,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":86,"samples_ran":48,"samples_unverified":38,"pointer_only_for_licence":15,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/u-net-convolutional-networks-for-biomedical","title":"U-Net: Convolutional Networks for Biomedical Image Segmentation","date":"2015-05-18","rows_on_this_dataset":1,"code_links":487,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":757,"samples_ran":510,"samples_unverified":247,"pointer_only_for_licence":426,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/fully-convolutional-networks-for-semantic-1","title":"Fully Convolutional Networks for Semantic Segmentation","date":"2014-11-14","rows_on_this_dataset":1,"code_links":51,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":4,"samples_ran":3,"samples_unverified":1,"pointer_only_for_licence":4,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":3,"samples_harvested":847,"samples_ran":561,"samples_unverified":286,"pointer_only_for_licence":445,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}