{"url":"/dataset/media-text","name":"Media-Text","full_name":"MediaText: a media industry-based dataset for scene text detetcion","description_markdown":"**Media-Text** dataset comprising images of banners, posters, covers and another images characterised for media industry.\r\n\r\n###  **DATASET DESCRIPTION**\r\n- 400 images\r\n- 7 744 annotated text instances\r\n- 973 annotations have been marked as illegible for the task of text recognition\r\n- 659 texts have been markes as do not care (###) for scene text detection.\r\n- Images are represented by 193 unique resolutions.\r\nAnnotation Format - Each image has corresponding  gt_*.txt file, which contains annotations in bounding box format (defined by 4 courners), transcription, and bool flag which determines that text is illegible for OCR. Proposed format is similar to ICDAR15 annotations.\r\n\r\nx1, x2, ..., x4, y4, transcription, OCR Flag \r\n\r\n**Example: **\r\n\r\n37,68,198,49,214,181,52,200,LADIES,False\r\n\r\n**Full paper: **  [ResearchGate](https://www.researchgate.net/publication/385351709_Media-Text_a_Media_Industry-Based_Dataset_for_Scene_Text_Detection)\r\n\r\n\r\nPlease cite the related works in your publications if it helps your research:\r\n<br>\r\n*S. Kalisz, M. Marczyk, J. Polańska, and R. Fagas, “Media-text: a media industry-based dataset for scene text detection,” in Modelling and simulation 2024. The 2024 European Simulation and Modelling Conference, M. Graña and J. D. Nuñez-Gonzalez, Eds., EUROSIS-ETI, 2024, pp. 138–144.*","description_withheld":null,"homepage":"https://zenodo.org/records/12796380","introduced_date":"2024-10-23","introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[{"name":"Images","url":"/datasets/modality/images"}],"tasks":[{"name":"Scene Text Detection","url":"/task/scene-text-detection","datasets_with_task":"/datasets/task/scene-text-detection"},{"name":"Text Detection","url":"/task/text-detection","datasets_with_task":"/datasets/task/text-detection"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["Media-Text"],"data_loaders":[],"num_papers_in_archive":0,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}