{"url":"/dataset/105941-images-natural-scenes-ocr-data-of-12","name":"105,941 Images Natural Scenes OCR Data of 12 Languages","full_name":"105,941 Images Natural Scenes OCR Data of 12 Languages","description_markdown":"Description:\r\n105,941 Images Natural Scenes OCR Data of 12 Languages. The data covers 12 languages (6 Asian languages, 6 European languages), multiple natural scenes, multiple photographic angles. For annotation, line-level quadrilateral bounding box annotation and transcription for the texts were annotated in the data. The data can be used for tasks such as OCR of multi-language.\r\n\r\nData size:\r\n105,941 images, including Asian language family: Japanese 9,997 images, Korean 10,231 images, Indonesian 7,591 images, Malay 5,650 images, Vietnamese 8,822 images, Thai 9,645 images; European language family: French 10,015 images, German 7,213 images, Italian 8,824 images, Portuguese 7,754 images, Russian 10,376 images and Spanish 9,823 images\r\n\r\nCollecting environment:\r\nincluding shop plaque, stop board, poster, ticket, road sign, comic, cover picture, prompt/reminder, warning, packing instruction, menu, building sign, etc.","description_withheld":null,"homepage":"https://bit.ly/3xF8rGT","introduced_date":"2022-06-21","introduced_date_note":null,"introduced_by":null,"license":{"name":"Commercial license","url":"https://drive.google.com/file/d/1saDCPm74D4UWfBL17VbkTsZLGfpOQj1J/view?usp=sharing"},"modalities":[{"name":"Images","url":"/datasets/modality/images"}],"tasks":[{"name":"Optical Character Recognition (OCR)","url":"/task/optical-character-recognition","datasets_with_task":"/datasets/task/optical-character-recognition"}],"languages":[],"variants":["105,941 Images Natural Scenes OCR Data of 12 Languages"],"data_loaders":[],"num_papers_in_archive":24,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-25T09:33:49+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}