{"url":"/dataset/codescan","name":"CodeSCAN","full_name":"ScreenCast ANalysis for Video Programming Tutorials","description_markdown":"CodeSCAN is the first large-scale and diverse dataset of coding screenshots with pixel-perfect annotations. It features:\r\n\r\n- 24 popular programming languages (according to Github)\r\n- 100 random repositories per language (with MIT, BSD-3 or WTFPL License), i.e. 2.400 repositories in total\r\n- Per repository we use 5 files, i.e. 12.000 files in total\r\n- ~100 different themes and 25 different fonts\r\n- Diverse layouts changes, such as menu bar visibility, sidebar position, output window content, etc.\r\n- Numerous realistic interactions such as searching, typing and selecting within a file, etc.\r\n\r\nCheck our project page (https://a-nau.github.io/codescan/) for details.","description_withheld":null,"homepage":"https://a-nau.github.io/codescan/","introduced_date":null,"introduced_date_note":null,"introduced_by":null,"license":{"name":"Other (Non-Commercial)","url":"https://zenodo.org/records/10939237"},"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Object Detection","url":"/task/object-detection","datasets_with_task":"/datasets/task/object-detection"},{"name":"Optical Character Recognition (OCR)","url":"/task/optical-character-recognition","datasets_with_task":"/datasets/task/optical-character-recognition"},{"name":"Code Search","url":"/task/code-search","datasets_with_task":"/datasets/task/code-search"},{"name":"Image Stylization","url":"/task/image-stylization","datasets_with_task":"/datasets/task/image-stylization"},{"name":"Code Classification","url":"/task/code-classification","datasets_with_task":"/datasets/task/code-classification"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["CodeSCAN"],"data_loaders":[],"num_papers_in_archive":0,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}