Datasets › CodeSCAN
CodeSCAN (ScreenCast ANalysis for Video Programming Tutorials)
CodeSCAN is the first large-scale and diverse dataset of coding screenshots with pixel-perfect annotations. It features:
- 24 popular programming languages (according to Github)
- 100 random repositories per language (with MIT, BSD-3 or WTFPL License), i.e. 2.400 repositories in total
- Per repository we use 5 files, i.e. 12.000 files in total
- ~100 different themes and 25 different fonts
- Diverse layouts changes, such as menu bar visibility, sidebar position, output window content, etc.
- Numerous realistic interactions such as searching, typing and selecting within a file, etc.
Check our project page (https://a-nau.github.io/codescan/) for details.
Benchmarks archive 2025-07-28
No leaderboard in the archive resolves to this dataset.
Papers archive 2025-07-28
No paper in the archive has a leaderboard row on this dataset.
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- CodeSCAN
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections