Datasets › CodeSCAN

CodeSCAN (ScreenCast ANalysis for Video Programming Tutorials)

archive 2025-07-28

CodeSCAN is the first large-scale and diverse dataset of coding screenshots with pixel-perfect annotations. It features:

  • 24 popular programming languages (according to Github)
  • 100 random repositories per language (with MIT, BSD-3 or WTFPL License), i.e. 2.400 repositories in total
  • Per repository we use 5 files, i.e. 12.000 files in total
  • ~100 different themes and 25 different fonts
  • Diverse layouts changes, such as menu bar visibility, sidebar position, output window content, etc.
  • Numerous realistic interactions such as searching, typing and selecting within a file, etc.

Check our project page (https://a-nau.github.io/codescan/) for details.

Benchmarks archive 2025-07-28

No leaderboard in the archive resolves to this dataset.

Papers archive 2025-07-28

No paper in the archive has a leaderboard row on this dataset.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

Other (Non-Commercial)

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • CodeSCAN

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections