Datasets › CASTLE Benchmark

CASTLE Benchmark (CASTLE Benchmark C@250)

Introduced by Richard A. Dubniczky et al. in CASTLE: Benchmarking Dataset for Static Code Analyzers and LLMs towards CWE Detection15 Feb 2025 archive 2025-07-28

The CASTLE Benchmark is a comprehensive dataset and a scoring method for evaluating single or combinations of static analyzers with a focus on security. It consists of a hand-crafted dataset of 250 micro-benchmark programs (almost 11,000 lines of C code), covering 25 common CWEs. We also introduce the novel CASTLE Score metric to enable fair and reliable comparisons, considering factors such as true positive and false positive rates, as well as the tools' ability to find more common issues. This dataset enables a comparison of single tools, as well as the effectiveness of tool combinations.

Benchmarks archive 2025-07-28

No leaderboard in the archive resolves to this dataset.

Papers archive 2025-07-28

No paper in the archive has a leaderboard row on this dataset; the archive counts 1 paper for it but never published that list.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

Variants archive 2025-07-28

  • CASTLE Benchmark

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections