Datasets › JEEBench

JEEBench

Introduced by Daman Arora et al. in Have LLMs Advanced Enough? A Challenging Problem Solving Benchmark For Large Language Models24 May 2023 archive 2025-07-28

JEEBench is a considerably more challenging benchmark dataset for evaluating the problem solving abilities of LLMs. It curates 515 challenging pre-engineering mathematics, physics and chemistry problems from the IIT JEE-Advanced Exam. Long-horizon reasoning on top of deep in-domain knowledge is essential for solving problems in this benchmark.

Source: Have LLMs Advanced Enough? A Challenging Problem Solving Benchmark For Large Language Models

Image Source: Have LLMs Advanced Enough? A Challenging Problem Solving Benchmark For Large Language Models

Benchmarks archive 2025-07-28

No leaderboard in the archive resolves to this dataset.

Papers archive 2025-07-28

No paper in the archive has a leaderboard row on this dataset; the archive counts 15 papers for it but never published that list.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • JEEBench

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections