{"url":"/dataset/pecc","name":"PECC","full_name":"PECC: Problem Extraction and Coding Challenges","description_markdown":"Recent advancements in large language models (LLMs) have showcased their exceptional abilities across various tasks, such as code generation, problem-solving and reasoning. Existing benchmarks evaluate tasks in isolation, yet the extent to which LLMs can understand prose-style tasks, identify the underlying problems, and then generate appropriate code solutions is still unexplored. Addressing this gap, we introduce PECC, a novel benchmark derived from Advent Of Code (AoC) challenges and Project Euler, including 2396 problems. Unlike conventional benchmarks, PECC requires LLMs to interpret narrative-embedded problems, extract requirements, and generate executable code. A key feature of our dataset is the complexity added by natural language prompting in chat-based evaluations, mirroring real-world instruction ambiguities. Results show varying model performance between narrative and neutral problems, with specific challenges in the Euler math-based subset with GPT-3.5-Turbo passing 50% of the AoC challenges and only 8% on the Euler problems. By probing the limits of LLMs' capabilities, our benchmark provides a framework to monitor and assess the subsequent progress of LLMs as a universal problem solver.","description_withheld":null,"homepage":"https://github.com/hallerpatrick/pecc","introduced_date":"2024-04-29","introduced_date_note":null,"introduced_by":{"paper":"/paper/pecc-problem-extraction-and-coding-challenges","title":"PECC: Problem Extraction and Coding Challenges","first_author":"Patrick Haller","url":null},"license":{"name":"MIT","url":"https://github.com/hallerpatrick/pecc?tab=MIT-1-ov-file"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Code Generation","url":"/task/code-generation","datasets_with_task":"/datasets/task/code-generation"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["PECC"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/code-generation-on-pecc","task":"Code Generation","dataset_variant":"PECC","rows":8,"metrics":["Pass@3"],"first_row_in_archive_order":{"model":"Claude 3 Haiku","paper":"/paper/pecc-problem-extraction-and-coding-challenges","metrics":{"Pass@3":"27.67"},"code_links":[{"title":"hallerpatrick/pecc","url":"https://github.com/hallerpatrick/pecc"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/pecc-problem-extraction-and-coding-challenges","title":"PECC: Problem Extraction and Coding Challenges","date":"2024-04-29","rows_on_this_dataset":8,"code_links":1,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}