{"url":"/dataset/oop","name":"OOP","full_name":null,"description_markdown":"The OOP benchmark features 431 Python programs that encompass essential OOP concepts and features like classes and encapsulation methods¹². The authors of the paper argue that current evaluation frameworks largely neglect OOP in favor of functional programming (FP), such as HumanEval and MBPP¹. To address this, they introduced this OOP-focused benchmark¹².\r\n\r\nIn addition to the benchmark, they also propose a novel evaluation metric, pass@o, tailored for OOP, enhancing traditional pass@k measures¹². This metric offers a more relevant and comprehensive assessment for OOP code generation¹.\r\n\r\nThe evaluation of 23 leading large language models (LLMs), including both general and code-specialized models, reveals three key insights¹²:\r\n1) pass@o offers a more relevant and comprehensive assessment for OOP code generation¹².\r\n2) Despite excelling in FP, code-specialized LLMs like WizardCoder lag in OOP compared to models like ChatGPT¹².\r\n3) The poor performance of all advanced LLMs on the OOP benchmark highlights a critical need for improvements in this field¹².\r\n\r\n(1) OOP: Object-Oriented Programming Evaluation Benchmark for Large .... https://arxiv.org/abs/2401.06628.\r\n(2) OOP: Object-Oriented Programming Evaluation Benchmark. https://arxiv.org/html/2401.06628v2.\r\n(3) OOP: Object-Oriented Programming Evaluation Benchmark .... https://www.x-mol.com/paper/1747067589511319552.\r\n(4) OOP：大型语言模型的面向对象编程评估基准,arXiv - CS .... https://www.x-mol.com/paper/1747067589511319552/t.\r\n(5) undefined. https://doi.org/10.48550/arXiv.2401.06628.\r\n(6) undefined. https://github.com/alphadl/OOP-eval.","description_withheld":null,"homepage":"https://github.com/alphadl/OOP-eval","introduced_date":"2024-01-12","introduced_date_note":null,"introduced_by":{"paper":"/paper/oop-object-oriented-programming-evaluation","title":"OOP: Object-Oriented Programming Evaluation Benchmark for Large Language Models","first_author":"Shuai Wang","url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["OOP"],"data_loaders":[],"num_papers_in_archive":3,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}