{"url":"/dataset/devops-eval","name":"DevOps-Eval","full_name":null,"description_markdown":"The **DevOps-Eval** is an industrial-first evaluation benchmark specifically designed for Large Language Models (LLMs) in the DevOps/AIOps domain¹. It was released by Ant Group in collaboration with Peking University³.\r\n\r\nThe goal of DevOps-Eval is to help developers, especially those in the DevOps field, track the progress and analyze the strengths and shortcomings of their models¹. The repository contains questions and exercises related to DevOps, including the AIOps and ToolLearning¹.\r\n\r\nHere are some key features of DevOps-Eval¹:\r\n- It currently contains **7486 multiple-choice questions** spanning 8 diverse general categories.\r\n- There are a total of **2840 samples** in the AIOps subcategory, covering scenarios such as log parsing, time series anomaly detection, time series classification, time series forecasting, and root cause analysis.\r\n- There are a total of **1509 samples** in the ToolLearning subcategory, covering 239 tool scenes across 59 fields.\r\n\r\nThe benchmark also includes a leaderboard that presents the zero-shot and five-shot accuracies from the models evaluated in the initial release¹. This allows for a comparison of different models' performance in the DevOps domain. \r\n\r\n(1) GitHub - codefuse-ai/codefuse-devops-eval: Industrial-first evaluation .... https://github.com/codefuse-ai/codefuse-devops-eval.\r\n(2) DevOps-Eval：蚂蚁集团联合北京大学发布首个面向DevOps .... https://developer.aliyun.com/article/1365893.\r\n(3) codefuse-devops-eval: A DevOps Domain Knowledge .... https://gitee.com/codefuse-ai/codefuse-devops-eval.\r\n(4) codefuse-devops-eval/resources/tutorial_zh.md at main .... https://github.com/codefuse-ai/codefuse-devops-eval/blob/main/resources/tutorial_zh.md.\r\n(5) DevOps-Eval：蚂蚁集团联合北京大学发布首个面向DevOps .... https://juejin.cn/post/7296513628332359692.","description_withheld":null,"homepage":"https://github.com/codefuse-ai/codefuse-devops-eval","introduced_date":null,"introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["DevOps-Eval"],"data_loaders":[],"num_papers_in_archive":0,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}