Browse State-of-the-Art › Benchmarking
Benchmarking
2,658 papers with code · 2 benchmarks · 10 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
2 leaderboard tables shown for this task, 2 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| CloudEval-YAML (1 row) | GPT-4 Turbo | CloudEval-YAML: A Practical Benchmark for Cloud Configuration Generation | code | — | Compare |
| Wiki-40B (1 row) | OutEffHop-Bert_base | Outlier-Efficient Hopfield Layers for Large Transformer-Based Models | code | Syntology ran 10 of 13 samples · 3 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
10 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 2,658 papers with code (5,548 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
17 Jun 2019 142 repositories listed Syntology ran 14 of 82 samples · 68 unverifiedIn this paper, we introduce the various features of this toolbox.
-
26 Feb 2021 82 repositories listed Syntology ran 16 of 20 samples · 4 unverified · 16 pointer-only (licence)State-of-the-art computer vision systems are trained to predict a fixed set of predetermined object categories.
-
25 Aug 2017 37 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedWe present Fashion-MNIST, a new dataset comprising of 28x28 grayscale images of 70, 000 fashion products from 10 categories, with 7, 000 images per category.
-
20 Nov 2014 24 repositories listed Syntology ran 12 of 32 samples · 20 unverified · 27 pointer-only (licence)We propose a novel paradigm for evaluating image descriptions that uses human consensus.
-
11 Feb 2019 23 repositories listed Syntology ran 6 of 15 samples · 9 unverified · 13 pointer-only (licence)In this paper, we propose the StarCraft Multi-Agent Challenge (SMAC) as a benchmark problem to fill this gap.
-
2 Mar 2020 15 repositories listed Syntology ran 1 of 23 samples · 22 unverifiedIn the last few years, graph neural networks (GNNs) have become the standard toolkit for analyzing and learning from data on graphs.
-
22 Apr 2016 15 repositories listed Syntology ran 0 of 4 samples · 4 unverifiedRecently, researchers have made significant progress combining the advances in deep learning for learning feature representations with reinforcement learning.
-
28 Mar 2019 14 repositories listed Syntology ran 2 of 3 samples · 1 unverifiedThen we propose a new dataset called ImageNet-P which enables researchers to benchmark a classifier's robustness to common perturbations.
-
28 Nov 2016 14 repositories listed Syntology ran 3 of 33 samples · 30 unverifiedThe size of the dataset and the fact that the questions are derived from real user search queries distinguishes MS MARCO from other well-known publicly available datasets for machine reading comprehension and…
-
2 Apr 2019 13 repositories listed Syntology ran 3 of 15 samples · 12 unverified · 15 pointer-only (licence)We present Habitat, a platform for research in embodied artificial intelligence (AI).
-
3 Oct 2018 13 repositories listed Syntology ran 3 of 11 samples · 8 unverified · 1 pointer-only (licence)Such architectural design and abstractions enable researchers and developers to extend the toolkit with their new algorithms and improvements, and to use it for performance benchmarking.
-
3 Oct 2016 13 repositories listed Syntology ran 3 of 26 samples · 23 unverified · 26 pointer-only (licence)An adversarial example library for constructing attacks, building defenses, and benchmarking both
-
25 Feb 2019 12 repositories listed Syntology ran 0 of 8 samples · 8 unverifiedSemantic segmentation of medical images aims to associate a pixel with a label in a medical image without human initialization.
-
29 Mar 2016 12 repositories listedWe introduce COCO, an open source platform for Comparing Continuous Optimizers in a black-box setting.
-
22 Mar 2017 11 repositories listed Syntology ran 4 of 40 samples · 36 unverified · 5 pointer-only (licence)Health care is one of the most exciting frontiers in data mining and machine learning.
-
16 Apr 2022 10 repositories listed Syntology ran 7 of 28 samples · 21 unverified · 4 pointer-only (licence)This large and diverse collection of tasks enables rigorous benchmarking of cross-task generalization under instructions -- training models to follow instructions on a subset of tasks and evaluating them on the…
-
18 Jul 2018 10 repositories listedSkillful mobile operation in three-dimensional environments is a primary topic of study in Artificial Intelligence.
-
14 Jun 2020 9 repositories listed Syntology ran 7 of 12 samples · 5 unverified · 7 pointer-only (licence)Multi-agent deep reinforcement learning (MARL) suffers from a lack of commonly-used evaluation tasks and criteria, making comparisons between approaches difficult.
-
22 Oct 2019 9 repositories listed Syntology ran 1 of 24 samples · 23 unverifiedPerson re-identification (re-ID), which aims to re-identify people across different camera views, has been significantly advanced by deep learning in recent years, particularly with convolutional neural networks (CNNs).
-
13 Mar 2019 9 repositories listedWe have recently seen the emergence of several publicly available Natural Language Understanding (NLU) toolkits, which map user utterances to structured, but more abstract, Dialogue Act (DA) or Intent specifications,…
-
11 Apr 2022 8 repositories listedWe also develop a metrics library, ivtmetrics, for model evaluation on surgical triplets.
-
15 Oct 2021 8 repositories listed Syntology ran 8 of 15 samples · 7 unverifiedLarge language models have recently been shown to attain reasonable zero-shot generalization on a diverse set of tasks (Brown et al., 2020).
-
9 Aug 2024 7 repositories listedThis document provides an overview of the challenge, including the registration process, rules, submission format, description of the datasets used, qualified team rankings, all team descriptions, and the benchmarking…
-
31 Mar 2019 7 repositories listed Syntology ran 3 of 4 samples · 1 unverified · 4 pointer-only (licence)The morphometry of a kidney tumor revealed by contrast-enhanced Computed Tomography (CT) imaging is an important factor in clinical decision making surrounding the lesion's diagnosis and treatment.
-
3 Dec 2018 7 repositories listed Syntology ran 1 of 3 samples · 2 unverifiedModern federated networks, such as those comprised of wearable devices, mobile phones, or autonomous vehicles, generate massive amounts of data each day.
-
28 Jan 2022 6 repositories listed Syntology ran 14 of 29 samples · 15 unverifiedDeep neural networks on 3D point cloud data have been widely used in the real world, especially in safety-critical applications.
-
12 Sep 2020 6 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedWe have publicly released the benchmarking code, evaluation protocols, and hyper-parameter settings of our work to promote reproducible research in this field.
-
5 Nov 2019 6 repositories listed Syntology ran 0 of 11 samples · 11 unverifiedTo make deliberate progress towards more intelligent and more human-like artificial systems, we need to be following an appropriate feedback signal: we need to be able to define and evaluate intelligence in a way that…
-
13 Jan 2019 6 repositories listed Syntology ran 0 of 8 samples · 8 unverifiedIn this work, we report the set-up and results of the Liver Tumor Segmentation Benchmark (LiTS), which was organized in conjunction with the IEEE International Symposium on Biomedical Imaging (ISBI) 2017 and the…
-
9 Feb 2024 5 repositories listedMachine-learning from a disparate set of tables, a data lake, requires assembling features by merging and aggregating tables.
Syntology lines on 24 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections