Browse State-of-the-Art › 16k
16k
87 papers with code · 1 benchmark · 1 dataset archive 2025-07-28
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| ConceptNet (1 row) | Suprime2 | 10 Years of the PCG workshop: Past and Future Trends | — | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
1 dataset whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
6 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 87 papers with code (146 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
27 May 2022 13 repositories listed Syntology ran 9 of 30 samples · 21 unverified · 1 pointer-only (licence)We also extend FlashAttention to block-sparse attention, yielding an approximate attention algorithm that is faster than any existing approximate attention method.
-
8 Nov 2020 5 repositories listed Syntology ran 2 of 3 samples · 1 unverifiedIn the recent months, a wide spectrum of efficient, fast Transformers have been proposed to tackle this problem, more often than not claiming superior or comparable model quality to vanilla Transformer models.
-
12 Sep 2019 4 repositories listedIn this work, we introduce the the Schema-Guided Dialogue (SGD) dataset, containing over 16k multi-domain conversations spanning 16 domains.
-
5 Jul 2024 3 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 1 pointer-only (licence)We evaluate our instantiations at the scale of 125M to 1.
-
27 Mar 2024 3 repositories listed Syntology ran 3 of 5 samples · 2 unverified · 4 pointer-only (licence)Empirically, we demonstrate that LLM agents can outperform crowdsourced human annotators - on a set of ~16k individual facts, SAFE agrees with crowdsourced human annotators 72% of the time, and on a random subset of 100…
-
28 Aug 2023 3 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedIn this paper, we introduce LongBench, the first bilingual, multi-task benchmark for long context understanding, enabling a more rigorous evaluation of long context understanding.
-
5 Jun 2025 2 repositories listed Syntology ran 2 of 7 samples · 5 unverifiedThe computational sparsity of Mixture-of-Experts (MoE) models enables sub-linear growth in compute cost as model size increases, thus offering a scalable path to training massive neural networks.
-
4 Oct 2024 2 repositories listedIn this paper, we present SwiftKV, a novel model transformation and distillation procedure specifically designed to reduce the time and cost of processing prompt tokens while preserving high quality of generated tokens.
-
3 Sep 2024 2 repositories listed Syntology ran 8 of 8 samples · 0 unverified · 8 pointer-only (licence)Current benchmarks like Needle-in-a-Haystack (NIAH), Ruler, and Needlebench focus on models' ability to understand long-context input sequences but fail to capture a critical dimension: the generation of high-quality…
-
17 May 2024 2 repositories listedThe datasets and scripts associated with CFLUE are openly accessible at https://github.
-
24 Aug 2023 2 repositories listed Syntology ran 3 of 5 samples · 2 unverified · 4 pointer-only (licence)We release Code Llama, a family of large language models for code based on Llama 2 providing state-of-the-art performance among open models, infilling capabilities, support for large input contexts, and zero-shot…
-
29 Jun 2023 2 repositories listed Syntology ran 6 of 6 samples · 0 unverified · 2 pointer-only (licence)Instruction tuning unlocks the superior capability of Large Language Models (LLM) to interact with humans.
-
1 Jun 2023 2 repositories listedWhile many works have proposed schemes to sparsify the attention patterns and reduce the computational overhead of self-attention, those are often limited by implementations concerns and end up imposing a simple and…
-
1 Jan 2023 2 repositories listedFor the cloth-changing problem, video-based ReID is rarely studied due to the lack of a suitable cloth-changing benchmark, and gait recognition is often researched under controlled conditions.
-
8 Aug 2022 2 repositories listedWhile large pretrained Transformer models have proven highly capable at tackling natural language tasks, handling long sequence inputs continues to be a significant challenge.
-
30 Apr 2020 2 repositories listedWith the emergence of the COVID-19 pandemic, the political and the medical aspects of disinformation merged as the problem got elevated to a whole new level to become the first global infodemic.
-
17 May 2015 2 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedIn this paper we introduce the problem of Visual Semantic Role Labeling: given an image we want to detect people doing actions and localize the objects of interaction.
-
12 Jun 2025 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)Furthermore, we build the Multi-Query Text Retrieval (MQTR) dataset, the first benchmark designed to evaluate the multi-query scene text retrieval capability of models, comprising four query types and 16k images.
-
8 Jun 2025 1 repository listedTo reduce the efficiency gap, we propose REO-RL, a class of Reinforcement Learning algorithms that minimizes REG by targeting a sparse set of token budgets.
-
28 May 2025 1 repository listedTo fill this gap, we introduce FAMA, the first family of open science SFMs for English and Italian, trained on 150k+ hours of OS speech data.
-
27 May 2025 1 repository listedSpeculative decoding is a widely adopted technique for accelerating inference in large language models (LLMs), but its performance degrades on long inputs due to increased attention cost and reduced draft accuracy.
-
27 May 2025 1 repository listedIn this paper, we introduce UI-Genie, a self-improving framework addressing two key challenges in GUI agents: verification of trajectory outcome is challenging and high-quality training data are not scalable.
-
24 May 2025 1 repository listed Syntology ran 1 of 3 samples · 2 unverified · 3 pointer-only (licence)Based on the variational form of softmax, we describe an efficient optimization-based algorithm to compute an approximate projection of softmax attention onto the class of Monarch matrices with Θ(N√(N) d) computational…
-
22 May 2025 1 repository listedWhile long-context large language models (LLMs) exhibit remarkable document processing capabilities, their prohibitively high training costs often hinder customized applications.
-
18 May 2025 1 repository listedThe core concept of those methods is to predefine or search for a set of factors to rescale the base frequencies of RoPE.
-
13 May 2025 1 repository listedPDDL-based symbolic task planning remains pivotal for robot autonomy yet struggles with dynamic human-robot collaboration due to scalability, re-planning demands, and delayed plan availability.
-
21 Mar 2025 1 repository listedWe present the KL3M tokenizers, a family of specialized tokenizers for legal, financial, and governmental text.
-
4 Feb 2025 1 repository listedAlgorithmic fairness has conventionally adopted the mathematically convenient perspective of racial color-blindness (i.
-
1 Feb 2025 1 repository listed Syntology ran 3 of 3 samples · 0 unverified · 2 pointer-only (licence)Equipping large language models (LLMs) with latent-space memory has attracted increasing attention as they can extend the context window of existing language models.
-
16 Jan 2025 1 repository listed Syntology ran 3 of 16 samples · 13 unverified · 16 pointer-only (licence)Hallucination remains a major challenge for Large Vision-Language Models (LVLMs).
Syntology lines on 14 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections