Browse State-of-the-Art › Automated Essay Scoring
Automated Essay Scoring
27 papers with code · 1 benchmark · 2 datasets archive 2025-07-28
Essay scoring: Automated Essay Scoring is the task of assigning a score to an essay, usually in the context of assessing the language ability of a language learner. The quality of an essay is affected by the following four primary dimensions: topic relevance, organization and coherence, word usage and sentence complexity, and grammar and mechanics.
Source: A Joint Model for Multimodal Document Quality Assessment
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| ASAP-AES (8 rows) | Neural Pairwise Contrastive Regression (NPCR) | Automated Essay Scoring via Pairwise Contrastive Regression | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
2 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
27 shown of 27 papers with code (104 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
23 Jan 2019 2 repositories listedCurrent state-of-art feature-engineered and end-to-end Automated Essay Score (AES) methods are proven to be unable to detect adversarial samples, e.
-
29 May 2024 1 repository listedWhile current Automated Essay Scoring (AES) methods demonstrate high scoring agreement with human raters, their decision-making mechanisms are not fully understood.
-
24 Apr 2024 1 repository listedWe evaluate both the AES performance that LLMs can achieve with prompting only and the helpfulness of the generated essay feedback.
-
13 Mar 2024 1 repository listedRecently, encoder-only pre-trained models such as BERT have been successfully applied in automated essay scoring (AES) to predict a single overall score.
-
10 Mar 2024 1 repository listedAlthough several methods were proposed to address the problem of automated essay scoring (AES) in the last 50 years, there is still much to desire in terms of effectiveness.
-
7 Feb 2024 1 repository listedWith an increasing focus in STEM education on critical thinking skills, science writing plays an ever more important role in curricula that stress inquiry skills.
-
12 Jan 2024 1 repository listedThrough extensive experiments with public and private datasets, we find that while LLMs do not surpass conventional state-of-the-art (SOTA) grading models in performance, they exhibit notable consistency,…
-
11 Jan 2024 1 repository listedAutomatic Essay Scoring (AES) is a well-established educational pursuit that employs machine learning to evaluate student-authored essays.
-
30 Oct 2023 1 repository listedHowever, among users are trolls, who provide training examples with incorrect labels.
-
10 Jun 2023 1 repository listedCoherence is an important aspect of text quality, and various approaches have been applied to coherence modeling.
-
26 May 2023 1 repository listed Syntology ran 0 of 6 samples · 6 unverifiedThus, predicting various trait scores of unseen-prompt essays (called cross-prompt essay trait scoring) is a remaining challenge of AES.
-
10 May 2023 1 repository listedHere, we propose WikiSQE, the first large-scale dataset for sentence quality estimation in Wikipedia.
-
28 Feb 2023 1 repository listedIn this study, we reproduce and compare state-of-the-art methods for AES in the Hindi domain.
-
1 Oct 2022 1 repository listedTo this end, in this paper we take inspiration from contrastive learning and propose a novel unified Neural Pairwise Contrastive Regression (NPCR) model in which both objectives are optimized simultaneously as a single…
-
8 May 2022 1 repository listedIn recent years, pre-trained models have become dominant in most natural language processing (NLP) tasks.
-
1 Nov 2021 1 repository listedIn this work, we first show that state-of-the-art systems, recent neural essay scoring systems, might be also influenced by the correlation between essay length and scores in a standard dataset.
-
13 Oct 2021 1 repository listedAutomated essay scoring (AES) is gaining increasing attention in the education sector as it significantly reduces the burden of manual scoring and allows ad hoc feedback for learners.
-
1 Aug 2021 1 repository listed“With the increasing popularity of learning Chinese as a second language (L2) the development of an automatic essay scoring (AES) method specially for Chinese L2 essays has become animportant task.
-
7 Apr 2021 1 repository listedAutomated text scoring (ATS) tasks, such as automated essay scoring and readability assessment, are important educational applications of natural language processing.
-
1 Feb 2021 1 repository listedTo find out which traits work best for different types of essays, we conduct ablation tests for each of the essay traits.
-
4 Aug 2020 1 repository listedCross-prompt automated essay scoring (AES) requires the system to use non target-prompt essays to award scores to a target-prompt essay.
-
14 Jul 2020 1 repository listedThis number is increasing further due to COVID-19 and the associated automation of education and testing.
-
18 Sep 2019 1 repository listedIn this paper, we present a new comparative study on automatic essay scoring (AES).
-
6 Aug 2019 1 repository listedThis paper presents an investigation of using a co-attention based neural network for source-dependent essay scoring.
-
18 Apr 2018 1 repository listedWe demonstrate that current state-of-the-art approaches to Automated Essay Scoring (AES) are not well-suited to capturing adversarially crafted input of grammatical but incoherent sequences of sentences.
-
14 Nov 2017 1 repository listedOur new method proposes a new \textsc{SkipFlow} mechanism that models relationships between snapshots of the hidden representations of a long short-term memory (LSTM) network as it reads.
-
1 Nov 2016 1 repository listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections