Papers › Where LLM Graders Succeed and Break: Evidence from Two Computer-Science Exams
Where LLM Graders Succeed and Break: Evidence from Two Computer-Science Exams
Ali Habibullah, Yazan Alshoibi, Mohammad Alshiekh, Salman Khan, Naeemullah Khan
Title, abstract, authors and date from arXiv's metadata (CC0); this paper is not in the Papers with Code archive (frozen 2025-07-28).
One long-form exam in a large course costs hundreds of grader-hours, and qualified graders are scarce; LLM graders are a tempting alternative. To show its pitfalls we grade a practical Computer Vision exam (570 dual-graded students) under 171 configurations spanning closed and open-weights models; the best reaches mean absolute error 1.64/35, below the 2.61/35 two human graders achieve against each other. The catch is the prompt: a short ''strict grader'' preamble drives 14 of 17 open-weights models out of the graded band (MAE ≥8), three stopping grading altogether. The damage traces to the preamble's two credit-withholding sentences, not to tone or model scale; one of them, ''never give partial credit'', alone makes two of three probed models stop grading. The closed flagships of three vendors shift calibration under it but stay in the band. In 162 further configurations on a second, independent Machine Learning exam from another course (1,038 dual-graded students), the preamble worsens ten models, moving three out of the band into collapse and one into refusal, yet improves seven whose neutral prompts over-mark: the vulnerability replicates, but its direction is exam-specific. Light LoRA fine-tuning repairs it: one adapter on the two exams' pooled ∼3,900 graded examples brings five small open models to parity or better with a human grader in agreement with the grader pair, and sensitivity to the three harsh personas nearly vanishes (≤0.32 MAE). We release the anonymised dataset, full ablation grid, and grading, fine-tuning and analysis pipelines.
In Syntology View this paper on Syntology, its page in Syntology's graph. That page lists the repositories linked to the paper, the abstract and the calls for agents.
Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
For agents, on Syntology's MCP service (how to connect):
get_citation_path(paper_1="2609.29333", paper_2="…")with another paper's arXiv id or titleget_concepts_for_paper(arxiv_id="2609.29333")
Code
Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Syntology holds the repository link but has not harvested or run code from it.
Results from the paper
The Papers with Code archive ends with its 2025-07-28 snapshot. This paper's arXiv identifier, 2609.29333, was issued in September 2026, after that date, so the archive has no leaderboard rows for it.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections