Papers › What Drives LLM Self-Reflection? A Controlled Ablation of Uncertainty Routing in Armed...
What Drives LLM Self-Reflection? A Controlled Ablation of Uncertainty Routing in Armed Conflict Forecasting
Poli Nemkova, Haeshitha Indukuri
Title, abstract, authors and date from arXiv's metadata (CC0); this paper is not in the Papers with Code archive (frozen 2025-07-28).
Self-reflection is widely assumed to improve LLM reasoning, yet which component drives the gain remains poorly understood. We present a controlled six-condition ablation isolating four components of LLM self-reflection: evidence exposure, diagnostic scaffolding, taxonomy vocabulary, and action routing. Two precise null results converge on a single mechanism. First, structured diagnostic questions add no measurable value over unstructured reflection (F1 = 0.296 vs $0.297$, p = 1.000, 95\% CI [-0.041, +0.040]). Second, presenting the full uncertainty taxonomy while collapsing the action space to a single generic action also adds no value (ΔF1 = +0.008, overlapping 95\% CIs), ruling out taxonomy vocabulary as the mechanism. Typed action routing provides consistent directional gains (F1 = 0.379 vs $0.296$); the conservative estimate controlling for taxonomy vocabulary is ΔF1 = +0.075, and the overall gain over the single-shot baseline is significant by bootstrap CI (ΔF1 = +0.101, 95\% CI [+0.020, +0.185]). The vocabulary-routing decomposition replicates on GPT-4o: taxonomy vocabulary adds no significant value over generic reflection (p = 0.773), while action routing provides significant gains (p = 0.025), confirming the mechanism holds across backbones. Gains concentrate on structurally novel conflicts: in Myanmar (F1: 0.000 →0.353) and Ukraine (0.167 →0.500), the vocabulary-only condition recovers no more than generic reflection while action routing breaks the degenerate prior. These findings identify typed action routing -- not diagnostic scaffolding or taxonomy vocabulary -- as a promising design principle for metacognitive LLM forecasting agents, while motivating larger-scale evaluation across conflict typologies.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Syntology holds the repository link but has not harvested or run code from it.
Results from the paper
The Papers with Code archive ends with its 2025-07-28 snapshot. This paper's arXiv identifier, 2608.12322, was issued in August 2026, after that date, so the archive has no leaderboard rows for it.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections