Papers › Adversarial Closed-Loop Curriculum for Evolving Role-Playing Agents

Adversarial Closed-Loop Curriculum for Evolving Role-Playing Agents

23 Sep 2026arXiv:2609.28609added by Syntology

Zheng Zhang, Liu Liu, Qi Chai, Deheng Ye, Peilin Zhao, Mao Zheng, Hao Wang

Title, abstract, authors and date from arXiv's metadata (CC0); this paper is not in the Papers with Code archive (frozen 2025-07-28).

Role-playing agents based on large language models have been widely applied in areas such as personalized assistance and social simulation. Recent RL methods typically train on a fixed scenario pool collected before learning begins. This creates a distributional bottleneck: as the agent improves, the scenarios where it performs poorly also change, while the training distribution remains static. Therefore, we propose AdvRole, an adversarial context rewriting framework that turns role-playing RL into a closed-loop curriculum. AdvRole alternates between an Actor that learns to role-play and a Rewriter that edits character profiles and dialogue contexts into actor-specific hard scenarios. The Rewriter is trained with a performance-gap reward, which favors rewrites that reduce the current Actor's score relative to the original scenario. As a result, the scenario pool evolves with the Actor and continuously targets under-mastered regions of the character-context space. Experiments on three role-playing benchmarks covering English and Chinese, as well as a new multilingual benchmark we release, show that AdvRole consistently outperforms baselines.

PaperPDF

In Syntology View this paper on Syntology, its page in Syntology's graph. That page lists the repositories linked to the paper, the abstract and the calls for agents.

Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

For agents, on Syntology's MCP service (how to connect):

Code

verl-project/verl found in paper text by SyntologySyntology: no sample linked to this paper (harvested for another paper). Syntology's graph links these samples to FARCA: Fact-Aligned Reliability-Aware Credit Assignment for Reinforcement Learning... (arXiv:2608.24350; that paper's own run record: 5 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 1 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 7 unverified); nothing checks that they implement this paper's method. report

Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

A paper named beside a repository with no sample linked to this paper is shown with that paper's own run record, not this paper's: “ran” means executed on a synthesized input, not that the code is correct, and “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. Papers are listed in arXiv-id order, at most three per repository.

Code Syntology ran Syntology

Syntology holds the repository link but has not harvested or run code from it.

Results from the paper

The Papers with Code archive ends with its 2025-07-28 snapshot. This paper's arXiv identifier, 2609.28609, was issued in September 2026, after that date, so the archive has no leaderboard rows for it.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections