{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2603-24139","title":"Tutor-Student Reinforcement Learning: A Dynamic Curriculum for Robust Deepfake Detection","arxiv_id":"2603.24139","date":"2026-03-25","proceeding":null,"authors":["Zhanhe Lei","Zhongyuan Wang","Jikang Cheng","Baojin Huang","Yuhong Yang","Zhen Han","Chao Liang","Dengpan Ye"],"abstract":"Standard supervised training for deepfake detection treats all samples with uniform importance, which can be suboptimal for learning robust and generalizable features. In this work, we propose a novel Tutor-Student Reinforcement Learning (TSRL) framework to dynamically optimize the training curriculum. Our method models the training process as a Markov Decision Process where a ``Tutor'' agent learns to guide a ``Student'' (the deepfake detector). The Tutor, implemented as a Proximal Policy Optimization (PPO) agent, observes a rich state representation for each training sample, encapsulating not only its visual features but also its historical learning dynamics, such as EMA loss and forgetting counts. Based on this state, the Tutor takes an action by assigning a continuous weight (0-1) to the sample's loss, thereby dynamically re-weighting the training batch. The Tutor is rewarded based on the Student's immediate performance change, specifically rewarding transitions from incorrect to correct predictions. This strategy encourages the Tutor to learn a curriculum that prioritizes high-value samples, such as hard-but-learnable examples, leading to a more efficient and effective training process. We demonstrate that this adaptive curriculum improves the Student's generalization capabilities against unseen manipulation techniques compared to traditional training methods. Code is available at https://github.com/wannac1/TSRL.","url_abs":"https://arxiv.org/abs/2603.24139","url_pdf":"https://arxiv.org/pdf/2603.24139","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2603.24139","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2603.24139"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/wannac1/TSRL","reach":null}],"summary":{"ran":2,"unverified":3},"by_repo_kind":{"found_in_text":{"samples":5,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":5,"samples":[{"code_sha256_prefix":"e54f5d97d7a522af","entry":"ActorCritic","repo":"wannac1/TSRL","repo_kind":"found_in_text","path":"training/agents/tutor_grpo.py","file_url":"https://github.com/wannac1/TSRL/blob/HEAD/training/agents/tutor_grpo.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"e54f5d97d7a522af"}},{"code_sha256_prefix":"ac5a79d84456dbca","entry":"RolloutBuffer","repo":"wannac1/TSRL","repo_kind":"found_in_text","path":"training/agents/tutor_grpo.py","file_url":"https://github.com/wannac1/TSRL/blob/HEAD/training/agents/tutor_grpo.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"ac5a79d84456dbca"}},{"code_sha256_prefix":"8231fc53cae3f502","entry":"ActorCriticGNN","repo":"wannac1/TSRL","repo_kind":"found_in_text","path":"training/agents/tutor_grpo.py","file_url":"https://github.com/wannac1/TSRL/blob/HEAD/training/agents/tutor_grpo.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"8231fc53cae3f502"}},{"code_sha256_prefix":"5414a73300368bc6","entry":"TutorGRPO","repo":"wannac1/TSRL","repo_kind":"found_in_text","path":"training/agents/tutor_grpo.py","file_url":"https://github.com/wannac1/TSRL/blob/HEAD/training/agents/tutor_grpo.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"5414a73300368bc6"}},{"code_sha256_prefix":"5ce2ff40c574e69c","entry":"TutorPPO","repo":"wannac1/TSRL","repo_kind":"found_in_text","path":"training/agents/tutor_grpo.py","file_url":"https://github.com/wannac1/TSRL/blob/HEAD/training/agents/tutor_grpo.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"5ce2ff40c574e69c"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.LG","source":"arxiv_2026.jsonl"},"syntology_extracted_results":null}