{"url":"/method/gtrxl","slug":"gtrxl","name":"GTrXL","full_name":"Gated Transformer-XL","full_name_withheld":false,"description_markdown":"**Gated Transformer-XL**, or **GTrXL**, is a [Transformer](https://paperswithcode.com/methods/category/transformers)-based architecture for reinforcement learning. It introduces architectural modifications that improve the stability and learning speed of the original Transformer and XL variant. Changes include:\r\n\r\n- Placing the [layer normalization](https://paperswithcode.com/method/layer-normalization) on only the input stream of the submodules. A key benefit to this reordering is that it now enables an identity map from the input of the transformer at the first layer to the output of the transformer after the last layer. This is in contrast to the canonical transformer, where there are a series of layer normalization operations that non-linearly transform the state encoding.\r\n- Replacing [residual connections](https://paperswithcode.com/method/residual-connection) with gating layers. The authors' experiments found that [GRUs](https://www.paperswithcode.com/method/gru) were the most effective form of gating.","description_state":"present","introduced_year":null,"introduced_by":{"title":"Stabilizing Transformers for Reinforcement Learning","paper":"/paper/stabilizing-transformers-for-reinforcement-1","first_author":"Emilio Parisotto","n_authors":13,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/stabilizing-transformers-for-reinforcement-1"},"source":{"url":"https://arxiv.org/abs/1910.06764v1","title":"Stabilizing Transformers for Reinforcement Learning","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Reinforcement Learning","area_id":"reinforcement-learning","collection":"RL Transformers","url":"/methods/category/rl-transformers","pwc_aliases":[]}],"n_papers_tagged":3,"archive_num_papers":3,"papers_newest_first":[{"paper":"/paper/recurrent-linear-transformers","title":"AGaLiTe: Approximate Gated Linear Transformers for Online Reinforcement Learning","date":"2023-10-24","arxiv_id":"2310.15719","n_code_links":2,"syntology":{"ran":7,"of":11,"unverified":4,"pointer_only":0}},{"paper":"/paper/coberl-contrastive-bert-for-reinforcement","title":"CoBERL: Contrastive BERT for Reinforcement Learning","date":"2021-07-12","arxiv_id":"2107.05431","n_code_links":2,"syntology":{"ran":3,"of":3,"unverified":0,"pointer_only":0}},{"paper":"/paper/stabilizing-transformers-for-reinforcement-1","title":"Stabilizing Transformers for Reinforcement Learning","date":"2019-10-13","arxiv_id":"1910.06764","n_code_links":5,"syntology":{"ran":3,"of":3,"unverified":0,"pointer_only":3}}],"papers_shown":3,"tasks":[{"task":"/task/reinforcement-learning","name":"Reinforcement Learning","papers":3},{"task":"/task/reinforcement-learning-2","name":"reinforcement-learning","papers":3},{"task":"/task/reinforcement-learning-1","name":"Reinforcement Learning (RL)","papers":2},{"task":"/task/diagnostic","name":"Diagnostic","papers":1},{"task":"/task/general-reinforcement-learning","name":"General Reinforcement Learning","papers":1},{"task":"/task/language-modeling","name":"Language Modeling","papers":1},{"task":"/task/language-modelling","name":"Language Modelling","papers":1},{"task":"/task/machine-translation","name":"Machine Translation","papers":1},{"task":"/task/partially-observable-reinforcement-learning","name":"Partially Observable Reinforcement Learning","papers":1}],"tasks_shown":9,"n_tasks":9,"usage_by_year":[{"year":"2019","papers":1},{"year":"2021","papers":1},{"year":"2023","papers":1}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/gtrxl"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}