{"url":"/method/permuteformer","slug":"permuteformer","name":"PermuteFormer","full_name":"PermuteFormer","full_name_withheld":false,"description_markdown":"**PermuteFormer** is a [Performer](https://paperswithcode.com/method/performer)-based model with relative position encoding that scales linearly on long sequences. PermuteFormer applies position-dependent transformation on queries and keys to encode positional information into the attention module. This transformation is carefully crafted so that the final output of self-attention is not affected by absolute positions of tokens.\r\n\r\nEach token’s query / key feature is illustrated as a row of blocks in the figure, and its elements are marked with different colors. The position-aware permutation permutes elements of each token’s query / key feature along the head size dimension in each attention head. Depending on the token’s position, the permutation applied to query / key feature is different.","description_state":"present","introduced_year":null,"introduced_by":{"title":"PermuteFormer: Efficient Relative Position Encoding for Long Sequences","paper":"/paper/permuteformer-efficient-relative-position","first_author":"Peng Chen","n_authors":1,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/permuteformer-efficient-relative-position"},"source":{"url":"https://arxiv.org/abs/2109.02377v2","title":"PermuteFormer: Efficient Relative Position Encoding for Long Sequences","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Transformers","url":"/methods/category/transformers","pwc_aliases":[]}],"n_papers_tagged":1,"archive_num_papers":1,"papers_newest_first":[{"paper":"/paper/permuteformer-efficient-relative-position","title":"PermuteFormer: Efficient Relative Position Encoding for Long Sequences","date":"2021-09-06","arxiv_id":"2109.02377","n_code_links":1,"syntology":{"ran":1,"of":1,"unverified":0,"pointer_only":1}}],"papers_shown":1,"tasks":[{"task":"/task/language-modeling","name":"Language Modeling","papers":1},{"task":"/task/language-modelling","name":"Language Modelling","papers":1},{"task":null,"name":"Position","papers":1}],"tasks_shown":3,"n_tasks":3,"usage_by_year":[{"year":"2021","papers":1}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/permuteformer"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}