{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2605-06683","title":"Toeplitz MLP Mixers are Low Complexity, Information-Rich Sequence Models","arxiv_id":"2605.06683","date":"2026-04-24","proceeding":null,"authors":["Benjamin L. Badger","Ethan Roland"],"abstract":"Transformer-based large language models are in some respects limited by the quadratic time and space computational complexity of attention. We introduce the Toeplitz MLP Mixer (TMM), a transformer-like architecture that swaps attention for triangular-masked Toeplitz matrix multiplication over the sequence dimension resulting in $\\mathcal{O} (dn \\log n)$ time and $\\mathcal O(dn)$ space complexity during training and $\\mathcal O(dn)$ time and space at inference prefill. Despite the lack of sophisticated input modulation or state maintenance present in other sub-quadratic architectures, TMMs yield greater training efficiency in terms of loss achieved per compute and device memory. We demonstrate that TMMs are capable of retaining more input information resulting in improved copying ability, which we argue results from a lack of architectural biases. Consistent with higher input information retention, TMMs exhibit superior information retrieval and in-context learning benchmark accuracy compared to comparable architectures. We conclude with an analysis from the perspective of operator index theory and show that, counterintuitively, trained Toeplitz layers of causal non-invertible models are more likely to be invertible or nearly so than models that are actually invertible over their inputs.","url_abs":"https://arxiv.org/abs/2605.06683","url_pdf":"https://arxiv.org/pdf/2605.06683","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2605.06683","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2605.06683"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/blbadger/lm-evaluation-harness","reach":null},{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/ethan-w-roland/ToeplitzMixers","reach":null}],"summary":{"ran":4,"unverified":3},"by_repo_kind":{"found_in_text":{"samples":7,"ran":4,"repositories":2}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"f5f5bcda6dc54604","entry":"FFTToeplitzCausalLinear","repo":"ethan-w-roland/ToeplitzMixers","repo_kind":"found_in_text","path":"experiments/003-multiheaded/toep_mixer_multiheaded.py","file_url":"https://github.com/ethan-w-roland/ToeplitzMixers/blob/HEAD/experiments/003-multiheaded/toep_mixer_multiheaded.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"f5f5bcda6dc54604"}},{"code_sha256_prefix":"0cd9fc5663ec8377","entry":"ToeplitzCausalLinear","repo":"ethan-w-roland/ToeplitzMixers","repo_kind":"found_in_text","path":"experiments/003-multiheaded/toep_mixer_multiheaded.py","file_url":"https://github.com/ethan-w-roland/ToeplitzMixers/blob/HEAD/experiments/003-multiheaded/toep_mixer_multiheaded.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"0cd9fc5663ec8377"}},{"code_sha256_prefix":"0ce645a05f6dcc8b","entry":"ToeplitzCausalLinear","repo":"blbadger/lm-evaluation-harness","repo_kind":"found_in_text","path":"lm_eval/models/tmm_model.py","file_url":"https://github.com/blbadger/lm-evaluation-harness/blob/HEAD/lm_eval/models/tmm_model.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"0ce645a05f6dcc8b"}},{"code_sha256_prefix":"ba8c9962b742e784","entry":"ToeplitzHeads","repo":"ethan-w-roland/ToeplitzMixers","repo_kind":"found_in_text","path":"experiments/003-multiheaded/toep_mixer_multiheaded.py","file_url":"https://github.com/ethan-w-roland/ToeplitzMixers/blob/HEAD/experiments/003-multiheaded/toep_mixer_multiheaded.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"ba8c9962b742e784"}},{"code_sha256_prefix":"0d0a7957d51396c1","entry":"MLPMixer","repo":"blbadger/lm-evaluation-harness","repo_kind":"found_in_text","path":"lm_eval/models/tmm_model.py","file_url":"https://github.com/blbadger/lm-evaluation-harness/blob/HEAD/lm_eval/models/tmm_model.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"0d0a7957d51396c1"}},{"code_sha256_prefix":"741adba8945a249a","entry":"MixerBlock","repo":"blbadger/lm-evaluation-harness","repo_kind":"found_in_text","path":"lm_eval/models/tmm_model.py","file_url":"https://github.com/blbadger/lm-evaluation-harness/blob/HEAD/lm_eval/models/tmm_model.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"741adba8945a249a"}},{"code_sha256_prefix":"61b29a9588d52a63","entry":"ToeplitzHeads","repo":"blbadger/lm-evaluation-harness","repo_kind":"found_in_text","path":"lm_eval/models/tmm_model.py","file_url":"https://github.com/blbadger/lm-evaluation-harness/blob/HEAD/lm_eval/models/tmm_model.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"61b29a9588d52a63"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.LG","source":"arxiv_2026.jsonl"},"syntology_extracted_results":null}