{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/quantum-doubly-stochastic-transformers","title":"Quantum Doubly Stochastic Transformers","arxiv_id":"2504.16275","date":"2025-04-22","proceeding":null,"authors":["Jannis Born","Filip Skogh","Kahn Rhrissorrakrai","Filippo Utro","Nico Wagner","Aleksandros Sobczyk"],"abstract":"At the core of the Transformer, the Softmax normalizes the attention matrix to be right stochastic. Previous research has shown that this often destabilizes training and that enforcing the attention matrix to be doubly stochastic (through Sinkhorn's algorithm) consistently improves performance across different tasks, domains and Transformer flavors. However, Sinkhorn's algorithm is iterative, approximative, non-parametric and thus inflexible w.r.t. the obtained doubly stochastic matrix (DSM). Recently, it has been proven that DSMs can be obtained with a parametric quantum circuit, yielding a novel quantum inductive bias for DSMs with no known classical analogue. Motivated by this, we demonstrate the feasibility of a hybrid classical-quantum doubly stochastic Transformer (QDSFormer) that replaces the Softmax in the self-attention layer with a variational quantum circuit. We study the expressive power of the circuit and find that it yields more diverse DSMs that better preserve information than classical operators. Across multiple small-scale object recognition tasks, we find that our QDSFormer consistently surpasses both a standard Vision Transformer and other doubly stochastic Transformers. Beyond the established Sinkformer, this comparison includes a novel quantum-inspired doubly stochastic Transformer (based on QR decomposition) that can be of independent interest. The QDSFormer also shows improved training stability and lower performance variation suggesting that it may mitigate the notoriously unstable training of ViTs on small-scale data.","url_abs":"https://arxiv.org/abs/2504.16275v1","url_pdf":"https://arxiv.org/pdf/2504.16275v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"inductive-bias","task_name":"Inductive Bias"},{"task_slug":"object-recognition","task_name":"Object Recognition"}],"methods":[{"method_slug":"absolute-position-encodings","method_name":"Absolute Position Encodings"},{"method_slug":"adam","method_name":"Adam"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"bpe","method_name":"BPE"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"label-smoothing","method_name":"Label Smoothing"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"position-wise-feed-forward-layer","method_name":"Position-Wise Feed-Forward Layer"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"transformer","method_name":"Transformer"},{"method_slug":"vision-transformer","method_name":"Vision Transformer"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2504.16275","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2504.16275"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/michaelsdr/sinkformers","reach":null},{"provenance":"deterministic:regex_extraction","url":"https://github.com/boschresearch/eurekaMoments","reach":null}],"summary":{"ran":2,"unverified":2},"by_repo_kind":{"found_in_text":{"samples":4,"ran":2,"repositories":2}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":1,"samples":[{"code_sha256_prefix":"27d7b9fb23e6948a","entry":"Attention","repo":"boschresearch/eurekaMoments","repo_kind":"found_in_text","path":"vit.py","file_url":"https://github.com/boschresearch/eurekaMoments/blob/HEAD/vit.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"AGPL-3.0","inline_ok":false,"mcp_get_code":{"code_sha256":"27d7b9fb23e6948a"}},{"code_sha256_prefix":"0541009c68476312","entry":"SinkhornDistance","repo":"michaelsdr/sinkformers","repo_kind":"found_in_text","path":"nlp-tutorial/text-classification-transformer/model_sym.py","file_url":"https://github.com/michaelsdr/sinkformers/blob/HEAD/nlp-tutorial/text-classification-transformer/model_sym.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"0541009c68476312"}},{"code_sha256_prefix":"5bcf16df9e4b0c26","entry":"MultiHeadAttention","repo":"michaelsdr/sinkformers","repo_kind":"found_in_text","path":"nlp-tutorial/text-classification-transformer/model_sym.py","file_url":"https://github.com/michaelsdr/sinkformers/blob/HEAD/nlp-tutorial/text-classification-transformer/model_sym.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"5bcf16df9e4b0c26"}},{"code_sha256_prefix":"92b973daf651dcf6","entry":"ScaledDotProductAttention","repo":"michaelsdr/sinkformers","repo_kind":"found_in_text","path":"nlp-tutorial/text-classification-transformer/model_sym.py","file_url":"https://github.com/michaelsdr/sinkformers/blob/HEAD/nlp-tutorial/text-classification-transformer/model_sym.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"92b973daf651dcf6"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}