{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/mvar-visual-autoregressive-modeling-with","title":"MVAR: Visual Autoregressive Modeling with Scale and Spatial Markovian Conditioning","arxiv_id":"2505.12742","date":"2025-05-19","proceeding":null,"authors":["Jinhua Zhang","Wei Long","Minghao Han","Weiyi You","Shuhang Gu"],"abstract":"Essential to visual generation is efficient modeling of visual data priors. Conventional next-token prediction methods define the process as learning the conditional probability distribution of successive tokens. Recently, next-scale prediction methods redefine the process to learn the distribution over multi-scale representations, significantly reducing generation latency. However, these methods condition each scale on all previous scales and require each token to consider all preceding tokens, exhibiting scale and spatial redundancy. To better model the distribution by mitigating redundancy, we propose Markovian Visual AutoRegressive modeling (MVAR), a novel autoregressive framework that introduces scale and spatial Markov assumptions to reduce the complexity of conditional probability modeling. Specifically, we introduce a scale-Markov trajectory that only takes as input the features of adjacent preceding scale for next-scale prediction, enabling the adoption of a parallel training strategy that significantly reduces GPU memory consumption. Furthermore, we propose spatial-Markov attention, which restricts the attention of each token to a localized neighborhood of size k at corresponding positions on adjacent scales, rather than attending to every token across these scales, for the pursuit of reduced modeling complexity. Building on these improvements, we reduce the computational complexity of attention calculation from O(N^2) to O(Nk), enabling training with just eight NVIDIA RTX 4090 GPUs and eliminating the need for KV cache during inference. Extensive experiments on ImageNet demonstrate that MVAR achieves comparable or superior performance with both small model trained from scratch and large fine-tuned models, while reducing the average GPU memory footprint by 3.0x.","url_abs":"https://arxiv.org/abs/2505.12742v1","url_pdf":"https://arxiv.org/pdf/2505.12742v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"mvar-visual-autoregressive-modeling-with","repo_url":"https://github.com/labshuhanggu/mvar","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":null,"task_name":"GPU"}],"methods":[{"method_slug":"attention","method_name":"Attention"},{"method_slug":"softmax","method_name":"Softmax"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2505.12742","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2505.12742"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/labshuhanggu/mvar","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran_draft_wrong":2,"ran_fixture":1,"ran":2,"unverified":3},"by_repo_kind":{"official":{"samples":8,"ran":5,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"9fcdaa6e423e8aa7","entry":"Normalize","repo":"labshuhanggu/mvar","repo_kind":"official","path":"models/basic_vae.py","file_url":"https://github.com/labshuhanggu/mvar/blob/HEAD/models/basic_vae.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"9fcdaa6e423e8aa7"}},{"code_sha256_prefix":"971ae8d2d8313e30","entry":"drop_path","repo":"labshuhanggu/mvar","repo_kind":"official","path":"models/helpers.py","file_url":"https://github.com/labshuhanggu/mvar/blob/HEAD/models/helpers.py","link_basis":"harvester_set","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"971ae8d2d8313e30"}},{"code_sha256_prefix":"e543491ea064c733","entry":"gumbel_softmax_with_rng","repo":"labshuhanggu/mvar","repo_kind":"official","path":"models/helpers.py","file_url":"https://github.com/labshuhanggu/mvar/blob/HEAD/models/helpers.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"e543491ea064c733"}},{"code_sha256_prefix":"3137073275f8c21a","entry":"nonlinearity","repo":"labshuhanggu/mvar","repo_kind":"official","path":"models/basic_vae.py","file_url":"https://github.com/labshuhanggu/mvar/blob/HEAD/models/basic_vae.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"3137073275f8c21a"}},{"code_sha256_prefix":"80371584448f2f6e","entry":"sample_with_top_k_top_p_","repo":"labshuhanggu/mvar","repo_kind":"official","path":"models/helpers.py","file_url":"https://github.com/labshuhanggu/mvar/blob/HEAD/models/helpers.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"80371584448f2f6e"}},{"code_sha256_prefix":"6aec3f6a3fa7cc1a","entry":"compile_model","repo":"labshuhanggu/mvar","repo_kind":"official","path":"main_cache.py","file_url":"https://github.com/labshuhanggu/mvar/blob/HEAD/main_cache.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"6aec3f6a3fa7cc1a"}},{"code_sha256_prefix":"ca42de1cd2be30bf","entry":"generate_backward_attn_mask_only_next","repo":"labshuhanggu/mvar","repo_kind":"official","path":"models/mvar.py","file_url":"https://github.com/labshuhanggu/mvar/blob/HEAD/models/mvar.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"ca42de1cd2be30bf"}},{"code_sha256_prefix":"793c7e2489dd8dfc","entry":"make_attn","repo":"labshuhanggu/mvar","repo_kind":"official","path":"models/basic_vae.py","file_url":"https://github.com/labshuhanggu/mvar/blob/HEAD/models/basic_vae.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"793c7e2489dd8dfc"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}