{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2605-22644","title":"Why SGD is not Brownian Motion: A New Perspective on Stochastic Dynamics","arxiv_id":"2605.22644","date":"2026-05-21","proceeding":null,"authors":["Igor Ignashin","Anna Radovskaya","Andrew Semenov","Egor Lopatin","Stanislav Potapov","Aleksandr Kovalenko","Andrey Veprikov","Aleksandr Shestakov","Andrey Leonidov","Aleksandr Beznosikov"],"abstract":"Stochastic Gradient Descent (SGD) is commonly modeled as a Langevin process, assuming that minibatch noise acts as Brownian motion. However, this approximation relies on a continuous-time limit and a sqrt(eta) noise scaling that does not match the discrete SGD update at finite learning rate. In this work, we propose an alternative formulation of SGD as deterministic dynamics in a fluctuating loss landscape induced by minibatch sampling. Starting directly from the discrete update, we derive a master equation for the parameter distribution and obtain a discrete Fokker--Planck equation that differs from the standard Langevin form at order eta^2. Using this framework, we analyze SGD dynamics near critical points of the loss. We show that the behavior decomposes along the eigenbasis of the mean Hessian into qualitatively distinct regimes. In particular, nearly-flat directions do not admit a stationary distribution: the variance grows over time, corresponding to effective diffusion along valleys with a coefficient proportional to the learning rate. We provide empirical evidence supporting these predictions on neural network models in computer vision and natural language processing, observing a clear qualitative separation between confined and diffusive modes.","url_abs":"https://arxiv.org/abs/2605.22644","url_pdf":"https://arxiv.org/pdf/2605.22644","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2605.22644","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2605.22644"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/brain-lab-research/SGDiffusion","reach":null}],"summary":{"ran":13,"unverified":4},"by_repo_kind":{"found_in_text":{"samples":17,"ran":13,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":17,"samples":[{"code_sha256_prefix":"950b82351b146029","entry":"alpha_estimator","repo":"brain-lab-research/SGDiffusion","repo_kind":"found_in_text","path":"src/visualise_utils.py","file_url":"https://github.com/brain-lab-research/SGDiffusion/blob/HEAD/src/visualise_utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"950b82351b146029"}},{"code_sha256_prefix":"3063a3c8f1afe8f6","entry":"alpha_estimator2","repo":"brain-lab-research/SGDiffusion","repo_kind":"found_in_text","path":"src/visualise_utils.py","file_url":"https://github.com/brain-lab-research/SGDiffusion/blob/HEAD/src/visualise_utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"3063a3c8f1afe8f6"}},{"code_sha256_prefix":"dfdedfdb8e7de832","entry":"compute_gradient_distances","repo":"brain-lab-research/SGDiffusion","repo_kind":"found_in_text","path":"src/statistic.py","file_url":"https://github.com/brain-lab-research/SGDiffusion/blob/HEAD/src/statistic.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"dfdedfdb8e7de832"}},{"code_sha256_prefix":"111ce930928fefee","entry":"flatten_grads","repo":"brain-lab-research/SGDiffusion","repo_kind":"found_in_text","path":"src/datamodelopt/core/tensor_utils.py","file_url":"https://github.com/brain-lab-research/SGDiffusion/blob/HEAD/src/datamodelopt/core/tensor_utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"111ce930928fefee"}},{"code_sha256_prefix":"c7d36593f7e19c5b","entry":"flatten_params","repo":"brain-lab-research/SGDiffusion","repo_kind":"found_in_text","path":"src/datamodelopt/core/tensor_utils.py","file_url":"https://github.com/brain-lab-research/SGDiffusion/blob/HEAD/src/datamodelopt/core/tensor_utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"c7d36593f7e19c5b"}},{"code_sha256_prefix":"a72e1393e15cbac1","entry":"format_time","repo":"brain-lab-research/SGDiffusion","repo_kind":"found_in_text","path":"src/datamodelopt/core/logging.py","file_url":"https://github.com/brain-lab-research/SGDiffusion/blob/HEAD/src/datamodelopt/core/logging.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"a72e1393e15cbac1"}},{"code_sha256_prefix":"340a7bfff5f1f4a5","entry":"get_device","repo":"brain-lab-research/SGDiffusion","repo_kind":"found_in_text","path":"src/datamodelopt/core/config.py","file_url":"https://github.com/brain-lab-research/SGDiffusion/blob/HEAD/src/datamodelopt/core/config.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"340a7bfff5f1f4a5"}},{"code_sha256_prefix":"cbc497fd8c503d83","entry":"get_generator","repo":"brain-lab-research/SGDiffusion","repo_kind":"found_in_text","path":"src/datamodelopt/core/seed.py","file_url":"https://github.com/brain-lab-research/SGDiffusion/blob/HEAD/src/datamodelopt/core/seed.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"cbc497fd8c503d83"}},{"code_sha256_prefix":"95ce2645335fb4b6","entry":"get_param_count","repo":"brain-lab-research/SGDiffusion","repo_kind":"found_in_text","path":"src/datamodelopt/core/tensor_utils.py","file_url":"https://github.com/brain-lab-research/SGDiffusion/blob/HEAD/src/datamodelopt/core/tensor_utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"95ce2645335fb4b6"}},{"code_sha256_prefix":"cb3865cbb3050b3c","entry":"get_torch_dtype","repo":"brain-lab-research/SGDiffusion","repo_kind":"found_in_text","path":"src/datamodelopt/core/config.py","file_url":"https://github.com/brain-lab-research/SGDiffusion/blob/HEAD/src/datamodelopt/core/config.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"cb3865cbb3050b3c"}},{"code_sha256_prefix":"e7898b70f6748622","entry":"init_data","repo":"brain-lab-research/SGDiffusion","repo_kind":"found_in_text","path":"src/visualise_utils.py","file_url":"https://github.com/brain-lab-research/SGDiffusion/blob/HEAD/src/visualise_utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"e7898b70f6748622"}},{"code_sha256_prefix":"645eff8e04b31c3e","entry":"load_model","repo":"brain-lab-research/SGDiffusion","repo_kind":"found_in_text","path":"src/datamodelopt/core/checkpointing.py","file_url":"https://github.com/brain-lab-research/SGDiffusion/blob/HEAD/src/datamodelopt/core/checkpointing.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"645eff8e04b31c3e"}},{"code_sha256_prefix":"d0d3b43ce3b9e921","entry":"train","repo":"brain-lab-research/SGDiffusion","repo_kind":"found_in_text","path":"src/statistic.py","file_url":"https://github.com/brain-lab-research/SGDiffusion/blob/HEAD/src/statistic.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"d0d3b43ce3b9e921"}},{"code_sha256_prefix":"43368a0a5b49f9e3","entry":"CIFAR","repo":"brain-lab-research/SGDiffusion","repo_kind":"found_in_text","path":"src/utils.py","file_url":"https://github.com/brain-lab-research/SGDiffusion/blob/HEAD/src/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"43368a0a5b49f9e3"}},{"code_sha256_prefix":"6b3e2abd8b904f36","entry":"MNIST","repo":"brain-lab-research/SGDiffusion","repo_kind":"found_in_text","path":"src/utils.py","file_url":"https://github.com/brain-lab-research/SGDiffusion/blob/HEAD/src/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"6b3e2abd8b904f36"}},{"code_sha256_prefix":"27c7edb2a414117f","entry":"evaluate_sgd_update","repo":"brain-lab-research/SGDiffusion","repo_kind":"found_in_text","path":"src/statistic.py","file_url":"https://github.com/brain-lab-research/SGDiffusion/blob/HEAD/src/statistic.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"27c7edb2a414117f"}},{"code_sha256_prefix":"40f7bef53b52ad30","entry":"load_similar_mnist_data","repo":"brain-lab-research/SGDiffusion","repo_kind":"found_in_text","path":"src/utils.py","file_url":"https://github.com/brain-lab-research/SGDiffusion/blob/HEAD/src/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"40f7bef53b52ad30"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.LG","source":"arxiv_2026.jsonl"},"syntology_extracted_results":{"kind":"leaderboard_placements","source":"Syntology's leaderboard-shaped extractor over the paper's own arXiv-HTML tables: a model pointed at a cell, the number was read from that cell and checked against the board's metric, dataset, split and scale, and an independent check accepted the entry; not reviewed by the paper's authors or the archive's editors","extractor_model":"global.anthropic.claude-sonnet-4-5-20250929-v1:0","verifier_model":null,"prompt_sha":"fa63d4bb9d755694","coverage":{"sentence":"Syntology has checked 6,264 of the 9,581 papers on this site that are newer than the archive; results from the others appear after they are checked.","papers_newer_than_archive":9581,"papers_checked":6264},"entries":[],"not_placed":{"boards":0,"rejected_by_independent_check":0,"refused_by_a_rule":0,"check_did_not_answer":0,"proposed_without_a_cell":0,"declined_by_site":0}}}