{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/in-context-learning-and-occam-s-razor","title":"In-context learning and Occam's razor","arxiv_id":"2410.14086","date":"2024-10-17","proceeding":null,"authors":["Eric Elmoznino","Tom Marty","Tejas Kasetty","Leo Gagnon","Sarthak Mittal","Mahan Fathi","Dhanya Sridhar","Guillaume Lajoie"],"abstract":"A central goal of machine learning is generalization. While the No Free Lunch Theorem states that we cannot obtain theoretical guarantees for generalization without further assumptions, in practice we observe that simple models which explain the training data generalize best: a principle called Occam's razor. Despite the need for simple models, most current approaches in machine learning only minimize the training error, and at best indirectly promote simplicity through regularization or architecture design. Here, we draw a connection between Occam's razor and in-context learning: an emergent ability of certain sequence models like Transformers to learn at inference time from past observations in a sequence. In particular, we show that the next-token prediction loss used to train in-context learners is directly equivalent to a data compression technique called prequential coding, and that minimizing this loss amounts to jointly minimizing both the training error and the complexity of the model that was implicitly learned from context. Our theory and the empirical experiments we use to support it not only provide a normative account of in-context learning, but also elucidate the shortcomings of current in-context learning methods, suggesting ways in which they can be improved. We make our code available at https://github.com/3rdCore/PrequentialCode.","url_abs":"https://arxiv.org/abs/2410.14086v3","url_pdf":"https://arxiv.org/pdf/2410.14086v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"in-context-learning-and-occam-s-razor","repo_url":"https://github.com/3rdcore/prequentialcode","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"CC0-1.0"}}],"tasks":[{"task_slug":"data-compression","task_name":"Data Compression"},{"task_slug":"in-context-learning","task_name":"In-Context Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2410.14086","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2410.14086"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/3rdcore/prequentialcode","reach":{"status":"ok","spdx":"CC0-1.0"}}],"summary":{"ran":3,"unverified":1},"by_repo_kind":{"official":{"samples":4,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"7b19c6e7617f256c","entry":"custom_collate_fn","repo":"3rdcore/prequentialcode","repo_kind":"official","path":"datasets/interfaces.py","file_url":"https://github.com/3rdcore/prequentialcode/blob/HEAD/datasets/interfaces.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"CC0-1.0","inline_ok":true,"mcp_get_code":{"code_sha256":"7b19c6e7617f256c"}},{"code_sha256_prefix":"adf6ca81990a0066","entry":"fastfood_torched_batched","repo":"3rdcore/prequentialcode","repo_kind":"official","path":"models/utils.py","file_url":"https://github.com/3rdcore/prequentialcode/blob/HEAD/models/utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"CC0-1.0","inline_ok":true,"mcp_get_code":{"code_sha256":"adf6ca81990a0066"}},{"code_sha256_prefix":"5120f3d68641ad15","entry":"make_fastfood_vars","repo":"3rdcore/prequentialcode","repo_kind":"official","path":"models/utils.py","file_url":"https://github.com/3rdcore/prequentialcode/blob/HEAD/models/utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"CC0-1.0","inline_ok":true,"mcp_get_code":{"code_sha256":"5120f3d68641ad15"}},{"code_sha256_prefix":"8a8b34172cffc441","entry":"fast_walsh_hadamard_torched_batched","repo":"3rdcore/prequentialcode","repo_kind":"official","path":"models/utils.py","file_url":"https://github.com/3rdcore/prequentialcode/blob/HEAD/models/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"CC0-1.0","inline_ok":true,"mcp_get_code":{"code_sha256":"8a8b34172cffc441"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}