{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/tokenization-and-the-noiseless-channel","title":"Tokenization and the Noiseless Channel","arxiv_id":"2306.16842","date":"2023-06-29","proceeding":null,"authors":["Vilém Zouhar","Clara Meister","Juan Luis Gastaldi","Li Du","Mrinmaya Sachan","Ryan Cotterell"],"abstract":"Subword tokenization is a key part of many NLP pipelines. However, little is known about why some tokenizer and hyperparameter combinations lead to better downstream model performance than others. We propose that good tokenizers lead to \\emph{efficient} channel usage, where the channel is the means by which some input is conveyed to the model and efficiency can be quantified in information-theoretic terms as the ratio of the Shannon entropy to the maximum possible entropy of the token distribution. Yet, an optimal encoding according to Shannon entropy assigns extremely long codes to low-frequency tokens and very short codes to high-frequency tokens. Defining efficiency in terms of R\\'enyi entropy, on the other hand, penalizes distributions with either very high or very low-frequency tokens. In machine translation, we find that across multiple tokenizers, the R\\'enyi entropy with $\\alpha = 2.5$ has a very strong correlation with \\textsc{Bleu}: $0.78$ in comparison to just $-0.32$ for compressed length.","url_abs":"https://arxiv.org/abs/2306.16842v1","url_pdf":"https://arxiv.org/pdf/2306.16842v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"tokenization-and-the-noiseless-channel","repo_url":"https://github.com/zouharvi/tokenization-scorer","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"machine-translation","task_name":"Machine Translation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2306.16842","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2306.16842"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/zouharvi/tokenization-scorer","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"unverified":2},"by_repo_kind":{"official":{"samples":2,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"4b7cef617d5cca71","entry":"get_metric","repo":"zouharvi/tokenization-scorer","repo_kind":"official","path":"tokenization_scorer/metrics.py","file_url":"https://github.com/zouharvi/tokenization-scorer/blob/HEAD/tokenization_scorer/metrics.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"4b7cef617d5cca71"}},{"code_sha256_prefix":"bf3b709a7d195f31","entry":"get_prob_distribution","repo":"zouharvi/tokenization-scorer","repo_kind":"official","path":"tokenization_scorer/metrics.py","file_url":"https://github.com/zouharvi/tokenization-scorer/blob/HEAD/tokenization_scorer/metrics.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"bf3b709a7d195f31"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}