{"url":"/method/gbst","slug":"gbst","name":"GBST","full_name":"Gradient-based Subword Tokenization Module","full_name_withheld":false,"description_markdown":"**GBST**, or **Gradient-based Subword Tokenization Module**, is a soft gradient-based subword tokenization module that automatically learns latent subword representations from characters in a data-driven fashion. Concretely, GBST enumerates candidate subword blocks and learns to score them in a position-wise fashion using a block scoring network.  \r\n\r\nGBST learns a position-wise soft selection over candidate subword blocks by scoring them with a scoring network. In contrast to prior tokenization-free methods, GBST learns interpretable latent subwords, which enables easy inspection of lexical representations and is more efficient than other byte-based models.","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"https://arxiv.org/abs/2106.12672v3","title":"Charformer: Fast Character Transformers via Gradient-based Subword Tokenization","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Subword Segmentation","url":"/methods/category/subword-segmentation","pwc_aliases":[]}],"n_papers_tagged":0,"archive_num_papers":null,"papers_newest_first":[],"papers_shown":0,"tasks":[],"tasks_shown":0,"n_tasks":0,"usage_by_year":[],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/gbst"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}