Methods › Natural Language Processing › Subword Segmentation › GBST

Gradient-based Subword Tokenization Module

GBST

0 papers tagged archive 2025-07-28

archive 2025-07-28 Description, source and code snippet are the archive's method entry.

GBST, or Gradient-based Subword Tokenization Module, is a soft gradient-based subword tokenization module that automatically learns latent subword representations from characters in a data-driven fashion. Concretely, GBST enumerates candidate subword blocks and learns to score them in a position-wise fashion using a block scoring network.

GBST learns a position-wise soft selection over candidate subword blocks by scoring them with a scoring network. In contrast to prior tokenization-free methods, GBST learns interpretable latent subwords, which enables easy inspection of lexical representations and is more efficient than other byte-based models.

Source: Charformer: Fast Character Transformers via...

Papers archive 2025-07-28

The archive tags no paper with this method.

Tasks archive 2025-07-28

The archive attaches no task to a paper tagged with this method.

Usage over time archive 2025-07-28

No dated papers to chart.

Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).

Categories archive 2025-07-28

Subword Segmentation

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections