Methods › Natural Language Processing › Tokenizers › SentencePiece › Papers, page 10
SentencePiece
Papers archive 2025-07-28
archive papers tagged: 908 · with a code link: 438 · where Syntology ran a sample: 111 (96 with a run with no instrument failure, 15 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (111 of 908 tagged: 96 with a run with no instrument failure, 15 where every run was a failure of Syntology's instrument)
Page 10 of 10: papers 901 to 908 of 908, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Reasoning Over Semantic-Level Graph for Fact Checking 9 Sep 2019 · 0 repositories · arXiv:1909.03745
-
Transfer Learning Robustness in Multi-Class Categorization by Fine-Tuning Pre-Trained Contextualized Language Models 8 Sep 2019 · 1 repository · arXiv:1909.03564
-
Integrating Multimodal Information in Large Pretrained Transformers 15 Aug 2019 · 1 repository · arXiv:1908.05787
-
Scalable Attentive Sentence-Pair Modeling via Distilled Sentence Embedding 14 Aug 2019 · 1 repository · arXiv:1908.05161
-
BioFLAIR: Pretrained Pooled Contextualized Embeddings for Biomedical Sequence Labeling Tasks 13 Aug 2019 · 1 repository · arXiv:1908.05760
-
ERNIE 2.0: A Continual Pre-training Framework for Language Understanding 29 Jul 2019 · 3 repositories · arXiv:1907.12412Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
XLNet: Generalized Autoregressive Pretraining for Language Understanding 19 Jun 2019 · 27 repositories · arXiv:1906.08237Syntology official (archive's flag): 1 ran · 15 ran (of which 2 constructed an object rather than computing a result; 10 with no instrument failure: 2 honoured, 0 violated, 8 with no contract checked; 5 where Syntology's instrument failed) · 9 unverified (of 24 harvested samples) · 4 pointer-only (licence)
-
SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing 19 Aug 2018 · 3 repositories · arXiv:1808.06226