Methods › Natural Language Processing › Tokenizers › SentencePiece › Papers, page 7
SentencePiece
Papers archive 2025-07-28
archive papers tagged: 908 · with a code link: 438 · where Syntology ran a sample: 111 (96 with a run with no instrument failure, 15 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (111 of 908 tagged: 96 with a run with no instrument failure, 15 where every run was a failure of Syntology's instrument)
Page 7 of 10: papers 601 to 700 of 908, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
AraBART: a Pretrained Arabic Sequence-to-Sequence Model for Abstractive Summarization 21 Mar 2022 · 0 repositories · arXiv:2203.10945
-
DQ-BART: Efficient Sequence-to-Sequence Model via Joint Distillation and Quantization 21 Mar 2022 · 2 repositories · arXiv:2203.11239Syntology official: harvested, nothing ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
Coloring the Blank Slate: Pre-training Imparts a Hierarchical Inductive Bias to Sequence-to-sequence Models 17 Mar 2022 · 1 repository · arXiv:2203.09397Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 3 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
SciNLI: A Corpus for Natural Language Inference on Scientific Text 13 Mar 2022 · 1 repository · arXiv:2203.06728
-
IT5: Text-to-text Pretraining for Italian Language Understanding and Generation 7 Mar 2022 · 3 repositories · arXiv:2203.03759
-
Parameter-Efficient Mixture-of-Experts Architecture for Pre-trained Language Models 2 Mar 2022 · 2 repositories · arXiv:2203.01104Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
E-LANG: Energy-Based Joint Inferencing of Super and Swift Language Models 1 Mar 2022 · 0 repositories · arXiv:2203.00748
-
HyperPrompt: Prompt-based Task-Conditioning of Transformers 1 Mar 2022 · 0 repositories · arXiv:2203.00759
-
Mixture-of-Experts with Expert Choice Routing 18 Feb 2022 · 0 repositories · arXiv:2202.09368
-
Probing Pretrained Models of Source Code 16 Feb 2022 · 1 repository · arXiv:2202.08975
-
Predicting on the Edge: Identifying Where a Larger Model Does Better 15 Feb 2022 · 0 repositories · arXiv:2202.07652
-
A multi-task semi-supervised framework for Text2Graph & Graph2Text 12 Feb 2022 · 1 repository · arXiv:2202.06041
-
HaT5: Hate Language Identification using Text-to-Text Transfer Transformer 11 Feb 2022 · 0 repositories · arXiv:2202.05690
-
Logical Reasoning for Task Oriented Dialogue Systems 8 Feb 2022 · 0 repositories · arXiv:2202.04161
-
Regression Transformer: Concurrent sequence regression and generation for molecular language modeling 1 Feb 2022 · 1 repository · arXiv:2202.01338Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples)
-
A Comparative Study on Language Models for Task-Oriented Dialogue Systems 21 Jan 2022 · 1 repository · arXiv:2201.08687
-
Cheating Automatic Short Answer Grading: On the Adversarial Usage of Adjectives and Adverbs 20 Jan 2022 · 1 repository · arXiv:2201.08318
-
GAP-Gen: Guided Automatic Python Code Generation 19 Jan 2022 · 1 repository · arXiv:2201.08810
-
AllWOZ: Towards Multilingual Task-Oriented Dialog Systems for All 16 Jan 2022 · 0 repositories
-
AraBART: a Pretrained Arabic Sequence-to-Sequence Model for Abstractive Summarization 16 Jan 2022 · 0 repositories
-
Are Pretrained Multilingual Models Equally Fair Across Languages? 16 Jan 2022 · 0 repositories
-
Auto-regressive Text Generation with Pre-Trained Language Models: An Empirical Study on Question-type Short Text Generation 16 Jan 2022 · 0 repositories
-
Efficient Zero-Shot Semantic Parsing with Paraphrasing from Pretrained Language Models 16 Jan 2022 · 0 repositories
-
Exploring Example Selection for Few-shot Text-to-SQL Semantic Parsing 16 Jan 2022 · 0 repositories
-
Exploring the Low-Resource Transfer-Learning with mT5 model 16 Jan 2022 · 0 repositories
-
Natural Language Deduction through Search over Statement Compositions 16 Jan 2022 · 0 repositories · arXiv:2201.06028
-
Re2G: Retrieve, Rerank, Generate 16 Jan 2022 · 0 repositories
-
UnifiedSKG: Unifying and Multi-Tasking Structured Knowledge Grounding with Text-to-Text Language Models 16 Jan 2022 · 1 repository · arXiv:2201.05966Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
Comparison of biomedical relationship extraction methods and models for knowledge graph creation 5 Jan 2022 · 0 repositories · arXiv:2201.01647
-
The University of Texas at Dallas HLTRI's Participation in EPIC-QA: Searching for Entailed Questions Revealing Novel Answer Nuggets 28 Dec 2021 · 0 repositories · arXiv:2112.13946
-
Your Answer is Incorrect... Would you like to know why? Introducing a Bilingual Short Answer Feedback Dataset 17 Dec 2021 · 0 repositories
-
CrossSum: Beyond English-Centric Cross-Lingual Summarization for 1,500+ Language Pairs 16 Dec 2021 · 1 repository · arXiv:2112.08804
-
DREAM: Improving Situational QA by First Elaborating the Situation 16 Dec 2021 · 1 repository · arXiv:2112.08656Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
AllWOZ: Towards Multilingual Task-Oriented Dialog Systems for All 15 Dec 2021 · 0 repositories · arXiv:2112.08333
-
LongT5: Efficient Text-To-Text Transformer for Long Sequences 15 Dec 2021 · 4 repositories · arXiv:2112.07916Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Named entity recognition architecture combining contextual and global features 15 Dec 2021 · 1 repository · arXiv:2112.08033
-
Improving Compositional Generalization with Latent Structure and Data Augmentation 14 Dec 2021 · 2 repositories · arXiv:2112.07610
-
Do Data-based Curricula Work? 13 Dec 2021 · 0 repositories · arXiv:2112.06510
-
Towards Neural Functional Program Evaluation 9 Dec 2021 · 0 repositories · arXiv:2112.04630
-
Dynamic Token Normalization Improves Vision Transformers 5 Dec 2021 · 1 repository · arXiv:2112.02624Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample) · 1 pointer-only (licence)
-
Controlling Conditional Language Models without Catastrophic Forgetting 1 Dec 2021 · 2 repositories · arXiv:2112.00791Syntology official (archive's flag): 3 ran · 7 ran (of which 1 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 4 pointer-only (licence)
-
Searching for Efficient Transformers for Language Modeling 1 Dec 2021 · 0 repositories
-
A Comparative Study of Transformers on Word Sense Disambiguation 30 Nov 2021 · 0 repositories · arXiv:2111.15417
-
Chemical Identification and Indexing in PubMed Articles via BERT and Text-to-Text Approaches 30 Nov 2021 · 0 repositories · arXiv:2111.15622
-
Text Mining Drug/Chemical-Protein Interactions using an Ensemble of BERT and T5 Based Models 30 Nov 2021 · 0 repositories · arXiv:2111.15617
-
ExT5: Towards Extreme Multi-Task Scaling for Transfer Learning 22 Nov 2021 · 4 repositories · arXiv:2111.10952Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples)
-
A Comparative Study on Transfer Learning and Distance Metrics in Semantic Clustering over the COVID-19 Tweets 16 Nov 2021 · 0 repositories · arXiv:2111.08658
-
ANNA: Enhanced Language Representation for Question Answering 16 Nov 2021 · 0 repositories
-
Coloring the Blank Slate: Pre-training Imparts a Hierarchical Inductive Bias to Sequence-to-sequence Models 16 Nov 2021 · 0 repositories
-
CST5: Data augmentation for Code-Switched Semantic Parsing 16 Nov 2021 · 0 repositories
-
DAML-ST5: Low Resource Style Transfer via Domain Adaptive Meta Learning 16 Nov 2021 · 0 repositories
-
FRSUM: Towards Faithful Abstractive Summarization via Enhancing Factual Robustness 16 Nov 2021 · 0 repositories
-
Generation of News Articles from Tweets : An Experiment 16 Nov 2021 · 0 repositories
-
GLM: General Language Model Pretraining with Autoregressive Blank Infilling 16 Nov 2021 · 0 repositories
-
Improving Compositional Generalization with Self-Training for Data-to-Text Generation 16 Nov 2021 · 0 repositories
-
Interpreting the Robustness of Neural NLP Models to Textual Perturbations 16 Nov 2021 · 0 repositories
-
Life after BERT: What do Other Muppets Understand about Language? 16 Nov 2021 · 0 repositories
-
Neural Keyphrase Generation: Analysis and Evaluation 16 Nov 2021 · 0 repositories
-
The Power of Prompt Tuning for Low-Resource Semantic Parsing 16 Nov 2021 · 0 repositories
-
Calculating Question Similarity is Enough: A New Method for KBQA Tasks 15 Nov 2021 · 0 repositories · arXiv:2111.07658
-
Say What? Collaborative Pop Lyric Generation Using Multitask Transfer Learning 15 Nov 2021 · 0 repositories · arXiv:2111.07592
-
Automated question generation and question answering from Turkish texts 11 Nov 2021 · 1 repository · arXiv:2111.06476
-
IBERT: Idiom Cloze-style reading comprehension with Attention 5 Nov 2021 · 0 repositories · arXiv:2112.02994
-
Backdoor Pre-trained Models Can Transfer to All 30 Oct 2021 · 1 repository · arXiv:2111.00197
-
Ask me in your own words: paraphrasing for multitask question answering 27 Oct 2021 · 1 repository
-
Fast Model Editing at Scale 21 Oct 2021 · 3 repositories · arXiv:2110.11309Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
EncT5: A Framework for Fine-tuning T5 as Non-autoregressive Models 16 Oct 2021 · 1 repository · arXiv:2110.08426
-
Evaluation of Transfer Learning for Polish with a text-to-text model 16 Oct 2021 · 0 repositories
-
Improving Compositional Generalization with Self-Training for Data-to-Text Generation 16 Oct 2021 · 1 repository · arXiv:2110.08467
-
Multi-Task End-to-End Training Improves Conversational Recommendation 16 Oct 2021 · 0 repositories
-
Semantic Tokenizer for Enhanced Natural Language Processing 16 Oct 2021 · 0 repositories
-
Sharpness-Aware Minimization Improves Language Model Generalization 16 Oct 2021 · 0 repositories · arXiv:2110.08529
-
The Power of Prompt Tuning for Low-Resource Semantic Parsing 16 Oct 2021 · 0 repositories · arXiv:2110.08525
-
Interpreting the Robustness of Neural NLP Models to Textual Perturbations 14 Oct 2021 · 0 repositories · arXiv:2110.07159
-
LFPT5: A Unified Framework for Lifelong Few-shot Language Learning Based on Prompt Tuning of T5 14 Oct 2021 · 1 repository · arXiv:2110.07298Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified; the one sample that ran constructed an object rather than computing a result (of 3 harvested samples)
-
SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing 14 Oct 2021 · 6 repositories · arXiv:2110.07205Syntology community repositories only · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
KG-FiD: Infusing Knowledge Graph in Fusion-in-Decoder for Open-Domain Question Answering 8 Oct 2021 · 0 repositories · arXiv:2110.04330
-
DeepA2: A Modular Framework for Deep Argument Analysis with Pretrained Neural Text2Text Language Models 4 Oct 2021 · 1 repository · arXiv:2110.01509
-
Perhaps PTLMs Should Go to School -- A Task to Assess Open Book and Closed Book QA 4 Oct 2021 · 0 repositories · arXiv:2110.01552
-
Low Frequency Names Exhibit Bias and Overfitting in Contextualizing Language Models 1 Oct 2021 · 0 repositories · arXiv:2110.00672
-
Scale Efficiently: Insights from Pretraining and Finetuning Transformers 29 Sep 2021 · 0 repositories
-
TransSlowDown: Efficiency Attacks on Neural Machine Translation Systems 29 Sep 2021 · 0 repositories
-
What to Prioritize? Natural Language Processing for the Development of a Modern Bug Tracking Solution in Hardware Development 28 Sep 2021 · 0 repositories · arXiv:2109.13825
-
Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers 22 Sep 2021 · 3 repositories · arXiv:2109.10686Syntology community repositories only · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 4 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 12 harvested samples) · 1 pointer-only (licence)
-
Hierarchy-Aware T5 with Path-Adaptive Mask Mechanism for Hierarchical Text Classification 17 Sep 2021 · 0 repositories · arXiv:2109.08585
-
Primer: Searching for Efficient Transformers for Language Modeling 17 Sep 2021 · 4 repositories · arXiv:2109.08668Syntology official: no sample here; runs from other or unrecorded repositories · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 2 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Remixers: A Mixer-Transformer Architecture with Compositional Operators for Natural Language Understanding 17 Sep 2021 · 0 repositories
-
Language Models are Few-shot Multilingual Learners 16 Sep 2021 · 1 repository · arXiv:2109.07684Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Attention Is Indeed All You Need: Semantically Attention-Guided Decoding for Data-to-Text NLG 15 Sep 2021 · 1 repository · arXiv:2109.07043
-
Prefix-to-SQL: Text-to-SQL Generation from Incomplete User Questions 15 Sep 2021 · 0 repositories · arXiv:2109.13066
-
Topic Transferable Table Question Answering 15 Sep 2021 · 1 repository · arXiv:2109.07377
-
Exploring a Unified Sequence-To-Sequence Transformer for Medical Product Safety Monitoring in Social Media 13 Sep 2021 · 1 repository · arXiv:2109.05815
-
Not All Models Localize Linguistic Knowledge in the Same Place: A Layer-wise Probing on BERToids' Representations 13 Sep 2021 · 0 repositories · arXiv:2109.05958
-
Investigating Numeracy Learning Ability of a Text-to-Text Transfer Model 10 Sep 2021 · 1 repository · arXiv:2109.04672Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
PICARD: Parsing Incrementally for Constrained Auto-Regressive Decoding from Language Models 10 Sep 2021 · 3 repositories · arXiv:2109.05093Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
Zero-Shot Dialogue State Tracking via Cross-Task Transfer 10 Sep 2021 · 1 repository · arXiv:2109.04655
-
Retrieve, Caption, Generate: Visual Grounding for Enhancing Commonsense in Text Generation Models 8 Sep 2021 · 0 repositories · arXiv:2109.03892
-
General-Purpose Question-Answering with Macaw 6 Sep 2021 · 2 repositories · arXiv:2109.02593
-
CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation 2 Sep 2021 · 5 repositories · arXiv:2109.00859Syntology official: harvested, nothing ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (of 11 harvested samples)
-
AraT5: Text-to-Text Transformers for Arabic Language Generation 31 Aug 2021 · 1 repository · arXiv:2109.12068