Methods › Natural Language Processing › Tokenizers › SentencePiece › Papers, page 8
SentencePiece
Papers archive 2025-07-28
archive papers tagged: 908 · with a code link: 438 · where Syntology ran a sample: 111 (96 with a run with no instrument failure, 15 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (111 of 908 tagged: 96 with a run with no instrument failure, 15 where every run was a failure of Syntology's instrument)
Page 8 of 10: papers 701 to 800 of 908, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Automated Mining of Leaderboards for Empirical AI Research 31 Aug 2021 · 1 repository · arXiv:2109.13089
-
Evaluating the Robustness of Neural Language Models to Input Perturbations 27 Aug 2021 · 1 repository · arXiv:2108.12237Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Sentence-T5: Scalable Sentence Encoders from Pre-trained Text-to-Text Models 19 Aug 2021 · 2 repositories · arXiv:2108.08877
-
Table Caption Generation in Scholarly Documents Leveraging Pre-trained Language Models 18 Aug 2021 · 1 repository · arXiv:2108.08111
-
EncT5: Fine-tuning T5 Encoder for Discriminative Tasks 17 Aug 2021 · 0 repositories
-
Misleading the Covid-19 vaccination discourse on Twitter: An exploratory study of infodemic around the pandemic 16 Aug 2021 · 1 repository · arXiv:2108.10735
-
SAPPHIRE: Approaches for Enhanced Concept-to-Text Generation 15 Aug 2021 · 0 repositories · arXiv:2108.06643
-
How Optimal is Greedy Decoding for Extractive Question Answering? 12 Aug 2021 · 1 repository · arXiv:2108.05857Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples)
-
Variable-Length Music Score Infilling via XLNet and Musically Specialized Positional Encoding 11 Aug 2021 · 1 repository · arXiv:2108.05064
-
End-to-End User Behavior Retrieval in Click-Through RatePrediction Model 10 Aug 2021 · 1 repository · arXiv:2108.04468
-
Exploring Listwise Evidence Reasoning with T5 for Fact Verification 1 Aug 2021 · 0 repositories
-
nmT5 - Is parallel data still relevant for pre-training massively multilingual language models? 1 Aug 2021 · 0 repositories
-
YoungSheldon at SemEval-2021 Task 5: Fine-tuning Pre-trained Language Models for Toxic Spans Detection using Token classification Objective 1 Aug 2021 · 1 repository
-
EmailSum: Abstractive Email Thread Summarization 30 Jul 2021 · 1 repository · arXiv:2107.14691
-
ReFormer: The Relational Transformer for Image Captioning 29 Jul 2021 · 1 repository · arXiv:2107.14178
-
Clinical Relation Extraction Using Transformer-based Models 19 Jul 2021 · 1 repository · arXiv:2107.08957
-
Turning Tables: Generating Examples from Semi-structured Tables for Endowing Language Models with Reasoning Skills 15 Jul 2021 · 1 repository · arXiv:2107.07261
-
Transformers with multi-modal features and post-fusion context for e-commerce session-based recommendation 11 Jul 2021 · 0 repositories · arXiv:2107.05124
-
ERNIE 3.0: Large-scale Knowledge Enhanced Pre-training for Language Understanding and Generation 5 Jul 2021 · 2 repositories · arXiv:2107.02137
-
What Helps Transformers Recognize Conversational Structure? Importance of Context, Punctuation, and Labels in Dialog Act Recognition 5 Jul 2021 · 1 repository · arXiv:2107.02294
-
Improving Factual Consistency of Abstractive Summarization on Customer Feedback 30 Jun 2021 · 0 repositories · arXiv:2106.16188
-
XL-Sum: Large-Scale Multilingual Abstractive Summarization for 44 Languages 25 Jun 2021 · 2 repositories · arXiv:2106.13822Syntology official: no sample here; runs from other or unrecorded repositories · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Winner Team Mia at TextVQA Challenge 2021: Vision-and-Language Representation Learning with Pre-trained Sequence-to-Sequence Model 24 Jun 2021 · 0 repositories · arXiv:2106.15332
-
CPM-2: Large-scale Cost-effective Pre-trained Language Models 20 Jun 2021 · 2 repositories · arXiv:2106.10715
-
JointGT: Graph-Text Joint Representation Learning for Text Generation from Knowledge Graphs 19 Jun 2021 · 1 repository · arXiv:2106.10502
-
Incorporating Word Sense Disambiguation in Neural Language Models 15 Jun 2021 · 2 repositories · arXiv:2106.07967
-
Improving Paraphrase Detection with the Adversarial Paraphrasing Task 14 Jun 2021 · 1 repository · arXiv:2106.07691
-
Memory-efficient Transformers via Top-k Attention 13 Jun 2021 · 2 repositories · arXiv:2106.06899
-
Target Model Agnostic Adversarial Attacks with Query Budgets on Language Understanding Models 13 Jun 2021 · 0 repositories · arXiv:2106.07047
-
Neural Combinatory Constituency Parsing 12 Jun 2021 · 1 repository · arXiv:2106.06689
-
FastSeq: Make Sequence Generation Faster 8 Jun 2021 · 1 repository · arXiv:2106.04718
-
TIMEDIAL: Temporal Commonsense Reasoning in Dialog 8 Jun 2021 · 1 repository · arXiv:2106.04571
-
nmT5 -- Is parallel data still relevant for pre-training massively multilingual language models? 3 Jun 2021 · 0 repositories · arXiv:2106.02171
-
Evidence-based Factual Error Correction 2 Jun 2021 · 1 repository · arXiv:2106.01072
-
Contextualized and Generalized Sentence Representations by Contrastive Self-Supervised Learning: A Case Study on Discourse Relation Analysis 1 Jun 2021 · 0 repositories
-
Implicit Representations of Meaning in Neural Language Models 1 Jun 2021 · 1 repository · arXiv:2106.00737
-
Multi-Grained Knowledge Distillation for Named Entity Recognition 1 Jun 2021 · 0 repositories
-
Towards a Comprehensive Understanding and Accurate Evaluation of Societal Biases in Pre-Trained Transformers 1 Jun 2021 · 0 repositories
-
How transfer learning impacts linguistic knowledge in deep NLP models? 31 May 2021 · 0 repositories · arXiv:2105.15179
-
ByT5: Towards a token-free future with pre-trained byte-to-byte models 28 May 2021 · 5 repositories · arXiv:2105.13626Syntology official (archive's flag): 4 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
SciFive: a text-to-text transformer model for biomedical literature 28 May 2021 · 1 repository · arXiv:2106.03598Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Killing One Bird with Two Stones: Model Extraction and Attribute Inference Attacks against BERT-based APIs 23 May 2021 · 0 repositories · arXiv:2105.10909
-
Attention as Inference via Fenchel Duality 21 May 2021 · 0 repositories
-
Explainable Tsetlin Machine framework for fake news detection with credibility score assessment 19 May 2021 · 6 repositories · arXiv:2105.09114
-
Exploring Text-to-Text Transformers for English to Hinglish Machine Translation with Synthetic Code-Mixing 18 May 2021 · 0 repositories · arXiv:2105.08807
-
Stage-wise Fine-tuning for Graph-to-Text Generation 17 May 2021 · 1 repository · arXiv:2105.08021
-
BERT Busters: Outlier Dimensions that Disrupt Transformers 14 May 2021 · 0 repositories · arXiv:2105.06990
-
Which transformer architecture fits my data? A vocabulary bottleneck in self-attention 9 May 2021 · 0 repositories · arXiv:2105.03928
-
MT6: Multilingual Pretrained Text-to-Text Transformer with Translation Pairs 18 Apr 2021 · 2 repositories · arXiv:2104.08692
-
The Power of Scale for Parameter-Efficient Prompt Tuning 18 Apr 2021 · 12 repositories · arXiv:2104.08691Syntology community repositories only · 13 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 15 harvested samples)
-
Decrypting Cryptic Crosswords: Semantically Complex Wordplay Puzzles as a Target for NLP 17 Apr 2021 · 2 repositories · arXiv:2104.08620
-
A Survey of Recent Abstract Summarization Techniques 15 Apr 2021 · 0 repositories · arXiv:2105.00824
-
ExplaGraphs: An Explanation Graph Generation Task for Structured Commonsense Reasoning 15 Apr 2021 · 1 repository · arXiv:2104.07644Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples)
-
NT5?! Training T5 to Perform Numerical Reasoning 15 Apr 2021 · 1 repository · arXiv:2104.07307
-
NAREOR: The Narrative Reordering Problem 14 Apr 2021 · 1 repository · arXiv:2104.06669
-
KI-BERT: Infusing Knowledge Context for Better Language and Domain Understanding 9 Apr 2021 · 0 repositories · arXiv:2104.08145
-
Transformers: "The End of History" for NLP? 9 Apr 2021 · 0 repositories · arXiv:2105.00813
-
CodeTrans: Towards Cracking the Language of Silicon's Code Through Self-Supervised Deep Learning and High Performance Computing 6 Apr 2021 · 1 repository · arXiv:2104.02443
-
Exploring Transformers in Emotion Recognition: a comparison of BERT, DistillBERT, RoBERTa, XLNet and ELECTRA 5 Apr 2021 · 0 repositories · arXiv:2104.02041
-
Russian Paraphrasers: Paraphrase with Transformers 1 Apr 2021 · 2 repositories
-
Through the Looking Glass: Learning to Attribute Synthetic Text Generated by Language Models 1 Apr 2021 · 0 repositories
-
Automatic Graph Partitioning for Very Large-scale Deep Learning 30 Mar 2021 · 0 repositories · arXiv:2103.16063
-
A Practical Survey on Faster and Lighter Transformers 26 Mar 2021 · 0 repositories · arXiv:2103.14636
-
K-XLNet: A General Method for Combining Explicit Knowledge with Language Model Pretraining 25 Mar 2021 · 0 repositories · arXiv:2104.10649
-
GLM: General Language Model Pretraining with Autoregressive Blank Infilling 18 Mar 2021 · 8 repositories · arXiv:2103.10360Syntology community repositories only · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Comparing the Performance of NLP Toolkits and Evaluation measures in Legal Tech 12 Mar 2021 · 0 repositories · arXiv:2103.11792
-
Orthogonal Attention: A Cloze-Style Approach to Negation Scope Resolution 7 Mar 2021 · 0 repositories · arXiv:2103.04294
-
Syntax-BERT: Improving Pre-trained Transformers with Syntax Trees 7 Mar 2021 · 1 repository · arXiv:2103.04350
-
NLP-CUET@DravidianLangTech-EACL2021: Investigating Visual and Textual Features to Identify Trolls from Multimodal Social Media Memes 28 Feb 2021 · 0 repositories · arXiv:2103.00466
-
NLP-CUET@LT-EDI-EACL2021: Multilingual Code-Mixed Hope Speech Detection using Cross-lingual Representation Learner 28 Feb 2021 · 1 repository · arXiv:2103.00464
-
Multi-task transfer learning for finding actionable information from crisis-related messages on social media 26 Feb 2021 · 0 repositories · arXiv:2102.13395
-
PADA: Example-based Prompt Learning for on-the-fly Adaptation to Unseen Domains 24 Feb 2021 · 1 repository · arXiv:2102.12206
-
Few Shot Learning for Information Verification 22 Feb 2021 · 0 repositories · arXiv:2102.10956
-
Quiz-Style Question Generation for News Stories 18 Feb 2021 · 2 repositories · arXiv:2102.09094
-
Exploring Transformers in Natural Language Generation: GPT, BERT, and XLNet 16 Feb 2021 · 1 repository · arXiv:2102.08036
-
WangchanBERTa: Pretraining transformer-based Thai Language Models 24 Jan 2021 · 2 repositories · arXiv:2101.09635
-
Transformer-Based Models for Question Answering on COVID19 16 Jan 2021 · 0 repositories · arXiv:2101.11432
-
Fake News Detection System using XLNet model with Topic Distributions: CONSTRAINT@AAAI2021 Shared Task 12 Jan 2021 · 0 repositories · arXiv:2101.11425
-
Cisco at AAAI-CAD21 shared task: Predicting Emphasis in Presentation Slides using Contextualized Embeddings 10 Jan 2021 · 1 repository · arXiv:2101.11422
-
Transformer-based approach towards music emotion recognition from lyrics 6 Jan 2021 · 1 repository · arXiv:2101.02051
-
Improving Sequence-to-Sequence Pre-training via Sequence Span Rewriting 2 Jan 2021 · 1 repository · arXiv:2101.00416
-
Pre-training Text-to-Text Transformers to Write and Reason with Concepts 1 Jan 2021 · 0 repositories
-
Syntactic Relevance XLNet Word Embedding Generation in Low-Resource Machine Translation 1 Jan 2021 · 0 repositories
-
EarlyBERT: Efficient BERT Training via Early-bird Lottery Tickets 31 Dec 2020 · 1 repository · arXiv:2101.00063Syntology official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Evidence-based Factual Error Correction 31 Dec 2020 · 3 repositories · arXiv:2012.15788
-
Leveraging ParsBERT and Pretrained mT5 for Persian Abstractive Text Summarization 21 Dec 2020 · 1 repository · arXiv:2012.11204
-
DialogXL: All-in-One XLNet for Multi-Party Conversation Emotion Recognition 16 Dec 2020 · 4 repositories · arXiv:2012.08695
-
Discriminative Pre-training for Low Resource Title Compression in Conversational Grocery 13 Dec 2020 · 0 repositories · arXiv:2012.06943
-
Yelp Review Rating Prediction: Machine Learning and Deep Learning Models 12 Dec 2020 · 1 repository · arXiv:2012.06690
-
How Can We Know When Language Models Know? On the Calibration of Language Models for Question Answering 2 Dec 2020 · 1 repository · arXiv:2012.00955
-
BERT at SemEval-2020 Task 8: Using BERT to Analyse Meme Emotions 1 Dec 2020 · 0 repositories
-
Comparing Probabilistic, Distributional and Transformer-Based Models on Logical Metonymy Interpretation 1 Dec 2020 · 0 repositories
-
Evaluating Unsupervised Representation Learning for Detecting Stances of Fake News 1 Dec 2020 · 0 repositories
-
Flight of the PEGASUS? Comparing Transformers on Few-shot and Zero-shot Multi-document Abstractive Summarization 1 Dec 2020 · 1 repository
-
Hitachi at SemEval-2020 Task 11: An Empirical Study of Pre-Trained Transformer Family for Propaganda Detection 1 Dec 2020 · 0 repositories
-
Hy-NLI: a Hybrid system for Natural Language Inference 1 Dec 2020 · 2 repositories
-
Label Representations in Modeling Classification as Text Generation 1 Dec 2020 · 0 repositories
-
Neural language models for text classification in evidence-based medicine 1 Dec 2020 · 0 repositories · arXiv:2012.00584
-
NLP@JUST at SemEval-2020 Task 4: Ensemble Technique for BERT and Roberta to Evaluate Commonsense Validation 1 Dec 2020 · 0 repositories
-
SINAI at SemEval-2020 Task 12: Offensive Language Identification Exploring Transfer Learning Models 1 Dec 2020 · 0 repositories