Methods › Natural Language Processing › Subword Segmentation › WordPiece › Papers, page 58
WordPiece
Papers archive 2025-07-28
archive papers tagged: 7,063 · with a code link: 2,910 · where Syntology ran a sample: 650 (529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (650 of 7,063 tagged: 529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument)
Page 58 of 71: papers 5,701 to 5,800 of 7,063, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Combining Deep Learning and String Kernels for the Localization of Swiss German Tweets 7 Oct 2020 · 0 repositories · arXiv:2010.03614
-
Detecting Fine-Grained Cross-Lingual Semantic Divergences without Supervision by Learning to Rank 7 Oct 2020 · 1 repository · arXiv:2010.03662
-
DiPair: Fast and Accurate Distillation for Trillion-Scale Text Matching and Pair Modeling 7 Oct 2020 · 0 repositories · arXiv:2010.03099
-
ELMo and BERT in semantic change detection for Russian 7 Oct 2020 · 0 repositories · arXiv:2010.03481
-
AxFormer: Accuracy-driven Approximation of Transformers for Faster, Smaller and more Accurate NLP Models 7 Oct 2020 · 1 repository · arXiv:2010.03688
-
Where Are the Facts? Searching for Fact-checked Information to Alleviate the Spread of Fake News 7 Oct 2020 · 2 repositories · arXiv:2010.03159Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Why do you think that? Exploring Faithful Sentence-Level Rationales Without Supervision 7 Oct 2020 · 1 repository · arXiv:2010.03384
-
Analyzing Individual Neurons in Pre-trained Language Models 6 Oct 2020 · 1 repository · arXiv:2010.02695
-
BERT Knows Punta Cana is not just beautiful, it's gorgeous: Ranking Scalar Adjectives with Contextualised Representations 6 Oct 2020 · 1 repository · arXiv:2010.02686
-
Cross-Lingual Text Classification with Minimal Resources by Transferring a Sparse Teacher 6 Oct 2020 · 1 repository · arXiv:2010.02562
-
Do Explicit Alignments Robustly Improve Multilingual Encoders? 6 Oct 2020 · 1 repository · arXiv:2010.02537
-
Exploring BERT's Sensitivity to Lexical Cues using Tests from Semantic Priming 6 Oct 2020 · 1 repository · arXiv:2010.03010Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Improving Efficient Neural Ranking Models with Cross-Architecture Knowledge Distillation 6 Oct 2020 · 1 repository · arXiv:2010.02666
-
Incorporating Behavioral Hypotheses for Query Generation 6 Oct 2020 · 0 repositories · arXiv:2010.02667
-
Intrinsic Probing through Dimension Selection 6 Oct 2020 · 1 repository · arXiv:2010.02812Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
LEGAL-BERT: The Muppets straight out of Law School 6 Oct 2020 · 0 repositories · arXiv:2010.02559
-
Neural Mask Generator: Learning to Generate Adaptive Word Maskings for Language Model Adaptation 6 Oct 2020 · 1 repository · arXiv:2010.02705
-
On the Interplay Between Fine-tuning and Sentence-level Probing for Linguistic Knowledge in Pre-trained Transformers 6 Oct 2020 · 0 repositories · arXiv:2010.02616
-
Poison Attacks against Text Datasets with Conditional Adversarially Regularized Autoencoder 6 Oct 2020 · 2 repositories · arXiv:2010.02684Syntology official (archive's flag): 2 ran · 2 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Pretrained Language Model Embryology: The Birth of ALBERT 6 Oct 2020 · 1 repository · arXiv:2010.02480
-
The Multilingual Amazon Reviews Corpus 6 Oct 2020 · 1 repository · arXiv:2010.02573
-
How Effective is Task-Agnostic Data Augmentation for Pretrained Transformers? 5 Oct 2020 · 0 repositories · arXiv:2010.01764
-
Improving AMR Parsing with Sequence-to-Sequence Pre-training 5 Oct 2020 · 1 repository · arXiv:2010.01771
-
InfoBERT: Improving Robustness of Language Models from An Information Theoretic Perspective 5 Oct 2020 · 2 repositories · arXiv:2010.02329Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 2 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Linguistic Profiling of a Neural Language Model 5 Oct 2020 · 0 repositories · arXiv:2010.01869
-
Mixup-Transformer: Dynamic Data Augmentation for NLP Tasks 5 Oct 2020 · 0 repositories · arXiv:2010.02394
-
PAIR: Planning and Iterative Refinement in Pre-trained Transformers for Long Text Generation 5 Oct 2020 · 0 repositories · arXiv:2010.02301
-
Pareto Probing: Trading Off Accuracy for Complexity 5 Oct 2020 · 1 repository · arXiv:2010.02180Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
PMI-Masking: Principled masking of correlated spans 5 Oct 2020 · 1 repository · arXiv:2010.01825
-
Pruning Redundant Mappings in Transformer Models via Spectral-Normalized Identity Prior 5 Oct 2020 · 1 repository · arXiv:2010.01791
-
PUM at SemEval-2020 Task 12: Aggregation of Transformer-based models' features for offensive language recognition 5 Oct 2020 · 0 repositories · arXiv:2010.01897
-
Self-training Improves Pre-training for Natural Language Understanding 5 Oct 2020 · 1 repository · arXiv:2010.02194
-
Unsupervised Reference-Free Summary Quality Evaluation via Contrastive Learning 5 Oct 2020 · 1 repository · arXiv:2010.01781
-
X-SRL: A Parallel Cross-Lingual Semantic Role Labeling Dataset 5 Oct 2020 · 1 repository · arXiv:2010.01998
-
An Empirical Study on Large-Scale Multi-Label Text Classification Including Few and Zero-Shot Labels 4 Oct 2020 · 1 repository · arXiv:2010.01653Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 1 pointer-only (licence)
-
On Losses for Modern Language Models 4 Oct 2020 · 1 repository · arXiv:2010.01694
-
Mining Knowledge for Natural Language Inference from Wikipedia Categories 3 Oct 2020 · 1 repository · arXiv:2010.01239Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Personality Trait Detection Using Bagged SVM over BERT Word Embedding Ensembles 3 Oct 2020 · 0 repositories · arXiv:2010.01309
-
Cost-effective Selection of Pretraining Data: A Case Study of Pretraining BERT on Social Media 2 Oct 2020 · 0 repositories · arXiv:2010.01150
-
Data-Efficient Pretraining via Contrastive Self-Supervision 2 Oct 2020 · 0 repositories · arXiv:2010.01061
-
LUKE: Deep Contextualized Entity Representations with Entity-aware Self-attention 2 Oct 2020 · 9 repositories · arXiv:2010.01057Syntology official (archive's flag): 3 ran · 3 ran (of which 2 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 7 unverified (of 10 harvested samples)
-
MultiCQA: Zero-Shot Transfer of Self-Supervised Text Matching Models on a Massive Scale 2 Oct 2020 · 1 repository · arXiv:2010.00980
-
STIL -- Simultaneous Slot Filling, Translation, Intent Classification, and Language Identification: Initial Results using mBART on MultiATIS++ 2 Oct 2020 · 1 repository · arXiv:2010.00760
-
A Technical Question Answering System with Transfer Learning 1 Oct 2020 · 1 repository
-
Beyond The Text: Analysis of Privacy Statements through Syntactic and Semantic Role Labeling 1 Oct 2020 · 0 repositories · arXiv:2010.00678
-
CoLAKE: Contextualized Language and Knowledge Embedding 1 Oct 2020 · 1 repository · arXiv:2010.00309Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
Detecting White Supremacist Hate Speech using Domain Specific Word Embedding with Deep Learning and BERT 1 Oct 2020 · 0 repositories · arXiv:2010.00357
-
Evaluating Multilingual BERT for Estonian 1 Oct 2020 · 0 repositories · arXiv:2010.00454
-
RefVOS: A Closer Look at Referring Expressions for Video Object Segmentation 1 Oct 2020 · 2 repositories · arXiv:2010.00263
-
RRF102: Meeting the TREC-COVID Challenge with a 100+ Runs Ensemble 1 Oct 2020 · 0 repositories · arXiv:2010.00200
-
Understanding tables with intermediate pre-training 1 Oct 2020 · 1 repository · arXiv:2010.00571
-
A Tale of Two Linkings: Dynamically Gating between Schema Linking and Structural Linking for Text-to-SQL Parsing 30 Sep 2020 · 1 repository · arXiv:2009.14809
-
A Vietnamese Dataset for Evaluating Machine Reading Comprehension 30 Sep 2020 · 0 repositories · arXiv:2009.14725
-
AUBER: Automated BERT Regularization 30 Sep 2020 · 0 repositories · arXiv:2009.14409
-
BERT for Monolingual and Cross-Lingual Reverse Dictionary 30 Sep 2020 · 1 repository · arXiv:2009.14790
-
Interpretable Machine Learning for COVID-19: An Empirical Study on Severity Prediction Task 30 Sep 2020 · 1 repository · arXiv:2010.02006
-
Pea-KD: Parameter-efficient and Accurate Knowledge Distillation on BERT 30 Sep 2020 · 0 repositories · arXiv:2009.14822
-
Contrastive Distillation on Intermediate Representations for Language Model Compression 29 Sep 2020 · 1 repository · arXiv:2009.14167Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples)
-
Cross-lingual Alignment Methods for Multilingual BERT: A Comparative Study 29 Sep 2020 · 0 repositories · arXiv:2009.14304
-
From Twitter to Traffic Predictor: Next-Day Morning Traffic Prediction Using Social Media Data 29 Sep 2020 · 0 repositories · arXiv:2009.13794
-
Gender prediction using limited Twitter Data 29 Sep 2020 · 0 repositories · arXiv:2010.02005
-
HINT3: Raising the bar for Intent Detection in the Wild 29 Sep 2020 · 1 repository · arXiv:2009.13833
-
MaP: A Matrix-based Prediction Approach to Improve Span Extraction in Machine Reading Comprehension 29 Sep 2020 · 0 repositories · arXiv:2009.14348
-
Neural Retrieval for Question Answering with Cross-Attention Supervised Data Augmentation 29 Sep 2020 · 0 repositories · arXiv:2009.13815
-
TEST_POSITIVE at W-NUT 2020 Shared Task-3: Joint Event Multi-task Learning for Slot Filling in Noisy Text 29 Sep 2020 · 0 repositories · arXiv:2009.14262
-
A Simple and Efficient Ensemble Classifier Combining Multiple Neural Network Models on Social Media Datasets in Vietnamese 28 Sep 2020 · 0 repositories · arXiv:2009.13060
-
Accelerating Multi-Model Inference by Merging DNNs of Different Weights 28 Sep 2020 · 0 repositories · arXiv:2009.13062
-
DialoGLUE: A Natural Language Understanding Benchmark for Task-Oriented Dialogue 28 Sep 2020 · 1 repository · arXiv:2009.13570
-
Fancy Man Lauches Zippo at WNUT 2020 Shared Task-1: A Bert Case Model for Wet Lab Entity Extraction 28 Sep 2020 · 0 repositories · arXiv:2009.12997
-
Knowledge-Aware Procedural Text Understanding with Multi-Stage Training 28 Sep 2020 · 0 repositories · arXiv:2009.13199
-
PIN: A Novel Parallel Interactive Network for Spoken Language Understanding 28 Sep 2020 · 0 repositories · arXiv:2009.13431
-
TernaryBERT: Distillation-aware Ultra-low Bit BERT 27 Sep 2020 · 5 repositories · arXiv:2009.12812
-
Metaphor Detection using Deep Contextualized Word Embeddings 26 Sep 2020 · 0 repositories · arXiv:2009.12565
-
Techniques to Improve Q&A Accuracy with Transformer-based models on Large Complex Documents 26 Sep 2020 · 0 repositories · arXiv:2009.12695
-
A little goes a long way: Improving toxic language classification despite data scarcity 25 Sep 2020 · 1 repository · arXiv:2009.12344
-
An Unsupervised Sentence Embedding Method by Mutual Information Maximization 25 Sep 2020 · 1 repository · arXiv:2009.12061
-
BET: A Backtranslation Approach for Easy Data Augmentation in Transformer-based Paraphrase Identification Context 25 Sep 2020 · 1 repository · arXiv:2009.12452
-
HetSeq: Distributed GPU Training on Heterogeneous Infrastructure 25 Sep 2020 · 1 repository · arXiv:2009.14783
-
A Comparative Study of Feature Types for Age-Based Text Classification 24 Sep 2020 · 1 repository · arXiv:2009.11898
-
Adapting BERT for Word Sense Disambiguation with Gloss Selection Objective and Example Sentences 24 Sep 2020 · 1 repository · arXiv:2009.11795Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
AnchiBERT: A Pre-Trained Model for Ancient ChineseLanguage Understanding and Generation 24 Sep 2020 · 1 repository · arXiv:2009.11473
-
A Token-wise CNN-based Method for Sentence Compression 23 Sep 2020 · 0 repositories · arXiv:2009.11260
-
AutoRC: Improving BERT Based Relation Classification Models via Architecture Search 22 Sep 2020 · 0 repositories · arXiv:2009.10680
-
Constructing interval variables via faceted Rasch measurement and multitask deep learning: a hate speech application 22 Sep 2020 · 2 repositories · arXiv:2009.10277Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
GRACE: Gradient Harmonized and Cascaded Labeling for Aspect-based Sentiment Analysis 22 Sep 2020 · 1 repository · arXiv:2009.10557
-
On Data Augmentation for Extreme Multi-label Classification 22 Sep 2020 · 0 repositories · arXiv:2009.10778
-
Latin BERT: A Contextual Language Model for Classical Philology 21 Sep 2020 · 1 repository · arXiv:2009.10053
-
Profile Consistency Identification for Open-domain Dialogue Agents 21 Sep 2020 · 1 repository · arXiv:2009.09680Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified; the one sample that ran constructed an object rather than computing a result (of 3 harvested samples)
-
"Listen, Understand and Translate": Triple Supervision Decouples End-to-end Speech-to-text Translation 21 Sep 2020 · 1 repository · arXiv:2009.09704
-
"When they say weed causes depression, but it's your fav antidepressant": Knowledge-aware Attention Framework for Relationship Extraction 21 Sep 2020 · 0 repositories · arXiv:2009.10155
-
Dual-path CNN with Max Gated block for Text-Based Person Re-identification 20 Sep 2020 · 1 repository · arXiv:2009.09343
-
Longformer for MS MARCO Document Re-ranking Task 20 Sep 2020 · 1 repository · arXiv:2009.09392
-
Persian Ezafe Recognition Using Transformers and Its Role in Part-Of-Speech Tagging 20 Sep 2020 · 1 repository · arXiv:2009.09474
-
Vicomtech at eHealth-KD Challenge 2020: Deep End-to-End Model for Entity and Relation Extraction in Medical Text 20 Sep 2020 · 0 repositories
-
VirtualFlow: Decoupling Deep Learning Models from the Underlying Hardware 20 Sep 2020 · 0 repositories · arXiv:2009.09523
-
BioALBERT: A Simple and Effective Pre-trained Language Model for Biomedical Named Entity Recognition 19 Sep 2020 · 0 repositories · arXiv:2009.09223
-
Conditionally Adaptive Multi-Task Learning: Improving Transfer Learning in NLP Using Fewer Parameters & Less Data 19 Sep 2020 · 1 repository · arXiv:2009.09139
-
Nominal Compound Chain Extraction: A New Task for Semantic-enriched Lexical Chain 19 Sep 2020 · 0 repositories · arXiv:2009.09173
-
Prior Art Search and Reranking for Generated Patent Text 19 Sep 2020 · 0 repositories · arXiv:2009.09132
-
fastHan: A BERT-based Multi-Task Toolkit for Chinese NLP 18 Sep 2020 · 1 repository · arXiv:2009.08633