Methods › Natural Language Processing › Subword Segmentation › WordPiece › Papers, page 50
WordPiece
Papers archive 2025-07-28
archive papers tagged: 7,063 · with a code link: 2,910 · where Syntology ran a sample: 650 (529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (650 of 7,063 tagged: 529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument)
Page 50 of 71: papers 4,901 to 5,000 of 7,063, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Editing Factual Knowledge in Language Models 16 Apr 2021 · 3 repositories · arXiv:2104.08164Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
Fast, Effective, and Self-Supervised: Transforming Masked Language Models into Universal Lexical and Sentence Encoders 16 Apr 2021 · 1 repository · arXiv:2104.08027Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Membership Inference Attack Susceptibility of Clinical Language Models 16 Apr 2021 · 0 repositories · arXiv:2104.08305
-
Probing Across Time: What Does RoBERTa Know and When? 16 Apr 2021 · 1 repository · arXiv:2104.07885Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Temporal Adaptation of BERT and Performance on Downstream Document Classification: Insights from Social Media 16 Apr 2021 · 2 repositories · arXiv:2104.08116
-
Towards Variable-Length Textual Adversarial Attacks 16 Apr 2021 · 0 repositories · arXiv:2104.08139
-
A Sample-Based Training Method for Distantly Supervised Relation Extraction with Pre-Trained Transformers 15 Apr 2021 · 0 repositories · arXiv:2104.07512
-
Are Multilingual BERT models robust? A Case Study on Adversarial Attacks for Multilingual Question Answering 15 Apr 2021 · 0 repositories · arXiv:2104.07646
-
BERT based Transformers lead the way in Extraction of Health Information from Social Media 15 Apr 2021 · 1 repository · arXiv:2104.07367
-
Does BERT Pretrained on Clinical Notes Reveal Sensitive Data? 15 Apr 2021 · 4 repositories · arXiv:2104.07762Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Emotion Dynamics Modeling via BERT 15 Apr 2021 · 0 repositories · arXiv:2104.07252
-
How to Train BERT with an Academic Budget 15 Apr 2021 · 4 repositories · arXiv:2104.07705
-
Natural Language Understanding with Privacy-Preserving BERT 15 Apr 2021 · 0 repositories · arXiv:2104.07504
-
SINA-BERT: A pre-trained Language Model for Analysis of Medical Texts in Persian 15 Apr 2021 · 0 repositories · arXiv:2104.07613
-
Text Guide: Improving the quality of long text classification by a text selection method based on feature importance 15 Apr 2021 · 1 repository · arXiv:2104.07225
-
TorontoCL at CMCL 2021 Shared Task: RoBERTa with Multi-Stage Fine-Tuning for Eye-Tracking Prediction 15 Apr 2021 · 1 repository · arXiv:2104.07244
-
Ultra-High Dimensional Sparse Representations with Binarization for Efficient Text Retrieval 15 Apr 2021 · 0 repositories · arXiv:2104.07198
-
UIT-E10dot3 at SemEval-2021 Task 5: Toxic Spans Detection with Named Entity Recognition and Question-Answering Approaches 15 Apr 2021 · 0 repositories · arXiv:2104.07376
-
An Interpretability Illusion for BERT 14 Apr 2021 · 0 repositories · arXiv:2104.07143
-
Demystifying BERT: Implications for Accelerator Design 14 Apr 2021 · 0 repositories · arXiv:2104.08335
-
Enhancing Word-Level Semantic Representation via Dependency Structure for Expressive Text-to-Speech Synthesis 14 Apr 2021 · 0 repositories · arXiv:2104.06835
-
Disentangling Representations of Text by Masking Transformers 14 Apr 2021 · 0 repositories · arXiv:2104.07155
-
On the Robustness of Intent Classification and Slot Labeling in Goal-oriented Dialog Systems to Real-world Noise 14 Apr 2021 · 1 repository · arXiv:2104.07149
-
Static Embeddings as Efficient Knowledge Bases? 14 Apr 2021 · 1 repository · arXiv:2104.07094
-
1-bit LAMB: Communication Efficient Large-Scale Large-Batch Training with LAMB's Convergence Speed 13 Apr 2021 · 1 repository · arXiv:2104.06069
-
Discourse Probing of Pretrained Language Models 13 Apr 2021 · 1 repository · arXiv:2104.05882
-
Large-Scale Contextualised Language Modelling for Norwegian 13 Apr 2021 · 2 repositories · arXiv:2104.06546Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples)
-
Mediators in Determining what Processing BERT Performs First 13 Apr 2021 · 1 repository · arXiv:2104.06400
-
Semantic maps and metrics for science Semantic maps and metrics for science using deep transformer encoders 13 Apr 2021 · 0 repositories · arXiv:2104.05928
-
Understanding Transformers for Bot Detection in Twitter 13 Apr 2021 · 1 repository · arXiv:2104.06182
-
BERT based freedom to operate patent analysis 12 Apr 2021 · 0 repositories · arXiv:2105.00817
-
Fighting the COVID-19 Infodemic with a Holistic BERT Ensemble 12 Apr 2021 · 1 repository · arXiv:2104.05745
-
Fine-Tuning Transformers for Identifying Self-Reporting Potential Cases and Symptoms of COVID-19 in Tweets 12 Apr 2021 · 1 repository · arXiv:2104.05501
-
Learning to Remove: Towards Isotropic Pre-trained BERT Embedding 12 Apr 2021 · 1 repository · arXiv:2104.05274
-
Multilingual Language Models Predict Human Reading Behavior 12 Apr 2021 · 1 repository · arXiv:2104.05433
-
WHOSe Heritage: Classification of UNESCO World Heritage "Outstanding Universal Value" Documents with Soft Labels 12 Apr 2021 · 1 repository · arXiv:2104.05547
-
Does syntax matter? A strong baseline for Aspect-based Sentiment Analysis with RoBERTa 11 Apr 2021 · 1 repository · arXiv:2104.04986
-
Fine-tuning Encoders for Improved Monolingual and Zero-shot Polylingual Neural Topic Modeling 11 Apr 2021 · 1 repository · arXiv:2104.05064
-
Innovative Bert-based Reranking Language Models for Speech Recognition 11 Apr 2021 · 0 repositories · arXiv:2104.04950
-
UniDrop: A Simple yet Effective Technique to Improve Transformer without Extra Cost 11 Apr 2021 · 0 repositories · arXiv:2104.04946
-
MIPT-NSU-UTMN at SemEval-2021 Task 5: Ensembling Learning with Pre-trained Language Models for Toxic Spans Detection 10 Apr 2021 · 1 repository · arXiv:2104.04739
-
Non-autoregressive Transformer-based End-to-end ASR using BERT 10 Apr 2021 · 0 repositories · arXiv:2104.04805
-
ZS-BERT: Towards Zero-Shot Relation Extraction with Attribute Representation Learning 10 Apr 2021 · 1 repository · arXiv:2104.04697
-
KI-BERT: Infusing Knowledge Context for Better Language and Domain Understanding 9 Apr 2021 · 0 repositories · arXiv:2104.08145
-
The Road to Know-Where: An Object-and-Room Informed Sequential BERT for Indoor Vision-Language Navigation 9 Apr 2021 · 1 repository · arXiv:2104.04167
-
Text2Chart: A Multi-Staged Chart Generator from Natural Language Text 9 Apr 2021 · 1 repository · arXiv:2104.04584
-
Transformers: "The End of History" for NLP? 9 Apr 2021 · 0 repositories · arXiv:2105.00813
-
Layer Reduction: Accelerating Conformer-Based Self-Supervised Model via Layer Consistency 8 Apr 2021 · 0 repositories · arXiv:2105.00812
-
Lone Pine at SemEval-2021 Task 5: Fine-Grained Detection of Hate Speech Using BERToxic 8 Apr 2021 · 1 repository · arXiv:2104.03506
-
Probing BERT in Hyperbolic Spaces 8 Apr 2021 · 1 repository · arXiv:2104.03869Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified; the one sample that ran constructed an object rather than computing a result (of 3 harvested samples) · 3 pointer-only (licence)
-
Uppsala NLP at SemEval-2021 Task 2: Multilingual Language Models for Fine-tuning and Feature Extraction in Word-in-Context Disambiguation 8 Apr 2021 · 0 repositories · arXiv:2104.03767
-
Better Neural Machine Translation by Extracting Linguistic Information from BERT 7 Apr 2021 · 1 repository · arXiv:2104.02831
-
Combining Pre-trained Word Embeddings and Linguistic Features for Sequential Metaphor Identification 7 Apr 2021 · 0 repositories · arXiv:2104.03285
-
Interpreting Verbal Metaphors by Paraphrasing 7 Apr 2021 · 0 repositories · arXiv:2104.03391
-
Speak or Chat with Me: End-to-End Spoken Language Understanding System with Flexible Inputs 7 Apr 2021 · 1 repository · arXiv:2104.05752
-
An Empirical Evaluation of Word Embedding Models for Subjectivity Analysis Tasks 6 Apr 2021 · 1 repository
-
Efficient transfer learning for NLP with ELECTRA 6 Apr 2021 · 1 repository · arXiv:2104.02756
-
HBert + BiasCorp -- Fighting Racism on the Web 6 Apr 2021 · 0 repositories · arXiv:2104.02242
-
MuSLCAT: Multi-Scale Multi-Level Convolutional Attention Transformer for Discriminative Music Modeling on Raw Waveforms 6 Apr 2021 · 0 repositories · arXiv:2104.02309
-
COVID-19 sentiment analysis via deep learning during the rise of novel cases 5 Apr 2021 · 0 repositories · arXiv:2104.10662
-
Exploring Transformers in Emotion Recognition: a comparison of BERT, DistillBERT, RoBERTa, XLNet and ELECTRA 5 Apr 2021 · 0 repositories · arXiv:2104.02041
-
Semantic Distance: A New Metric for ASR Performance Analysis Towards Spoken Language Understanding 5 Apr 2021 · 0 repositories · arXiv:2104.02138
-
What's the best place for an AI conference, Vancouver or ______: Why completing comparative questions is difficult 5 Apr 2021 · 0 repositories · arXiv:2104.01940
-
Improving Pretrained Models for Zero-shot Multi-label Text Classification through Reinforced Label Hierarchy Reasoning 4 Apr 2021 · 1 repository · arXiv:2104.01666
-
MCL@IITK at SemEval-2021 Task 2: Multilingual and Cross-lingual Word-in-Context Disambiguation using Augmented Data, Signals, and Transformers 4 Apr 2021 · 0 repositories · arXiv:2104.01567
-
ReCAM@IITK at SemEval-2021 Task 4: BERT and ALBERT based Ensemble for Abstract Word Prediction 4 Apr 2021 · 1 repository · arXiv:2104.01563
-
Exploring the Role of BERT Token Representations to Explain Sentence Probing Results 3 Apr 2021 · 1 repository · arXiv:2104.01477
-
Unsupervised Domain Adaptation with Global and Local Graph Neural Networks in Limited Labeled Data Scenario: Application to Disaster Management 3 Apr 2021 · 0 repositories · arXiv:2104.01436
-
IITK@LCP at SemEval 2021 Task 1: Classification for Lexical Complexity Regression Task 2 Apr 2021 · 1 repository · arXiv:2104.01046
-
The Coronavirus is a Bioweapon: Analysing Coronavirus Fact-Checked Stories 2 Apr 2021 · 0 repositories · arXiv:2104.01215
-
A Dashboard for Mitigating the COVID-19 Misinfodemic 1 Apr 2021 · 0 repositories
-
Are Neural Networks Extracting Linguistic Properties or Memorizing Training Data? An Observation with a Multilingual Probe for Predicting Tense 1 Apr 2021 · 1 repository
-
BERT meets Cranfield: Uncovering the Properties of Full Ranking on Fully Labeled Data 1 Apr 2021 · 0 repositories
-
BERT Prescriptions to Avoid Unwanted Headaches: A Comparison of Transformer Architectures for Adverse Drug Event Detection 1 Apr 2021 · 1 repository
-
BERTective: Language Models and Contextual Information for Deception Detection 1 Apr 2021 · 0 repositories
-
BERxiT: Early Exiting for BERT with Better Fine-Tuning and Extension to Regression 1 Apr 2021 · 1 repository
-
Complex Question Answering on knowledge graphs using machine translation and multi-task learning 1 Apr 2021 · 0 repositories
-
Content-based Models of Quotation 1 Apr 2021 · 0 repositories
-
Detecting Scenes in Fiction: A new Segmentation Task 1 Apr 2021 · 0 repositories
-
ENPAR:Enhancing Entity and Entity Pair Representations for Joint Entity Relation Extraction 1 Apr 2021 · 1 repository
-
Evaluating language models for the retrieval and categorization of lexical collocations 1 Apr 2021 · 1 repository
-
Evaluating Neural Model Robustness for Machine Comprehension 1 Apr 2021 · 0 repositories
-
HLE-UPC at SemEval-2021 Task 5: Multi-Depth DistilBERT for Toxic Spans Detection 1 Apr 2021 · 1 repository · arXiv:2104.00639
-
How Fast can BERT Learn Simple Natural Language Inference? 1 Apr 2021 · 0 repositories
-
Keep Learning: Self-supervised Meta-learning for Learning from Inference 1 Apr 2021 · 0 repositories
-
Maximal Multiverse Learning for Promoting Cross-Task Generalization of Fine-Tuned Language Models 1 Apr 2021 · 0 repositories
-
Multilingual Entity and Relation Extraction Dataset and Model 1 Apr 2021 · 1 repository
-
Neural-Driven Search-Based Paraphrase Generation 1 Apr 2021 · 0 repositories
-
NLQuAD: A Non-Factoid Long Question Answering Data Set 1 Apr 2021 · 1 repository
-
On the (In)Effectiveness of Images for Text Classification 1 Apr 2021 · 0 repositories
-
Probing for idiomaticity in vector space models 1 Apr 2021 · 1 repository
-
Retrieval, Re-ranking and Multi-task Learning for Knowledge-Base Question Answering 1 Apr 2021 · 0 repositories
-
An In-depth Analysis of Passage-Level Label Transfer for Contextual Document Ranking 30 Mar 2021 · 1 repository · arXiv:2103.16669
-
Automatic Graph Partitioning for Very Large-scale Deep Learning 30 Mar 2021 · 0 repositories · arXiv:2103.16063
-
Grounding Dialogue Systems via Knowledge Graph Aware Decoding with Pre-trained Transformers 30 Mar 2021 · 1 repository · arXiv:2103.16289
-
Kaleido-BERT: Vision-Language Pre-training on Fashion Domain 30 Mar 2021 · 1 repository · arXiv:2103.16110Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 2 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples)
-
Multi-Scale Vision Longformer: A New Vision Transformer for High-Resolution Image Encoding 29 Mar 2021 · 3 repositories · arXiv:2103.15358Syntology official (archive's flag): 3 ran · 8 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 2 violated, 2 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples) · 2 pointer-only (licence)
-
Contextual Text Embeddings for Twi 29 Mar 2021 · 0 repositories · arXiv:2103.15963
-
Retraining DistilBERT for a Voice Shopping Assistant by Using Universal Dependencies 29 Mar 2021 · 0 repositories · arXiv:2103.15737
-
Whitening Sentence Representations for Better Semantics and Faster Retrieval 29 Mar 2021 · 3 repositories · arXiv:2103.15316Syntology community repositories only · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 5 harvested samples)