Methods › Natural Language Processing › Subword Segmentation › WordPiece › Papers, page 41
WordPiece
Papers archive 2025-07-28
archive papers tagged: 7,063 · with a code link: 2,910 · where Syntology ran a sample: 650 (529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (650 of 7,063 tagged: 529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument)
Page 41 of 71: papers 4,001 to 4,100 of 7,063, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
BERT got a Date: Introducing Transformers to Temporal Tagging 16 Nov 2021 · 0 repositories
-
BERT is Robust! A Case Against Synonym-Based Adversarial Examples in Text Classification 16 Nov 2021 · 0 repositories
-
BigFive: A Dataset of Coarse- and Fine-Grained Personality Characteristics 16 Nov 2021 · 0 repositories
-
BORT: Back and Denoising Reconstruction for End-to-End Task-Oriented Dialog 16 Nov 2021 · 0 repositories
-
Building Chinese Biomedical Language Models via Multi-Level Text Discrimination 16 Nov 2021 · 1 repository
-
Chinese Word Attention based on Valid Division of Sentence 16 Nov 2021 · 0 repositories
-
Contextualized Sensorimotor Norms: multi-dimensional measures of sensorimotor strength for ambiguous English words, in context 16 Nov 2021 · 0 repositories
-
Cross-domain Named Entity Recognition via Graph Matching 16 Nov 2021 · 0 repositories
-
CVSS-BERT: Explainable Natural Language Processing to Determine the Severity of a Computer Security Vulnerability from its Description 16 Nov 2021 · 1 repository · arXiv:2111.08510
-
Data Contamination: From Memorization to Exploitation 16 Nov 2021 · 0 repositories
-
DAWSON: Data Augmentation using Weak Supervision On Natural Language 16 Nov 2021 · 0 repositories
-
Deep-to-bottom Weights Decay: A Systemic Knowledge Review Learning Technique for Transformer Layers in Knowledge Distillation 16 Nov 2021 · 0 repositories
-
Discontinuous Constituency and BERT: A Case Study of Dutch 16 Nov 2021 · 0 repositories
-
ElitePLM: An Empirical Study on General Language Ability Evaluation of Pretrained Language Models 16 Nov 2021 · 0 repositories
-
ELLE: Efficient Lifelong Pre-training for Emerging Data 16 Nov 2021 · 0 repositories
-
End-to-end Task-oriented Dialog Policy Learning based on Pre-trained Language Model 16 Nov 2021 · 0 repositories
-
Event Detection via Derangement Question Answering 16 Nov 2021 · 0 repositories
-
EventBERT 16 Nov 2021 · 0 repositories
-
Explicit Modeling the Context for Chinese NER 16 Nov 2021 · 0 repositories
-
Eye Gaze and Self-attention: How Humans and Transformers Attend Words in Sentences 16 Nov 2021 · 0 repositories
-
Feature-rich Open-vocabulary Interpretable Neural Representations for All of the World’s 7000 Languages 16 Nov 2021 · 0 repositories
-
Feature Structure Distillation for BERT Transferring 16 Nov 2021 · 0 repositories
-
gaBERT — an Irish Language Model 16 Nov 2021 · 0 repositories
-
Get the Point! Graph Enhanced Candidate Retrieval for Zero-shot Entity Linking 16 Nov 2021 · 0 repositories
-
GLM: General Language Model Pretraining with Autoregressive Blank Infilling 16 Nov 2021 · 0 repositories
-
Graph-based Fine-grained Multimodal Attention Mechanism for Sentiment Analysis 16 Nov 2021 · 0 repositories
-
How does the pre-training objective affect what large language models learn about linguistic properties? 16 Nov 2021 · 0 repositories
-
Impact of Tokenization on Language Models: An Analysis for Turkish 16 Nov 2021 · 0 repositories
-
Improving Neural Models for Radiology Report Retrieval with Lexicon-based Automated Annotation 16 Nov 2021 · 0 repositories
-
Improving Unsupervised Sentence Simplification Using Fine-Tuned Masked Language Models 16 Nov 2021 · 0 repositories
-
Input-specific Attention Subnetworks for Adversarial Detection 16 Nov 2021 · 0 repositories
-
Integrated Semantic and Phonetic Post-correction for Chinese Speech Recognition 16 Nov 2021 · 1 repository · arXiv:2111.08400
-
Interpreting Language Models Through Knowledge Graph Extraction 16 Nov 2021 · 1 repository · arXiv:2111.08546
-
Interpreting the Robustness of Neural NLP Models to Textual Perturbations 16 Nov 2021 · 0 repositories
-
Investigating the Use of BERT Anchors for Bilingual Lexicon Induction with Minimal Supervision 16 Nov 2021 · 0 repositories
-
Is Neural Topic Modelling Better than Clustering? An Empirical Study on Clustering with Contextual Embeddings for Topics 16 Nov 2021 · 0 repositories
-
"Is Whole Word Masking Always Better for Chinese BERT?": Probing on Chinese Grammatical Error Correction 16 Nov 2021 · 0 repositories
-
KinyaBERT: a Morphology-aware Kinyarwanda Language Model 16 Nov 2021 · 0 repositories
-
Knowledge Enhanced Embedding: Improve Model Generalization Through Knowledge Graphs 16 Nov 2021 · 0 repositories
-
Language Level Classification on German Texts using a Neural Approach 16 Nov 2021 · 0 repositories
-
Learning to Ignore Adversarial Attacks 16 Nov 2021 · 0 repositories
-
Life after BERT: What do Other Muppets Understand about Language? 16 Nov 2021 · 0 repositories
-
Looking Into the Black Box - How Are Idioms Processed in BERT? 16 Nov 2021 · 0 repositories
-
LordBERT: Embedding Long Text by Segment Ordering with BERT 16 Nov 2021 · 0 repositories
-
MarCQAp: Effective Context Modeling for Conversational Question Answering 16 Nov 2021 · 0 repositories
-
MarkBERT: Marking Word Boundaries Improves Chinese BERT 16 Nov 2021 · 1 repository
-
MDERank: A Masked Document Embedding Rank Approach for Unsupervised Keyphrase Extraction 16 Nov 2021 · 0 repositories
-
Metadata Shaping: Natural Language Annotations for the Long Tail 16 Nov 2021 · 0 repositories
-
NSP-BERT: A Prompt-based Zero-Shot Learner Through an Original Pre-training Task —— Next Sentence Prediction 16 Nov 2021 · 0 repositories
-
On the Robustness of Reading Comprehension Models to Entity Renaming 16 Nov 2021 · 0 repositories
-
PALBERT: Teaching ALBERT to Ponder 16 Nov 2021 · 0 repositories
-
PARE: A Simple and Strong Baseline for Monolingual and Multilingual Distantly Supervised Relation Extraction 16 Nov 2021 · 0 repositories
-
Perturbations in the Wild: Leveraging Human-Written Text Perturbations for Realistic Adversarial Attack and Defense 16 Nov 2021 · 0 repositories
-
Pinyin-bert: A new solution to Chinese pinyin to character conversion task 16 Nov 2021 · 0 repositories
-
Probing BERT’s priors with serial reproduction chains 16 Nov 2021 · 0 repositories
-
PromptBERT: Improving BERT Sentence Embeddings with Prompts 16 Nov 2021 · 0 repositories
-
ReCo: Reliable Multi-hop Causal Reasoning via Structural Causal Recurrent Unit 16 Nov 2021 · 0 repositories
-
Representation of Ambiguity in Pre-Trained Sentence Embeddings 16 Nov 2021 · 0 repositories
-
SAMBERT: Improve Aspect Sentiment Triplet Extraction by Segmenting the Attention Maps of BERT 16 Nov 2021 · 0 repositories
-
Self-Supervised Contrastive Learning with Adversarial Perturbations for Robust Pretrained Language Models 16 Nov 2021 · 0 repositories
-
SHIELD: Defending Textual Neural Networks against Black-Box Adversarial Attacks with Stochastic Multi-Expert Patcher 16 Nov 2021 · 0 repositories
-
Softmax Bottleneck Makes Language Models Unable to Represent Multi-mode Word Distributions 16 Nov 2021 · 0 repositories
-
TACO: Pre-training of Deep Transformers with Attention Convolution using Disentangled Positional Representation 16 Nov 2021 · 0 repositories
-
The impact of lexical and grammatical processing on generating code from natural language 16 Nov 2021 · 0 repositories
-
Towards Fully Self-Supervised Learning of Knowledge from Unstructured Text 16 Nov 2021 · 0 repositories
-
Towards Improving Topic Models with the BERT-based Neural Topic Encoder 16 Nov 2021 · 0 repositories
-
Understanding Attention in Machine Reading Comprehension 16 Nov 2021 · 0 repositories
-
UNICON: Unsupervised Intent Discovery via Semantic-level Contrastive Learning 16 Nov 2021 · 0 repositories
-
Unsupervised multiple-choice question generation for out-of-domain Q&A fine-tuning 16 Nov 2021 · 0 repositories
-
Weight Squeezing: Reparameterization for Knowledge Transfer and Model Compression 16 Nov 2021 · 0 repositories
-
When classifying grammatical role, BERT doesn't care about word order... except when it matters 16 Nov 2021 · 0 repositories
-
Assessing gender bias in medical and scientific masked language models with StereoSet 15 Nov 2021 · 0 repositories · arXiv:2111.08088
-
Exploring Story Generation with Multi-task Objectives in Variational Autoencoders 15 Nov 2021 · 0 repositories · arXiv:2111.08133
-
IIITT@Dravidian-CodeMix-FIRE2021: Transliterate or translate? Sentiment analysis of code-mixed text in Dravidian languages 15 Nov 2021 · 1 repository · arXiv:2111.07906
-
Improving Prosody for Unseen Texts in Speech Synthesis by Utilizing Linguistic Information and Noisy Data 15 Nov 2021 · 0 repositories · arXiv:2111.07549
-
Scaling Law for Recommendation Models: Towards General-purpose User Representations 15 Nov 2021 · 0 repositories · arXiv:2111.11294
-
"Will You Find These Shortcuts?" A Protocol for Evaluating the Faithfulness of Input Salience Methods for Text Classification 14 Nov 2021 · 0 repositories · arXiv:2111.07367
-
SocialBERT -- Transformers for Online SocialNetwork Language Modelling 13 Nov 2021 · 0 repositories · arXiv:2111.07148
-
MS-LaTTE: A Dataset of Where and When To-do Tasks are Completed 12 Nov 2021 · 1 repository · arXiv:2111.06902
-
Character-level HyperNetworks for Hate Speech Detection 11 Nov 2021 · 1 repository · arXiv:2111.06336
-
Improving Large-scale Language Models and Resources for Filipino 11 Nov 2021 · 0 repositories · arXiv:2111.06053
-
Amazon SageMaker Model Parallelism: A General and Flexible Framework for Large Model Training 10 Nov 2021 · 0 repositories · arXiv:2111.05972
-
BagBERT: BERT-based bagging-stacking for multi-topic classification 10 Nov 2021 · 1 repository · arXiv:2111.05808
-
CEHR-BERT: Incorporating temporal information from structured EHR data to improve prediction tasks 10 Nov 2021 · 0 repositories · arXiv:2111.08585
-
Prune Once for All: Sparse Pre-Trained Language Models 10 Nov 2021 · 2 repositories · arXiv:2111.05754
-
DSBERT:Unsupervised Dialogue Structure learning with BERT 9 Nov 2021 · 0 repositories · arXiv:2111.04933
-
FPM: A Collection of Large-scale Foundation Pre-trained Language Models 9 Nov 2021 · 0 repositories · arXiv:2111.04909
-
Human-in-the-Loop Disinformation Detection: Stance, Sentiment, or Something Else? 9 Nov 2021 · 0 repositories · arXiv:2111.05139
-
AI-UPV at IberLEF-2021 DETOXIS task: Toxicity Detection in Immigration-Related Web News Comments Using Transformers and Statistical Models 8 Nov 2021 · 1 repository · arXiv:2111.04530
-
Chemical detection and indexing in PubMed full text articles using deep learning and rule-based methods 8 Nov 2021 · 0 repositories
-
Detecting Depression in Thai Blog Posts: a Dataset and a Baseline 8 Nov 2021 · 0 repositories · arXiv:2111.04574
-
Guiding Multi-Step Rearrangement Tasks with Natural Language Instructions 8 Nov 2021 · 2 repositories
-
Sexism Prediction in Spanish and English Tweets Using Monolingual and Multilingual BERT and Ensemble Models 8 Nov 2021 · 1 repository · arXiv:2111.04551
-
TACCL: Guiding Collective Algorithm Synthesis using Communication Sketches 8 Nov 2021 · 2 repositories · arXiv:2111.04867
-
TaCL: Improving BERT Pre-training with Token-aware Contrastive Learning 7 Nov 2021 · 2 repositories · arXiv:2111.04198
-
Profitable Trade-Off Between Memory and Performance In Multi-Domain Chatbot Architectures 6 Nov 2021 · 0 repositories · arXiv:2111.03963
-
Context-Aware Transformer Transducer for Speech Recognition 5 Nov 2021 · 0 repositories · arXiv:2111.03250
-
Effective Cross-Utterance Language Modeling for Conversational Speech Recognition 5 Nov 2021 · 0 repositories · arXiv:2111.03333
-
IBERT: Idiom Cloze-style reading comprehension with Attention 5 Nov 2021 · 0 repositories · arXiv:2112.02994
-
Sexism Identification in Tweets and Gabs using Deep Neural Networks 5 Nov 2021 · 0 repositories · arXiv:2111.03612