Methods › Natural Language Processing › Subword Segmentation › WordPiece › Papers, page 40
WordPiece
Papers archive 2025-07-28
archive papers tagged: 7,063 · with a code link: 2,910 · where Syntology ran a sample: 650 (529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (650 of 7,063 tagged: 529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument)
Page 40 of 71: papers 3,901 to 4,000 of 7,063, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Can Machine Learning Tools Support the Identification of Sustainable Design Leads From Product Reviews? Opportunities and Challenges 17 Dec 2021 · 0 repositories · arXiv:2112.09391
-
Challenging America: Modeling language in longer time scales 17 Dec 2021 · 0 repositories
-
Explain, Edit, and Understand: Rethinking User Study Design for Evaluating Model Explanations 17 Dec 2021 · 1 repository · arXiv:2112.09669
-
Joint Chinese Word Segmentation and Part-of-speech Tagging via Two-stage Span Labeling 17 Dec 2021 · 0 repositories · arXiv:2112.09488
-
Learning to Win Lottery Tickets in BERT Transfer via Task-agnostic Mask Training 17 Dec 2021 · 0 repositories
-
Rank4Class: A Ranking Formulation for Multiclass Classification 17 Dec 2021 · 0 repositories · arXiv:2112.09727
-
Towards Faithful Personalized Response Selection in Retrieval Based Dialog Systems 17 Dec 2021 · 0 repositories
-
An Empirical Study on Transfer Learning for Privilege Review 16 Dec 2021 · 0 repositories · arXiv:2112.08606
-
Knowledge-Augmented Language Models for Cause-Effect Relation Classification 16 Dec 2021 · 1 repository · arXiv:2112.08615Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Does Pre-training Induce Systematic Inference? How Masked Language Models Acquire Commonsense Knowledge 16 Dec 2021 · 0 repositories · arXiv:2112.08583
-
3D Question Answering 15 Dec 2021 · 0 repositories · arXiv:2112.08359
-
Applying SoftTriple Loss for Supervised Language Model Fine Tuning 15 Dec 2021 · 0 repositories · arXiv:2112.08462
-
Fine-Tuning Large Neural Language Models for Biomedical Natural Language Processing 15 Dec 2021 · 0 repositories · arXiv:2112.07869
-
One size does not fit all: Investigating strategies for differentially-private learning across NLP tasks 15 Dec 2021 · 1 repository · arXiv:2112.08159Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
One System to Rule them All: a Universal Intent Recognition System for Customer Service Chatbots 15 Dec 2021 · 0 repositories · arXiv:2112.08261
-
Tracing Text Provenance via Context-Aware Lexical Substitution 15 Dec 2021 · 0 repositories · arXiv:2112.07873
-
ACE-BERT: Adversarial Cross-modal Enhanced BERT for E-commerce Retrieval 14 Dec 2021 · 0 repositories · arXiv:2112.07209
-
Building on Huang et al. GlossBERT for Word Sense Disambiguation 14 Dec 2021 · 0 repositories · arXiv:2112.07089
-
Classifying Emails into Human vs Machine Category 14 Dec 2021 · 0 repositories · arXiv:2112.07742
-
CoCo-BERT: Improving Video-Language Pre-training with Contrastive Cross-modal Matching and Denoising 14 Dec 2021 · 0 repositories · arXiv:2112.07515
-
Epigenomic language models powered by Cerebras 14 Dec 2021 · 0 repositories · arXiv:2112.07571
-
From Dense to Sparse: Contrastive Pruning for Better Pre-trained Language Model Compression 14 Dec 2021 · 2 repositories · arXiv:2112.07198
-
Measuring Fairness with Biased Rulers: A Survey on Quantifying Biases in Pretrained Language Models 14 Dec 2021 · 1 repository · arXiv:2112.07447
-
Text Classification Models for Form Entity Linking 14 Dec 2021 · 1 repository · arXiv:2112.07443
-
Towards a Unified Foundation Model: Jointly Pre-Training Transformers on Unpaired Images and Text 14 Dec 2021 · 0 repositories · arXiv:2112.07074
-
A Study on Token Pruning for ColBERT 13 Dec 2021 · 0 repositories · arXiv:2112.06540
-
Measuring Context-Word Biases in Lexical Semantic Datasets 13 Dec 2021 · 0 repositories · arXiv:2112.06733
-
Do Data-based Curricula Work? 13 Dec 2021 · 0 repositories · arXiv:2112.06510
-
Keyphrase Generation Beyond the Boundaries of Title and Abstract 13 Dec 2021 · 1 repository · arXiv:2112.06776
-
Roof-Transformer: Divided and Joined Understanding with Knowledge Enhancement 13 Dec 2021 · 0 repositories · arXiv:2112.06736
-
WECHSEL: Effective initialization of subword embeddings for cross-lingual transfer of monolingual language models 13 Dec 2021 · 1 repository · arXiv:2112.06598Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Findings on Conversation Disentanglement 10 Dec 2021 · 0 repositories · arXiv:2112.05346
-
Multimodal Interactions Using Pretrained Unimodal Models for SIMMC 2.0 10 Dec 2021 · 1 repository · arXiv:2112.05328
-
Detecting potentially harmful and protective suicide-related content on twitter: A machine learning approach 9 Dec 2021 · 2 repositories · arXiv:2112.04796
-
From Scattered Sources to Comprehensive Technology Landscape: A Recommendation-based Retrieval Approach 9 Dec 2021 · 0 repositories · arXiv:2112.04810
-
Semantic Search as Extractive Paraphrase Span Detection 9 Dec 2021 · 1 repository · arXiv:2112.04886
-
Improving language models by retrieving from trillions of tokens 8 Dec 2021 · 2 repositories · arXiv:2112.04426Syntology 16 ran (of which 5 constructed an object rather than computing a result; 14 with no instrument failure: 0 honoured, 3 violated, 11 with no contract checked; 2 where Syntology's instrument failed) · 7 unverified (of 23 harvested samples) · 3 pointer-only (licence)
-
JABER and SABER: Junior and Senior Arabic BERt 8 Dec 2021 · 1 repository · arXiv:2112.04329
-
A Transferable Approach for Partitioning Machine Learning Models on Multi-Chip-Modules 7 Dec 2021 · 0 repositories · arXiv:2112.04041
-
raceBERT -- A Transformer-based Model for Predicting Race and Ethnicity from Names 7 Dec 2021 · 1 repository · arXiv:2112.03807
-
BERTMap: A BERT-based Ontology Alignment System 5 Dec 2021 · 1 repository · arXiv:2112.02682
-
Causal Distillation for Language Models 5 Dec 2021 · 1 repository · arXiv:2112.02505Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
DIBERT: Dependency Injected Bidirectional Encoder Representations from Transformers 5 Dec 2021 · 1 repository
-
VarCLR: Variable Semantic Representation Pre-training via Contrastive Learning 5 Dec 2021 · 1 repository · arXiv:2112.02650Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Bridging Pre-trained Models and Downstream Tasks for Source Code Understanding 4 Dec 2021 · 1 repository · arXiv:2112.02268Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 13 harvested samples)
-
Representation Learning for Conversational Data using Discourse Mutual Information Maximization 4 Dec 2021 · 0 repositories · arXiv:2112.05787
-
Unraveling Social Perceptions & Behaviors towards Migrants on Twitter 4 Dec 2021 · 0 repositories · arXiv:2112.06642
-
A Novel Deep Parallel Time-series Relation Network for Fault Diagnosis 3 Dec 2021 · 0 repositories · arXiv:2112.03405
-
Augmenting Customer Support with an NLP-based Receptionist 3 Dec 2021 · 0 repositories · arXiv:2112.01959
-
Given Users Recommendations Based on Reviews on Yelp 3 Dec 2021 · 1 repository · arXiv:2112.01762
-
NN-LUT: Neural Approximation of Non-Linear Operations for Efficient Transformer Inference 3 Dec 2021 · 0 repositories · arXiv:2112.02191
-
Siamese BERT-based Model for Web Search Relevance Ranking Evaluated on a New Czech Dataset 3 Dec 2021 · 1 repository · arXiv:2112.01810
-
BEVT: BERT Pretraining of Video Transformers 2 Dec 2021 · 1 repository · arXiv:2112.01529
-
PLSUM: Generating PT-BR Wikipedia by Summarizing Multiple Websites 2 Dec 2021 · 1 repository · arXiv:2112.01591
-
Unsupervised Law Article Mining based on Deep Pre-Trained Language Representation Models with Application to the Italian Civil Code 2 Dec 2021 · 0 repositories · arXiv:2112.03033
-
Domain-oriented Language Pre-training with Adaptive Hybrid Masking and Optimal Transport Alignment 1 Dec 2021 · 0 repositories · arXiv:2112.03024
-
DRONE: Data-aware Low-rank Compression for Large NLP Models 1 Dec 2021 · 0 repositories
-
NER-BERT: A Pre-trained Model for Low-Resource Entity Tagging 1 Dec 2021 · 0 repositories · arXiv:2112.00405
-
TriBERT: Human-centric Audio-visual Representation Learning 1 Dec 2021 · 1 repository
-
Wiki to Automotive: Understanding the Distribution Shift and its impact on Named Entity Recognition 1 Dec 2021 · 0 repositories · arXiv:2112.00283
-
A Comparative Study of Transformers on Word Sense Disambiguation 30 Nov 2021 · 0 repositories · arXiv:2111.15417
-
Generating Rich Product Descriptions for Conversational E-commerce Systems 30 Nov 2021 · 0 repositories · arXiv:2111.15298
-
KARL-Trans-NER: Knowledge Aware Representation Learning for Named Entity Recognition using Transformers 30 Nov 2021 · 0 repositories · arXiv:2111.15436
-
NLP Techniques for Water Quality Analysis in Social Media Content 30 Nov 2021 · 0 repositories · arXiv:2112.11441
-
Sentiment Analysis and Effect of COVID-19 Pandemic using College SubReddit Data 30 Nov 2021 · 1 repository · arXiv:2112.04351
-
SpaceEdit: Learning a Unified Editing Space for Open-Domain Image Editing 30 Nov 2021 · 0 repositories · arXiv:2112.00180
-
Text classification problems via BERT embedding method and graph convolutional neural network 30 Nov 2021 · 0 repositories · arXiv:2111.15379
-
Customer Sentiment Analysis using Weak Supervision for Customer-Agent Chat 29 Nov 2021 · 0 repositories · arXiv:2111.14282
-
Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point Modeling 29 Nov 2021 · 3 repositories · arXiv:2111.14819Syntology official (archive's flag): 4 ran · 5 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 2 pointer-only (licence)
-
Speech Tasks Relevant to Sleepiness Determined with Deep Transfer Learning 29 Nov 2021 · 0 repositories · arXiv:2111.14684
-
Tapping BERT for Preposition Sense Disambiguation 27 Nov 2021 · 0 repositories · arXiv:2111.13972
-
Predicting Document Coverage for Relation Extraction 26 Nov 2021 · 0 repositories · arXiv:2111.13611
-
Does constituency analysis enhance domain-specific pre-trained BERT models for relation extraction? 25 Nov 2021 · 0 repositories · arXiv:2112.02955
-
Evaluating the Robustness of Retrieval Pipelines with Query Variation Generators 25 Nov 2021 · 1 repository · arXiv:2111.13057
-
New Approaches to Long Document Summarization: Fourier Transform Based Attention in a Transformer Model 25 Nov 2021 · 0 repositories · arXiv:2111.15473
-
Probabilistic Impact Score Generation using Ktrain-BERT to Identify Hate Words from Twitter Discussions 25 Nov 2021 · 0 repositories · arXiv:2111.12939
-
Recommending Multiple Positive Citations for Manuscript via Content-Dependent Modeling and Multi-Positive Triplet 25 Nov 2021 · 0 repositories · arXiv:2111.12899
-
Transformer-based Korean Pretrained Language Models: A Survey on Three Years of Progress 25 Nov 2021 · 0 repositories · arXiv:2112.03014
-
PeCo: Perceptual Codebook for BERT Pre-training of Vision Transformers 24 Nov 2021 · 1 repository · arXiv:2111.12710
-
DABS: A Domain-Agnostic Benchmark for Self-Supervised Learning 23 Nov 2021 · 1 repository · arXiv:2111.12062Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Variational Learning for Unsupervised Knowledge Grounded Dialogs 23 Nov 2021 · 1 repository · arXiv:2112.00653
-
Can depth-adaptive BERT perform better on binary classification tasks 22 Nov 2021 · 0 repositories · arXiv:2111.10951
-
Does BERT look at sentiment lexicon? 19 Nov 2021 · 0 repositories · arXiv:2111.10100
-
Lexicon-based Methods vs. BERT for Text Sentiment Analysis 19 Nov 2021 · 0 repositories · arXiv:2111.10097
-
DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing 18 Nov 2021 · 3 repositories · arXiv:2111.09543Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
Dynamic-TinyBERT: Boost TinyBERT's Inference Efficiency by Dynamic Sequence Length 18 Nov 2021 · 0 repositories · arXiv:2111.09645
-
How Emotionally Stable is ALBERT? Testing Robustness with Stochastic Weight Averaging on a Sentiment Analysis Task 18 Nov 2021 · 1 repository · arXiv:2111.09612
-
LAnoBERT: System Log Anomaly Detection based on BERT Masked Language Model 18 Nov 2021 · 0 repositories · arXiv:2111.09564
-
RoBERTuito: a pre-trained language model for social media text in Spanish 18 Nov 2021 · 1 repository · arXiv:2111.09453Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples)
-
The Power of Selecting Key Blocks with Local Pre-ranking for Long Document Information Retrieval 18 Nov 2021 · 1 repository · arXiv:2111.09852
-
A Comparative Study on Transfer Learning and Distance Metrics in Semantic Clustering over the COVID-19 Tweets 16 Nov 2021 · 0 repositories · arXiv:2111.08658
-
A Flexible Multi-Task Model for BERT Serving 16 Nov 2021 · 0 repositories
-
A Graph Enhanced BERT Model for Event Prediction 16 Nov 2021 · 0 repositories
-
A Sentence is Worth 128 Pseudo Tokens: A Semantic-Aware Contrastive Learning Framework for Sentence Embeddings 16 Nov 2021 · 0 repositories
-
A Structured Semantic Reinforcement method for Task-Oriented Dialogue 16 Nov 2021 · 0 repositories
-
AdapLeR: Speeding up Inference by Adaptive Length Reduction 16 Nov 2021 · 0 repositories
-
An Isotropy Analysis in the Multilingual BERT Embedding Space 16 Nov 2021 · 0 repositories
-
ANNA: Enhanced Language Representation for Question Answering 16 Nov 2021 · 0 repositories
-
Assessing the Coherence Modeling Capabilities of Pretrained Transformer-based Language Models 16 Nov 2021 · 0 repositories
-
Attention-based Multi-hypothesis Fusion for Speech Summarization 16 Nov 2021 · 2 repositories · arXiv:2111.08201