Methods › Natural Language Processing › Subword Segmentation › WordPiece › Papers, page 49
WordPiece
Papers archive 2025-07-28
archive papers tagged: 7,063 · with a code link: 2,910 · where Syntology ran a sample: 650 (529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (650 of 7,063 tagged: 529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument)
Page 49 of 71: papers 4,801 to 4,900 of 7,063, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
DaLAJ - a dataset for linguistic acceptability judgments for Swedish: Format, baseline, sharing 14 May 2021 · 0 repositories · arXiv:2105.06681
-
Distilling BERT for low complexity network training 13 May 2021 · 0 repositories · arXiv:2105.06514
-
BertGCN: Transductive Text Classification by Combining GCN and BERT 12 May 2021 · 1 repository · arXiv:2105.05727
-
Better than BERT but Worse than Baseline 12 May 2021 · 0 repositories · arXiv:2105.05915
-
Building a Question and Answer System for News Domain 12 May 2021 · 0 repositories · arXiv:2105.05744
-
Evaluating Gender Bias in Natural Language Inference 12 May 2021 · 1 repository · arXiv:2105.05541
-
Go Beyond Plain Fine-tuning: Improving Pretrained Models for Social Commonsense 12 May 2021 · 0 repositories · arXiv:2105.05913
-
Kleister: Key Information Extraction Datasets Involving Long Documents with Complex Layouts 12 May 2021 · 0 repositories · arXiv:2105.05796
-
MATE-KD: Masked Adversarial TExt, a Companion to Knowledge Distillation 12 May 2021 · 1 repository · arXiv:2105.05912
-
OCHADAI-KYOTO at SemEval-2021 Task 1: Enhancing Model Generalization and Robustness for Lexical Complexity Prediction 12 May 2021 · 0 repositories · arXiv:2105.05535
-
Playing Codenames with Language Graphs and Word Embeddings 12 May 2021 · 1 repository · arXiv:2105.05885
-
Priberam at MESINESP Multi-label Classification of Medical Texts Task 12 May 2021 · 1 repository · arXiv:2105.05614
-
Priberam Labs at the NTCIR-15 SHINRA2020-ML: Classification Task 12 May 2021 · 0 repositories · arXiv:2105.05605
-
UIUC_BioNLP at SemEval-2021 Task 11: A Cascade of Neural Models for Structuring Scholarly NLP Contributions 12 May 2021 · 1 repository · arXiv:2105.05435
-
Addressing "Documentation Debt" in Machine Learning Research: A Retrospective Datasheet for BookCorpus 11 May 2021 · 1 repository · arXiv:2105.05241Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
BERT is to NLP what AlexNet is to CV: Can Pre-Trained Language Models Identify Analogies? 11 May 2021 · 1 repository · arXiv:2105.04949
-
Integrating extracted information from bert and multiple embedding methods with the deep neural network for humour detection 11 May 2021 · 0 repositories · arXiv:2105.05112
-
Role of Artificial Intelligence in Detection of Hateful Speech for Hinglish Data on Social Media 11 May 2021 · 0 repositories · arXiv:2105.04913
-
Assessing the Syntactic Capabilities of Transformer-based Multilingual Language Models 10 May 2021 · 0 repositories · arXiv:2105.04688
-
Automatic Classification of Human Translation and Machine Translation: A Study from the Perspective of Lexical Diversity 10 May 2021 · 0 repositories · arXiv:2105.04616
-
SRLF: A Stance-aware Reinforcement Learning Framework for Content-based Rumor Detection on Social Media 10 May 2021 · 0 repositories · arXiv:2105.04098
-
FNet: Mixing Tokens with Fourier Transforms 9 May 2021 · 12 repositories · arXiv:2105.03824Syntology community repositories only · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 1 pointer-only (licence)
-
Which transformer architecture fits my data? A vocabulary bottleneck in self-attention 9 May 2021 · 0 repositories · arXiv:2105.03928
-
Improving Named Entity Recognition by External Context Retrieving and Cooperative Learning 8 May 2021 · 3 repositories · arXiv:2105.03654Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
NLP-IIS@UT at SemEval-2021 Task 4: Machine Reading Comprehension using the Long Document Transformer 8 May 2021 · 0 repositories · arXiv:2105.03775
-
Adapting by Pruning: A Case Study on BERT 7 May 2021 · 1 repository · arXiv:2105.03343
-
Empirical Evaluation of Pre-trained Transformers for Human-Level NLP: The Role of Sample Size and Dimensionality 7 May 2021 · 1 repository · arXiv:2105.03484
-
Understanding by Understanding Not: Modeling Negation in Language Models 7 May 2021 · 1 repository · arXiv:2105.03519
-
Adapting Monolingual Models: Data can be Scarce when Language Similarity is High 6 May 2021 · 1 repository · arXiv:2105.02855
-
Aligning Subtitles in Sign Language Videos 6 May 2021 · 0 repositories · arXiv:2105.02877
-
Bird's Eye: Probing for Linguistic Graph Structures with a Simple Information-Theoretic Approach 6 May 2021 · 1 repository · arXiv:2105.02629
-
Introducing Information Retrieval for Biomedical Informatics Students 6 May 2021 · 1 repository · arXiv:2105.02746
-
TABBIE: Pretrained Representations of Tabular Data 6 May 2021 · 2 repositories · arXiv:2105.02584
-
Goldilocks: Just-Right Tuning of BERT for Technology-Assisted Review 3 May 2021 · 0 repositories · arXiv:2105.01044
-
SmoothI: Smooth Rank Indicators for Differentiable IR Metrics 3 May 2021 · 1 repository · arXiv:2105.00942
-
Unreasonable Effectiveness of Rule-Based Heuristics in Solving Russian SuperGLUE Tasks 3 May 2021 · 0 repositories · arXiv:2105.01192
-
MathBERT: A Pre-Trained Model for Mathematical Formula Understanding 2 May 2021 · 0 repositories · arXiv:2105.00377
-
MRCBert: A Machine Reading ComprehensionApproach for Unsupervised Summarization 1 May 2021 · 1 repository · arXiv:2105.00239
-
When to Fold'em: How to answer Unanswerable questions 1 May 2021 · 1 repository · arXiv:2105.00328
-
BERT Meets Relational DB: Contextual Representations of Relational Databases 30 Apr 2021 · 0 repositories · arXiv:2104.14914
-
Word Sense Disambiguation with Transformer Models 30 Apr 2021 · 0 repositories
-
Let's Play Mono-Poly: BERT Can Reveal Words' Polysemy Level and Partitionability into Senses 29 Apr 2021 · 1 repository · arXiv:2104.14694
-
Improving BERT Model Using Contrastive Learning for Biomedical Relation Extraction 28 Apr 2021 · 1 repository · arXiv:2104.13913
-
MelBERT: Metaphor Detection via Contextualized Late Interaction using Metaphorical Identification Theories 28 Apr 2021 · 1 repository · arXiv:2104.13615Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Multi-Task Learning of Query Intent and Named Entities using Transfer Learning 28 Apr 2021 · 0 repositories · arXiv:2105.03316
-
Societal Biases in Retrieved Contents: Measurement Framework and Adversarial Mitigation for BERT Rankers 28 Apr 2021 · 1 repository · arXiv:2104.13640
-
Multi-class Text Classification using BERT-based Active Learning 27 Apr 2021 · 0 repositories · arXiv:2104.14289
-
Semi-supervised Interactive Intent Labeling 27 Apr 2021 · 0 repositories · arXiv:2104.13406
-
UoT-UWF-PartAI at SemEval-2021 Task 5: Self Attention Based Bi-GRU with Multi-Embedding Representation for Toxicity Highlighter 27 Apr 2021 · 0 repositories · arXiv:2104.13164
-
Diverse Image Inpainting with Bidirectional and Autoregressive Transformers 26 Apr 2021 · 0 repositories · arXiv:2104.12335
-
MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding 26 Apr 2021 · 5 repositories · arXiv:2104.12763Syntology community repositories only · 7 ran (of which 4 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 1 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples)
-
Phrase break prediction with bidirectional encoder representations in Japanese text-to-speech synthesis 26 Apr 2021 · 1 repository · arXiv:2104.12395
-
Potential Idiomatic Expression (PIE)-English: Corpus for Classes of Idioms 25 Apr 2021 · 2 repositories · arXiv:2105.03280
-
Extract then Distill: Efficient and Effective Task-Agnostic BERT Distillation 24 Apr 2021 · 0 repositories · arXiv:2104.11928
-
Learning Passage Impacts for Inverted Indexes 24 Apr 2021 · 1 repository · arXiv:2104.12016Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Analysing Cyberbullying using Natural Language Processing by Understanding Jargon in Social Media 23 Apr 2021 · 0 repositories · arXiv:2107.08902
-
BERT-CoQAC: BERT-based Conversational Question Answering in Context 23 Apr 2021 · 0 repositories · arXiv:2104.11394
-
Comparative Analysis of Machine Learning and Deep Learning Algorithms for Detection of Online Hate Speech 23 Apr 2021 · 0 repositories · arXiv:2108.01063
-
Multimodal Fusion with BERT and Attention Mechanism for Fake News Detection 23 Apr 2021 · 1 repository · arXiv:2104.11476
-
Optimizing small BERTs trained for German NER 23 Apr 2021 · 2 repositories · arXiv:2104.11559
-
Towards Trustworthy Deception Detection: Benchmarking Model Robustness across Domains, Modalities, and Languages 23 Apr 2021 · 0 repositories · arXiv:2104.11761
-
On Geodesic Distances and Contextual Embedding Compression for Text Classification 22 Apr 2021 · 1 repository · arXiv:2104.11295
-
Discriminative Self-training for Punctuation Prediction 21 Apr 2021 · 0 repositories · arXiv:2104.10339
-
Disfluency Detection with Unlabeled Data and Small BERT Models 21 Apr 2021 · 0 repositories · arXiv:2104.10769
-
B-PROP: Bootstrapped Pre-training with Representative Words Prediction for Ad-hoc Retrieval 20 Apr 2021 · 1 repository · arXiv:2104.09791Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 2 honoured, 0 violated, 5 with no contract checked; 3 where Syntology's instrument failed) · 6 unverified (of 16 harvested samples) · 3 pointer-only (licence)
-
Efficient pre-training objectives for Transformers 20 Apr 2021 · 0 repositories · arXiv:2104.09694
-
Measuring Shifts in Attitudes Towards COVID-19 Measures in Belgium Using Multilingual BERT 20 Apr 2021 · 1 repository · arXiv:2104.09947
-
Subsentence Extraction from Text Using Coverage-Based Deep Learning Language Models 20 Apr 2021 · 1 repository · arXiv:2104.09777
-
UIT-ISE-NLP at SemEval-2021 Task 5: Toxic Spans Detection with BiLSTM-CRF and ToxicBERT Comment Classification 20 Apr 2021 · 1 repository · arXiv:2104.10100
-
WASSA@IITK at WASSA 2021: Multi-task Learning and Transformer Finetuning for Emotion Classification and Empathy Prediction 20 Apr 2021 · 0 repositories · arXiv:2104.09827
-
BigGreen at SemEval-2021 Task 1: Lexical Complexity Prediction with Assembly Models 19 Apr 2021 · 1 repository · arXiv:2104.09040
-
ELECTRAMed: a new pre-trained language representation model for biomedical NLP 19 Apr 2021 · 2 repositories · arXiv:2104.09585
-
Modeling "Newsworthiness" for Lead-Generation Across Corpora 19 Apr 2021 · 0 repositories · arXiv:2104.09653
-
Neural Language Models with Distant Supervision to Identify Major Depressive Disorder from Clinical Notes 19 Apr 2021 · 0 repositories · arXiv:2104.09644
-
OCTIS: Comparing and Optimizing Topic models is Simple! 19 Apr 2021 · 1 repository
-
Operationalizing a National Digital Library: The Case for a Norwegian Transformer Model 19 Apr 2021 · 2 repositories · arXiv:2104.09617
-
Probing for Bridging Inference in Transformer Language Models 19 Apr 2021 · 1 repository · arXiv:2104.09400
-
Sentiment Classification in Swahili Language Using Multilingual BERT 19 Apr 2021 · 0 repositories · arXiv:2104.09006
-
TeamUNCC@LT-EDI-EACL2021: Hope Speech Detection using Transfer Learning with Transformers 19 Apr 2021 · 1 repository
-
CEAR: Cross-Entity Aware Reranker for Knowledge Base Completion 18 Apr 2021 · 0 repositories · arXiv:2104.08741
-
Dual-View Distilled BERT for Sentence Embedding 18 Apr 2021 · 0 repositories · arXiv:2104.08675
-
FedNLP: Benchmarking Federated Learning Methods for Natural Language Processing Tasks 18 Apr 2021 · 1 repository · arXiv:2104.08815
-
Knowledge Neurons in Pretrained Transformers 18 Apr 2021 · 3 repositories · arXiv:2104.08696Syntology official (archive's flag): 1 ran · 2 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Language in a (Search) Box: Grounding Language Learning in Real-World Human-Machine Interaction 18 Apr 2021 · 0 repositories · arXiv:2104.08874
-
Rethinking Network Pruning -- under the Pre-train and Fine-tune Paradigm 18 Apr 2021 · 1 repository · arXiv:2104.08682Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample) · 1 pointer-only (licence)
-
SimCSE: Simple Contrastive Learning of Sentence Embeddings 18 Apr 2021 · 23 repositories · arXiv:2104.08821Syntology community repositories only · 17 ran (of which 9 constructed an object rather than computing a result; 12 with no instrument failure: 1 honoured, 0 violated, 11 with no contract checked; 5 where Syntology's instrument failed) · 13 unverified (of 30 harvested samples) · 19 pointer-only (licence)
-
Zero-shot Cross-lingual Transfer of Neural Machine Translation with Multilingual Pretrained Encoders 18 Apr 2021 · 1 repository · arXiv:2104.08757Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
A multilabel approach to morphosyntactic probing 17 Apr 2021 · 0 repositories · arXiv:2104.08464
-
ASBERT: Siamese and Triplet network embedding for open question answering 17 Apr 2021 · 0 repositories · arXiv:2104.08558
-
Co-BERT: A Context-Aware BERT Retrieval Model Incorporating Local and Query-specific Context 17 Apr 2021 · 0 repositories · arXiv:2104.08523
-
Frequency-based Distortions in Contextualized Word Embeddings 17 Apr 2021 · 0 repositories · arXiv:2104.08465
-
Three-level Hierarchical Transformer Networks for Long-sequence and Multiple Clinical Documents Classification 17 Apr 2021 · 1 repository · arXiv:2104.08444
-
Identifying the Limits of Cross-Domain Knowledge Transfer for Pretrained Models 17 Apr 2021 · 1 repository · arXiv:2104.08410
-
Improving Zero-Shot Cross-Lingual Transfer Learning via Robust Training 17 Apr 2021 · 1 repository · arXiv:2104.08645
-
Multi-source Neural Topic Modeling in Multi-view Embedding Spaces 17 Apr 2021 · 1 repository · arXiv:2104.08551
-
The Topic Confusion Task: A Novel Scenario for Authorship Attribution 17 Apr 2021 · 0 repositories · arXiv:2104.08530
-
UPB at SemEval-2021 Task 5: Virtual Adversarial Training for Toxic Spans Detection 17 Apr 2021 · 0 repositories · arXiv:2104.08635
-
Zero-shot Slot Filling with DPR and RAG 17 Apr 2021 · 2 repositories · arXiv:2104.08610
-
An Analysis of a BERT Deep Learning Strategy on a Technology Assisted Review Task 16 Apr 2021 · 0 repositories · arXiv:2104.08340
-
Memorisation versus Generalisation in Pre-trained Language Models 16 Apr 2021 · 1 repository · arXiv:2105.00828Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)