Methods › Natural Language Processing › Subword Segmentation › WordPiece › Papers, page 28
WordPiece
Papers archive 2025-07-28
archive papers tagged: 7,063 · with a code link: 2,910 · where Syntology ran a sample: 650 (529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (650 of 7,063 tagged: 529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument)
Page 28 of 71: papers 2,701 to 2,800 of 7,063, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Calibration of Transformer-based Models for Identifying Stress and Depression in Social Media 26 May 2023 · 0 repositories · arXiv:2305.16797
-
GeoVLN: Learning Geometry-Enhanced Visual Representation with Slot Attention for Vision-and-Language Navigation 26 May 2023 · 1 repository · arXiv:2305.17102
-
Incorporating Distributions of Discourse Structure for Long Document Abstractive Summarization 26 May 2023 · 0 repositories · arXiv:2305.16784
-
KNSE: A Knowledge-aware Natural Language Inference Framework for Dialogue Symptom Status Recognition 26 May 2023 · 0 repositories · arXiv:2305.16833
-
Theoretical and Practical Perspectives on what Influence Functions Do 26 May 2023 · 0 repositories · arXiv:2305.16971
-
Zero is Not Hero Yet: Benchmarking Zero-Shot Performance of LLMs for Financial Tasks 26 May 2023 · 1 repository · arXiv:2305.16633
-
Comparative Study of Pre-Trained BERT Models for Code-Mixed Hindi-English Data 25 May 2023 · 0 repositories · arXiv:2305.15722
-
Context-aware attention layers coupled with optimal transport domain adaptation and multimodal fusion methods for recognizing dementia from spontaneous speech 25 May 2023 · 0 repositories · arXiv:2305.16406
-
Not wacky vs. definitely wacky: A study of scalar adverbs in pretrained language models 25 May 2023 · 0 repositories · arXiv:2305.16426
-
Text-to-Motion Retrieval: Towards Joint Understanding of Human Motion Data and Natural Language 25 May 2023 · 1 repository · arXiv:2305.15842
-
A Causal View of Entity Bias in (Large) Language Models 24 May 2023 · 1 repository · arXiv:2305.14695Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Complex Mathematical Symbol Definition Structures: A Dataset and Model for Coordination Resolution in Definition Extraction 24 May 2023 · 1 repository · arXiv:2305.14660
-
Context-Aware Transformer Pre-Training for Answer Sentence Selection 24 May 2023 · 0 repositories · arXiv:2305.15358
-
Dynamic Masking Rate Schedules for MLM Pretraining 24 May 2023 · 0 repositories · arXiv:2305.15096
-
Extracting Psychological Indicators Using Question Answering 24 May 2023 · 0 repositories · arXiv:2305.14891
-
Ghostbuster: Detecting Text Ghostwritten by Large Language Models 24 May 2023 · 2 repositories · arXiv:2305.15047
-
How to Distill your BERT: An Empirical Study on the Impact of Weight Initialisation and Distillation Objectives 24 May 2023 · 1 repository · arXiv:2305.15032Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (of 11 harvested samples)
-
Neural Summarization of Electronic Health Records 24 May 2023 · 0 repositories · arXiv:2305.15222
-
Revisiting Token Dropping Strategy in Efficient BERT Pretraining 24 May 2023 · 1 repository · arXiv:2305.15273
-
Few-shot Adaptation to Distribution Shifts By Mixing Source and Target Embeddings 23 May 2023 · 0 repositories · arXiv:2305.14521
-
When your Cousin has the Right Connections: Unsupervised Bilingual Lexicon Induction for Related Data-Imbalanced Languages 23 May 2023 · 1 repository · arXiv:2305.14012
-
All Roads Lead to Rome? Exploring the Invariance of Transformers' Representations 23 May 2023 · 1 repository · arXiv:2305.14555Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Assessing Linguistic Generalisation in Language Models: A Dataset for Brazilian Portuguese 23 May 2023 · 0 repositories · arXiv:2305.14070
-
AxomiyaBERTa: A Phonologically-aware Transformer Model for Assamese 23 May 2023 · 1 repository · arXiv:2305.13641
-
Connecting the Dots: What Graph-Based Text Representations Work Best for Text Classification Using Graph Neural Networks? 23 May 2023 · 1 repository · arXiv:2305.14578
-
Exploring Large Language Models for Classical Philology 23 May 2023 · 1 repository · arXiv:2305.13698
-
Handling Realistic Label Noise in BERT Text Classification 23 May 2023 · 0 repositories · arXiv:2305.16337
-
On Robustness of Finetuned Transformer-based NLP Models 23 May 2023 · 1 repository · arXiv:2305.14453
-
Text Is All You Need: Learning Language Representations for Sequential Recommendation 23 May 2023 · 1 repository · arXiv:2305.13731Syntology 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Training Transitive and Commutative Multimodal Transformers with LoReTTa 23 May 2023 · 0 repositories · arXiv:2305.14243
-
Exploring Energy-based Language Models with Different Architectures and Training Methods for Speech Recognition 22 May 2023 · 2 repositories · arXiv:2305.12676
-
GATology for Linguistics: What Syntactic Dependencies It Knows 22 May 2023 · 0 repositories · arXiv:2305.13403
-
SimCSE++: Improving Contrastive Learning for Sentence Embeddings from Two Perspectives 22 May 2023 · 0 repositories · arXiv:2305.13192
-
Language-Agnostic Bias Detection in Language Models with Bias Probing 22 May 2023 · 1 repository · arXiv:2305.13302
-
Atomic Inference for NLI with Generated Facts as Atoms 22 May 2023 · 1 repository · arXiv:2305.13214
-
Stock and market index prediction using Informer network 22 May 2023 · 0 repositories · arXiv:2305.14382
-
Syntactic Knowledge via Graph Attention with BERT in Machine Translation 22 May 2023 · 0 repositories · arXiv:2305.13413
-
A Deeper (Autoregressive) Approach to Non-Convergent Discourse Parsing 21 May 2023 · 0 repositories · arXiv:2305.12510
-
A Symbolic Framework for Evaluating Mathematical Reasoning and Generalisation with Transformers 21 May 2023 · 0 repositories · arXiv:2305.12563
-
BertRLFuzzer: A BERT and Reinforcement Learning Based Fuzzer 21 May 2023 · 1 repository · arXiv:2305.12534
-
F-PABEE: Flexible-patience-based Early Exiting for Single-label and Multi-label text Classification Tasks 21 May 2023 · 0 repositories · arXiv:2305.11916
-
Infor-Coef: Information Bottleneck-based Dynamic Token Downsampling for Compact and Efficient language model 21 May 2023 · 0 repositories · arXiv:2305.12458
-
IR Models and the COVID-19 Pandemic: A Comparative Study of Performance and Challenges 21 May 2023 · 0 repositories · arXiv:2305.12528
-
CDJUR-BR -- A Golden Collection of Legal Document from Brazilian Justice with Fine-Grained Named Entities 20 May 2023 · 0 repositories · arXiv:2305.18315
-
SEntFiN 1.0: Entity-Aware Sentiment Analysis for Financial News 20 May 2023 · 0 repositories · arXiv:2305.12257
-
A Sequence-to-Sequence Approach for Arabic Pronoun Resolution 19 May 2023 · 0 repositories · arXiv:2305.11529
-
Eye-SpatialNet: Spatial Information Extraction from Ophthalmology Notes 19 May 2023 · 0 repositories · arXiv:2305.11948
-
Federated Foundation Models: Privacy-Preserving and Collaborative Learning for Large Models 19 May 2023 · 0 repositories · arXiv:2305.11414
-
Language-Universal Phonetic Representation in Multilingual Speech Pretraining for Low-Resource Speech Recognition 19 May 2023 · 0 repositories · arXiv:2305.11569
-
Ahead-of-Time P-Tuning 18 May 2023 · 0 repositories · arXiv:2305.10835
-
Ditto: A Simple and Efficient Approach to Improve Sentence Embeddings 18 May 2023 · 1 repository · arXiv:2305.10786
-
PDP: Parameter-free Differentiable Pruning is All You Need 18 May 2023 · 0 repositories · arXiv:2305.11203
-
Trading Syntax Trees for Wordpieces: Target-oriented Opinion Words Extraction with Wordpieces and Aspect Enhancement 18 May 2023 · 0 repositories · arXiv:2305.11034
-
A quantitative study of NLP approaches to question difficulty estimation 17 May 2023 · 1 repository · arXiv:2305.10236
-
AD-KD: Attribution-Driven Knowledge Distillation for Language Model Compression 17 May 2023 · 1 repository · arXiv:2305.10010Syntology official (archive's flag): 3 ran · 3 ran (of which 1 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Explaining black box text modules in natural language with language models 17 May 2023 · 2 repositories · arXiv:2305.09863Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Solving Cosine Similarity Underestimation between High Frequency Words by L2 Norm Discounting 17 May 2023 · 1 repository · arXiv:2305.10610
-
CWTM: Leveraging Contextualized Word Embeddings from BERT for Neural Topic Modeling 16 May 2023 · 1 repository · arXiv:2305.09329
-
Measuring Dimensions of Self-Presentation in Twitter Bios and their Links to Misinformation Sharing 16 May 2023 · 1 repository · arXiv:2305.09548
-
Weight-Inherited Distillation for Task-Agnostic BERT Compression 16 May 2023 · 1 repository · arXiv:2305.09098
-
Coreference-aware Double-channel Attention Network for Multi-party Dialogue Reading Comprehension 15 May 2023 · 1 repository · arXiv:2305.08348
-
Knowledge Rumination for Pre-trained Language Models 15 May 2023 · 1 repository · arXiv:2305.08732
-
Private Training Set Inspection in MLaaS 15 May 2023 · 0 repositories · arXiv:2305.09058
-
Text2Gender: A Deep Learning Architecture for Analysis of Blogger's Age and Gender 15 May 2023 · 0 repositories · arXiv:2305.08633
-
MatSci-NLP: Evaluating Scientific Language Models on Materials Science Language Tasks Using Text-to-Schema Modeling 14 May 2023 · 1 repository · arXiv:2305.08264Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample) · 1 pointer-only (licence)
-
GPT-Sentinel: Distinguishing Human and ChatGPT Generated Content 13 May 2023 · 2 repositories · arXiv:2305.07969
-
PESTS: Persian_English Cross Lingual Corpus for Semantic Textual Similarity 13 May 2023 · 0 repositories · arXiv:2305.07893
-
A General-Purpose Multilingual Document Encoder 11 May 2023 · 1 repository · arXiv:2305.07016
-
A Method to Automate the Discharge Summary Hospital Course for Neurology Patients 10 May 2023 · 0 repositories · arXiv:2305.06416
-
Enriching language models with graph-based context information to better understand textual data 10 May 2023 · 1 repository · arXiv:2305.11070
-
A Review of Vision-Language Models and their Performance on the Hateful Memes Challenge 9 May 2023 · 1 repository · arXiv:2305.06159
-
Alleviating Over-smoothing for Unsupervised Sentence Representation 9 May 2023 · 1 repository · arXiv:2305.06154Syntology official: harvested, nothing ran · 0 ran · 4 unverified (of 4 harvested samples)
-
Attack Named Entity Recognition by Entity Boundary Interference 9 May 2023 · 0 repositories · arXiv:2305.05253
-
Detection of depression on social networks using transformers and ensembles 9 May 2023 · 1 repository · arXiv:2305.05325
-
Effects of sub-word segmentation on performance of transformer language models 9 May 2023 · 0 repositories · arXiv:2305.05480
-
StrAE: Autoencoding for Pre-Trained Embeddings using Explicit Structure 9 May 2023 · 0 repositories · arXiv:2305.05588
-
GersteinLab at MEDIQA-Chat 2023: Clinical Note Summarization from Doctor-Patient Conversations through Fine-tuning and In-context Learning 8 May 2023 · 0 repositories · arXiv:2305.05001
-
PreCog: Exploring the Relation between Memorization and Performance in Pre-trained Language Models 8 May 2023 · 0 repositories · arXiv:2305.04673
-
Vulnerability Detection Using Two-Stage Deep Learning Models 8 May 2023 · 0 repositories · arXiv:2305.09673
-
Stanford MLab at SemEval-2023 Task 10: Exploring GloVe- and Transformer-Based Methods for the Explainable Detection of Online Sexism 7 May 2023 · 0 repositories · arXiv:2305.04356
-
On the Usage of Continual Learning for Out-of-Distribution Generalization in Pre-trained Language Models of Code 6 May 2023 · 0 repositories · arXiv:2305.04106
-
Pre-training Language Model as a Multi-perspective Course Learner 6 May 2023 · 0 repositories · arXiv:2305.03981
-
Rhetorical Role Labeling of Legal Documents using Transformers and Graph Neural Networks 6 May 2023 · 0 repositories · arXiv:2305.04100
-
Block the Label and Noise: An N-Gram Masked Speller for Chinese Spell Checking 5 May 2023 · 0 repositories · arXiv:2305.03314
-
CLaC at SemEval-2023 Task 2: Comparing Span-Prediction and Sequence-Labeling approaches for NER 5 May 2023 · 0 repositories · arXiv:2305.03845
-
Harnessing the Power of BERT in the Turkish Clinical Domain: Pretraining Approaches for Limited Data Scenarios 5 May 2023 · 0 repositories · arXiv:2305.03788
-
Predicting COVID-19 and pneumonia complications from admission texts 5 May 2023 · 0 repositories · arXiv:2305.03661
-
Using ChatGPT for Entity Matching 5 May 2023 · 1 repository · arXiv:2305.03423
-
Enhancing Pashto Text Classification using Language Processing Techniques for Single And Multi-Label Analysis 4 May 2023 · 0 repositories · arXiv:2305.03201
-
Improving Code Example Recommendations on Informal Documentation Using BERT and Query-Aware LSH: A Comparative Study 4 May 2023 · 1 repository · arXiv:2305.03017
-
Leveraging BERT Language Model for Arabic Long Document Classification 4 May 2023 · 0 repositories · arXiv:2305.03519
-
A Novel Plagiarism Detection Approach Combining BERT-based Word Embedding, Attention-based LSTMs and an Improved Differential Evolution Algorithm 3 May 2023 · 0 repositories · arXiv:2305.02374
-
evaluating bert and parsbert for analyzing persian advertisement data 3 May 2023 · 0 repositories · arXiv:2305.02426
-
Evaluating BERT-based Scientific Relation Classifiers for Scholarly Knowledge Graph Construction on Digital Library Collections 3 May 2023 · 0 repositories · arXiv:2305.02291
-
Exploring Linguistic Properties of Monolingual BERTs with Typological Classification among Languages 3 May 2023 · 0 repositories · arXiv:2305.02215
-
Improving Cancer Hallmark Classification with BERT-based Deep Learning Approach 2 May 2023 · 0 repositories · arXiv:2305.03501
-
Unlimiformer: Long-Range Transformers with Unlimited Length Input 2 May 2023 · 2 repositories · arXiv:2305.01625Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
Logion: Machine Learning for Greek Philology 1 May 2023 · 0 repositories · arXiv:2305.01099
-
Retrieving Comparative Arguments using Ensemble Methods and Neural Information Retrieval 1 May 2023 · 0 repositories · arXiv:2305.01513
-
SafeWebUH at SemEval-2023 Task 11: Learning Annotator Disagreement in Derogatory Text: Comparison of Direct Training vs Aggregation 1 May 2023 · 1 repository · arXiv:2305.01050