Methods › Natural Language Processing › Subword Segmentation › WordPiece › Papers, page 31
WordPiece
Papers archive 2025-07-28
archive papers tagged: 7,063 · with a code link: 2,910 · where Syntology ran a sample: 650 (529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (650 of 7,063 tagged: 529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument)
Page 31 of 71: papers 3,001 to 3,100 of 7,063, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
A Multi-View Joint Learning Framework for Embedding Clinical Codes and Text Using Graph Neural Networks 27 Jan 2023 · 0 repositories · arXiv:2301.11608
-
Can We Use Probing to Better Understand Fine-tuning and Knowledge Distillation of the BERT NLU? 27 Jan 2023 · 0 repositories · arXiv:2301.11688
-
Context Matters: A Strategy to Pre-train Language Model for Science Education 27 Jan 2023 · 0 repositories · arXiv:2301.12031
-
Predicting Sentence-Level Factuality of News and Bias of Media Outlets 27 Jan 2023 · 1 repository · arXiv:2301.11850
-
Understanding INT4 Quantization for Transformer Models: Latency Speedup, Composability, and Failure Cases 27 Jan 2023 · 1 repository · arXiv:2301.12017
-
A benchmark for toxic comment classification on Civil Comments dataset 26 Jan 2023 · 1 repository · arXiv:2301.11125
-
BERT-Embedding and Citation Network Analysis based Query Expansion Technique for Scholarly Search 26 Jan 2023 · 0 repositories · arXiv:2301.11069
-
A Stability Analysis of Fine-Tuning a Pre-Trained Model 24 Jan 2023 · 0 repositories · arXiv:2301.09820
-
Multitask Instruction-based Prompting for Fallacy Recognition 24 Jan 2023 · 0 repositories · arXiv:2301.09992
-
Injecting the BM25 Score as Text Improves BERT-Based Re-rankers 23 Jan 2023 · 1 repository · arXiv:2301.09728
-
StockEmotions: Discover Investor Emotions for Financial Sentiment Analysis and Multivariate Time Series 23 Jan 2023 · 2 repositories · arXiv:2301.09279
-
Stress Test for BERT and Deep Models: Predicting Words from Italian Poetry 21 Jan 2023 · 0 repositories · arXiv:2302.09303
-
Phoneme-Level BERT for Enhanced Prosody of Text-to-Speech with Grapheme Predictions 20 Jan 2023 · 2 repositories · arXiv:2301.08810Syntology 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples)
-
Which Features are Learned by CodeBert: An Empirical Study of the BERT-based Source Code Representation Learning 20 Jan 2023 · 0 repositories · arXiv:2301.08427
-
An Error-Guided Correction Model for Chinese Spelling Error Correction 16 Jan 2023 · 1 repository · arXiv:2301.06323
-
TEDB System Description to a Shared Task on Euphemism Detection 2022 16 Jan 2023 · 1 repository · arXiv:2301.06602
-
Improving Noise Robustness for Spoken Content Retrieval using Semi-supervised ASR and N-best Transcripts for BERT-based Ranking Models 15 Jan 2023 · 0 repositories · arXiv:2301.06056
-
NarrowBERT: Accelerating Masked Language Model Pretraining and Inference 11 Jan 2023 · 1 repository · arXiv:2301.04761Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Topics in Contextualised Attention Embeddings 11 Jan 2023 · 0 repositories · arXiv:2301.04339
-
Language Models sounds the Death Knell of Knowledge Graphs 10 Jan 2023 · 0 repositories · arXiv:2301.03980
-
Designing BERT for Convolutional Networks: Sparse and Hierarchical Masked Modeling 9 Jan 2023 · 2 repositories · arXiv:2301.03580Syntology official (archive's flag): 9 ran · 13 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 14 harvested samples) · 2 pointer-only (licence)
-
Online Fake Review Detection Using Supervised Machine Learning And BERT Model 9 Jan 2023 · 0 repositories · arXiv:2301.03225
-
App Review Driven Collaborative Bug Finding 7 Jan 2023 · 1 repository · arXiv:2301.02818
-
RLAS-BIABC: A Reinforcement Learning-Based Answer Selection Using the BERT Model Boosted by an Improved ABC Algorithm 7 Jan 2023 · 0 repositories · arXiv:2301.02807
-
Learning Trajectory-Word Alignments for Video-Language Tasks 5 Jan 2023 · 0 repositories · arXiv:2301.01953
-
PIE-QG: Paraphrased Information Extraction for Unsupervised Question Generation from Small Corpora 3 Jan 2023 · 0 repositories · arXiv:2301.01064
-
Understanding Political Polarisation using Language Models: A dataset and method 2 Jan 2023 · 0 repositories · arXiv:2301.00891
-
Floods Relevancy and Identification of Location from Twitter Posts using NLP Techniques 1 Jan 2023 · 0 repositories · arXiv:2301.00321
-
Leveraging Semantic Representations Combined with Contextual Word Representations for Recognizing Textual Entailment in Vietnamese 1 Jan 2023 · 0 repositories · arXiv:2301.00422
-
Sentiment Analysis of COVID-19 Public Activity Restriction (PPKM) Impact using BERT Method 31 Dec 2022 · 0 repositories · arXiv:2301.00096
-
An Analysis of Attention via the Lens of Exchangeability and Latent Variable Models 30 Dec 2022 · 0 repositories · arXiv:2212.14852
-
Distant Reading of the German Coalition Deal: Recognizing Policy Positions with BERT-based Text Classification 30 Dec 2022 · 0 repositories · arXiv:2212.14648
-
Cramming: Training a Language Model on a Single GPU in One Day 28 Dec 2022 · 1 repository · arXiv:2212.14034Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 2 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
A Survey on Knowledge-Enhanced Pre-trained Language Models 27 Dec 2022 · 0 repositories · arXiv:2212.13428
-
Benchmark for Uncertainty & Robustness in Self-Supervised Learning 23 Dec 2022 · 1 repository · arXiv:2212.12411
-
Finetuning for Sarcasm Detection with a Pruned Dataset 23 Dec 2022 · 1 repository · arXiv:2212.12213
-
CAMeMBERT: Cascading Assistant-Mediated Multilingual BERT 22 Dec 2022 · 0 repositories · arXiv:2212.11456
-
Enhancing the prediction of disease outcomes using electronic health records and pretrained deep learning models 22 Dec 2022 · 0 repositories · arXiv:2212.12067
-
Analyzing Semantic Faithfulness of Language Models via Input Intervention on Question Answering 21 Dec 2022 · 1 repository · arXiv:2212.10696
-
Automatic Emotion Modelling in Written Stories 21 Dec 2022 · 1 repository · arXiv:2212.11382
-
Cross-Linguistic Syntactic Difference in Multilingual BERT: How Good is It and How Does It Affect Transfer? 21 Dec 2022 · 1 repository · arXiv:2212.10879
-
Towards Efficient Visual Simplification of Computational Graphs in Deep Neural Networks 21 Dec 2022 · 0 repositories · arXiv:2212.10774
-
A Twitter BERT Approach for Offensive Language Detection in Marathi 20 Dec 2022 · 0 repositories · arXiv:2212.10039
-
Hybrid Rule-Neural Coreference Resolution System based on Actor-Critic Learning 20 Dec 2022 · 0 repositories · arXiv:2212.10087
-
Identifying and Manipulating the Personality Traits of Language Models 20 Dec 2022 · 0 repositories · arXiv:2212.10276
-
In and Out-of-Domain Text Adversarial Robustness via Label Smoothing 20 Dec 2022 · 0 repositories · arXiv:2212.10258
-
Parameter-efficient Zero-shot Transfer for Cross-Language Dense Retrieval with Adapters 20 Dec 2022 · 0 repositories · arXiv:2212.10448
-
Pretraining Without Attention 20 Dec 2022 · 1 repository · arXiv:2212.10544Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Do CoNLL-2003 Named Entity Taggers Still Work Well in 2023? 19 Dec 2022 · 1 repository · arXiv:2212.09747
-
Enriching Relation Extraction with OpenIE 19 Dec 2022 · 0 repositories · arXiv:2212.09376
-
Less is More: Parameter-Free Text Classification with Gzip 19 Dec 2022 · 0 repositories · arXiv:2212.09410
-
MANTIS at TSAR-2022 Shared Task: Improved Unsupervised Lexical Simplification with Pretrained Encoders 19 Dec 2022 · 0 repositories · arXiv:2212.09855
-
Bort: Towards Explainable Neural Networks with Bounded Orthogonal Constraint 18 Dec 2022 · 1 repository · arXiv:2212.09062Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Neural Coreference Resolution based on Reinforcement Learning 18 Dec 2022 · 0 repositories · arXiv:2212.09028
-
Neural Rankers for Effective Screening Prioritisation in Medical Systematic Review Literature Search 18 Dec 2022 · 0 repositories · arXiv:2212.09017
-
Exploiting Rich Textual User-Product Context for Improving Sentiment Analysis 17 Dec 2022 · 0 repositories · arXiv:2212.08888
-
LegalRelectra: Mixed-domain Language Modeling for Long-range Legal Text Comprehension 16 Dec 2022 · 0 repositories · arXiv:2212.08204
-
Plansformer: Generating Symbolic Plans using Transformers 16 Dec 2022 · 0 repositories · arXiv:2212.08681
-
POIBERT: A Transformer-based Model for the Tour Recommendation Problem 16 Dec 2022 · 0 repositories · arXiv:2212.13900
-
ReCo: Reliable Causal Chain Reasoning via Structural Causal Recurrent Neural Networks 16 Dec 2022 · 1 repository · arXiv:2212.08322Syntology official (archive's flag): 4 ran · 4 ran (of which 2 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Utilizing distilBert transformer model for sentiment classification of COVID-19's Persian open-text responses 16 Dec 2022 · 0 repositories · arXiv:2212.08407
-
Efficient Pre-training of Masked Language Model via Concept-based Curriculum Masking 15 Dec 2022 · 1 repository · arXiv:2212.07617Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Visually-augmented pretrained language models for NLP tasks without images 15 Dec 2022 · 1 repository · arXiv:2212.07937
-
Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and Language 14 Dec 2022 · 5 repositories · arXiv:2212.07525Syntology community repositories only · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 8 harvested samples) · 2 pointer-only (licence)
-
Evaluating Byte and Wordpiece Level Models for Massively Multilingual Semantic Parsing 14 Dec 2022 · 0 repositories · arXiv:2212.07223
-
Explainability of Text Processing and Retrieval Methods: A Critical Survey 14 Dec 2022 · 0 repositories · arXiv:2212.07126
-
DexBERT: Effective, Task-Agnostic and Fine-grained Representation Learning of Android Bytecode 12 Dec 2022 · 1 repository · arXiv:2212.05976
-
Classifying the Ideological Orientation of User-Submitted Texts in Social Media 12 Dec 2022 · 1 repository
-
Punctuation Restoration for Singaporean Spoken Languages: English, Malay, and Mandarin 10 Dec 2022 · 1 repository · arXiv:2212.05356
-
Incorporating Emotions into Health Mention Classification Task on Social Media 9 Dec 2022 · 1 repository · arXiv:2212.05039
-
Explain to me like I am five -- Sentence Simplification Using Transformers 8 Dec 2022 · 1 repository · arXiv:2212.04595
-
Memorization of Named Entities in Fine-tuned BERT Models 7 Dec 2022 · 1 repository · arXiv:2212.03749
-
Learning-To-Embed: Adopting Transformer based models for E-commerce Products Representation Learning 7 Dec 2022 · 0 repositories · arXiv:2212.03725
-
SimVTP: Simple Video Text Pre-training with Masked Autoencoders 7 Dec 2022 · 0 repositories · arXiv:2212.03490
-
TweetDrought: A Deep-Learning Drought Impacts Recognizer based on Twitter Data 7 Dec 2022 · 0 repositories · arXiv:2212.04001
-
CySecBERT: A Domain-Adapted Language Model for the Cybersecurity Domain 6 Dec 2022 · 0 repositories · arXiv:2212.02974
-
Vision Transformer Computation and Resilience for Dynamic Inference 6 Dec 2022 · 0 repositories · arXiv:2212.02687
-
LUNA: Language Understanding with Number Augmentations on Transformers via Number Plugins and Pre-training 6 Dec 2022 · 1 repository · arXiv:2212.02691Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Modern French Poetry Generation with RoBERTa and GPT-2 6 Dec 2022 · 0 repositories · arXiv:2212.02911
-
Style transfer and classification in hebrew news items 6 Dec 2022 · 0 repositories · arXiv:2212.03019
-
Video Games as a Corpus: Sentiment Analysis using Fallout New Vegas Dialog 5 Dec 2022 · 0 repositories · arXiv:2212.02168
-
ColD Fusion: Collaborative Descent for Distributed Multitask Finetuning 2 Dec 2022 · 0 repositories · arXiv:2212.01378
-
Event knowledge in large language models: the gap between the impossible and the unlikely 2 Dec 2022 · 1 repository · arXiv:2212.01488
-
Adapted Multimodal BERT with Layer-wise Fusion for Sentiment Analysis 1 Dec 2022 · 0 repositories · arXiv:2212.00678
-
BudgetLongformer: Can we Cheaply Pretrain a SotA Legal Language Model From Scratch? 30 Nov 2022 · 0 repositories · arXiv:2211.17135
-
ExtremeBERT: A Toolkit for Accelerating Pretraining of Customized BERT 30 Nov 2022 · 1 repository · arXiv:2211.17201Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
HEAT: Hardware-Efficient Automatic Tensor Decomposition for Transformer Compression 30 Nov 2022 · 0 repositories · arXiv:2211.16749
-
Composition based oxidation state prediction of materials using deep learning 29 Nov 2022 · 1 repository · arXiv:2211.15895
-
Diverse Multi-Answer Retrieval with Determinantal Point Processes 29 Nov 2022 · 0 repositories · arXiv:2211.16029
-
Outfit Generation and Recommendation -- An Experimental Study 29 Nov 2022 · 0 repositories · arXiv:2211.16353
-
Automatically Extracting Information in Medical Dialogue: Expert System And Attention for Labelling 28 Nov 2022 · 0 repositories · arXiv:2211.15544
-
DiffusionBERT: Improving Generative Masked Language Models with Diffusion Models 28 Nov 2022 · 1 repository · arXiv:2211.15029Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples) · 2 pointer-only (licence)
-
Handling and extracting key entities from customer conversations using Speech recognition and Named Entity recognition 28 Nov 2022 · 0 repositories · arXiv:2211.17107
-
Is it Required? Ranking the Skills Required for a Job-Title 28 Nov 2022 · 0 repositories · arXiv:2212.08553
-
Revisiting Distance Metric Learning for Few-Shot Natural Language Classification 28 Nov 2022 · 0 repositories · arXiv:2211.15202
-
Scientific and Creative Analogies in Pretrained Language Models 28 Nov 2022 · 2 repositories · arXiv:2211.15268
-
ESIE-BERT: Enriching Sub-words Information Explicitly with BERT for Joint Intent Classification and SlotFilling 27 Nov 2022 · 0 repositories · arXiv:2211.14829
-
Understanding BLOOM: An empirical study on diverse NLP tasks 27 Nov 2022 · 0 repositories · arXiv:2211.14865
-
An Analysis of Social Biases Present in BERT Variants Across Multiple Languages 25 Nov 2022 · 1 repository · arXiv:2211.14402Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Finetuning BERT on Partially Annotated NER Corpora 25 Nov 2022 · 1 repository · arXiv:2211.14360