Methods › Natural Language Processing › Subword Segmentation › WordPiece › Papers, page 26
WordPiece
Papers archive 2025-07-28
archive papers tagged: 7,063 · with a code link: 2,910 · where Syntology ran a sample: 650 (529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (650 of 7,063 tagged: 529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument)
Page 26 of 71: papers 2,501 to 2,600 of 7,063, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Optimizing Multi-Class Text Classification: A Diverse Stacking Ensemble Framework Utilizing Transformers 19 Aug 2023 · 0 repositories · arXiv:2308.11519
-
Learning Representations on Logs for AIOps 18 Aug 2023 · 1 repository · arXiv:2308.11526
-
Predictive Authoring for Brazilian Portuguese Augmentative and Alternative Communication 18 Aug 2023 · 1 repository · arXiv:2308.09497
-
A Comparative Study of Text Embedding Models for Semantic Text Similarity in Bug Reports 17 Aug 2023 · 1 repository · arXiv:2308.09193
-
BIOptimus: Pre-training an Optimal Biomedical Language Model with Curriculum Learning for Named Entity Recognition 16 Aug 2023 · 1 repository · arXiv:2308.08625
-
"Beware of deception": Detecting Half-Truth and Debunking it through Controlled Claim Editing 15 Aug 2023 · 0 repositories · arXiv:2308.07973
-
DS4DH at #SMM4H 2023: Zero-Shot Adverse Drug Events Normalization using Sentence Transformers and Reciprocal-Rank Fusion 15 Aug 2023 · 0 repositories · arXiv:2308.12877
-
Finding Stakeholder-Material Information from 10-K Reports using Fine-Tuned BERT and LSTM Models 15 Aug 2023 · 0 repositories · arXiv:2308.07522
-
MultiSChuBERT: Effective Multimodal Fusion for Scholarly Document Quality Prediction 15 Aug 2023 · 0 repositories · arXiv:2308.07971
-
SPM: Structured Pretraining and Matching Architectures for Relevance Modeling in Meituan Search 15 Aug 2023 · 0 repositories · arXiv:2308.07711
-
Leveraging Codebook Knowledge with NLI and ChatGPT for Zero-Shot Political Relation Classification 15 Aug 2023 · 1 repository · arXiv:2308.07876
-
Ternary Singular Value Decomposition as a Better Parameterized Form in Linear Mapping 15 Aug 2023 · 1 repository · arXiv:2308.07641Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
An Ensemble Approach to Question Classification: Integrating Electra Transformer, GloVe, and LSTM 13 Aug 2023 · 0 repositories · arXiv:2308.06828
-
Improving Face Recognition from Caption Supervision with Multi-Granular Contextual Feature Aggregation 13 Aug 2023 · 0 repositories · arXiv:2308.06866
-
Enhancing Phenotype Recognition in Clinical Notes Using Large Language Models: PhenoBCBERT and PhenoGPT 11 Aug 2023 · 1 repository · arXiv:2308.06294
-
Identification of the Relevance of Comments in Codes Using Bag of Words and Transformer Based Models 11 Aug 2023 · 1 repository · arXiv:2308.06144
-
Task Conditioned BERT for Joint Intent Detection and Slot-filling 11 Aug 2023 · 0 repositories · arXiv:2308.06165
-
Bringing order into the realm of Transformer-based language models for artificial intelligence and law 10 Aug 2023 · 0 repositories · arXiv:2308.05502
-
Exploring Machine Learning and Transformer-based Approaches for Deceptive Text Classification: A Comparative Analysis 10 Aug 2023 · 0 repositories · arXiv:2308.05476
-
MetRoBERTa: Leveraging Traditional Customer Relationship Management Data to Develop a Transit-Topic-Aware Language Model 9 Aug 2023 · 0 repositories · arXiv:2308.05012
-
Performance Analysis of Transformer Based Models (BERT, ALBERT and RoBERTa) in Fake News Detection 9 Aug 2023 · 1 repository · arXiv:2308.04950
-
A Cross-Domain Evaluation of Approaches for Causal Knowledge Extraction 7 Aug 2023 · 1 repository · arXiv:2308.03891
-
Analysis of the Evolution of Advanced Transformer-Based Language Models: Experiments on Opinion Mining 7 Aug 2023 · 1 repository · arXiv:2308.03235
-
Detecting Spells in Fantasy Literature with a Transformer Based Artificial Intelligence 7 Aug 2023 · 0 repositories · arXiv:2308.03660
-
Trusting Language Models in Education 7 Aug 2023 · 0 repositories · arXiv:2308.03866
-
Training BERT Models to Carry Over a Coding System Developed on One Corpus to Another 7 Aug 2023 · 0 repositories · arXiv:2308.03742
-
End-to-End Query Term Weighting 6 Aug 2023 · 1 repository
-
DaMSTF: Domain Adversarial Learning Enhanced Meta Self-Training for Domain Adaptation 5 Aug 2023 · 0 repositories · arXiv:2308.02753Syntology 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 7 harvested samples)
-
An APT Event Extraction Method Based on BERT-BiGRU-CRF for APT Attack Detection 4 Aug 2023 · 0 repositories
-
Explaining Relation Classification Models with Semantic Extents 4 Aug 2023 · 2 repositories · arXiv:2308.02193
-
Meta-Tsallis-Entropy Minimization: A New Self-Training Approach for Domain Adaptation on Text Classification 4 Aug 2023 · 0 repositories · arXiv:2308.02746
-
Baby's CoThought: Leveraging Large Language Models for Enhanced Reasoning in Compact Models 3 Aug 2023 · 1 repository · arXiv:2308.01684
-
Comparing scalable strategies for generating numerical perspectives 3 Aug 2023 · 0 repositories · arXiv:2308.01535
-
Food Classification using Joint Representation of Visual and Textual Data 3 Aug 2023 · 0 repositories · arXiv:2308.02562
-
Local Large Language Models for Complex Structured Medical Tasks 3 Aug 2023 · 1 repository · arXiv:2308.01727
-
Bio+Clinical BERT, BERT Base, and CNN Performance Comparison for Predicting Drug-Review Satisfaction 2 Aug 2023 · 0 repositories · arXiv:2308.03782
-
Contextual Emotion Recognition Using Transformer-Based Models 2 Aug 2023 · 1 repository
-
Towards Better Query Classification with Multi-Expert Knowledge Condensation in JD Ads Search 2 Aug 2023 · 0 repositories · arXiv:2308.01098
-
Retrieval Augmented Generation and Representative Vector Summarization for large unstructured textual data in Medical Education 1 Aug 2023 · 2 repositories · arXiv:2308.00479
-
Self-Supervised Contrastive BERT Fine-tuning for Fusion-based Reviewed-Item Retrieval 1 Aug 2023 · 2 repositories · arXiv:2308.00762
-
Classifying multilingual party manifestos: Domain transfer across country, time, and genre 31 Jul 2023 · 1 repository · arXiv:2307.16511
-
Noisy Self-Training with Data Augmentations for Offensive and Hate Speech Detection Tasks 31 Jul 2023 · 1 repository · arXiv:2307.16609
-
VacancySBERT: the approach for representation of titles and skills for semantic similarity search in the recruitment domain 31 Jul 2023 · 1 repository · arXiv:2307.16638
-
LaFiCMIL: Rethinking Large File Classification from the Perspective of Correlated Multiple Instance Learning 30 Jul 2023 · 0 repositories · arXiv:2308.01413
-
Tutorials on Stance Detection using Pre-trained Language Models: Fine-tuning BERT and Prompting Large Language Models 28 Jul 2023 · 0 repositories · arXiv:2307.15331
-
Evaluating Generative Models for Graph-to-Text Generation 27 Jul 2023 · 1 repository · arXiv:2307.14712
-
New Interaction Paradigm for Complex EDA Software Leveraging GPT 27 Jul 2023 · 1 repository · arXiv:2307.14740Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample) · 1 pointer-only (licence)
-
TextManiA: Enriching Visual Feature by Text-driven Manifold Augmentation 27 Jul 2023 · 0 repositories · arXiv:2307.14611
-
Comparative Analysis of Libraries for the Sentimental Analysis 26 Jul 2023 · 0 repositories · arXiv:2307.14311
-
Developing and Evaluating Tiny to Medium-Sized Turkish BERT Models 26 Jul 2023 · 0 repositories · arXiv:2307.14134
-
DPBERT: Efficient Inference for BERT based on Dynamic Planning 26 Jul 2023 · 0 repositories · arXiv:2308.00108
-
A Hybrid Machine Learning Model for Classifying Gene Mutations in Cancer using LSTM, BiLSTM, CNN, GRU, and GloVe 24 Jul 2023 · 0 repositories · arXiv:2307.14361
-
How Does Naming Affect LLMs on Code Analysis Tasks? 24 Jul 2023 · 0 repositories · arXiv:2307.12488
-
Gradient-Based Word Substitution for Obstinate Adversarial Examples Generation in Language Models 24 Jul 2023 · 0 repositories · arXiv:2307.12507
-
Transformer-based Joint Source Channel Coding for Textual Semantic Communication 23 Jul 2023 · 0 repositories · arXiv:2307.12266
-
Identifying Misinformation on YouTube through Transcript Contextual Analysis with Transformer Models 22 Jul 2023 · 1 repository · arXiv:2307.12155
-
DEFTri: A Few-Shot Label Fused Contextual Representation Learning For Product Defect Triage in e-Commerce 21 Jul 2023 · 0 repositories · arXiv:2307.11344
-
The Looming Threat of Fake and LLM-generated LinkedIn Profiles: Challenges and Opportunities for Detection and Prevention 21 Jul 2023 · 1 repository · arXiv:2307.11864
-
An In-Depth Evaluation of Federated Learning on Biomedical Natural Language Processing 20 Jul 2023 · 2 repositories · arXiv:2307.11254
-
Generative Language Models on Nucleotide Sequences of Human Genes 20 Jul 2023 · 1 repository · arXiv:2307.10634
-
Mood Classification of Bangla Songs Based on Lyrics 19 Jul 2023 · 0 repositories · arXiv:2307.10314
-
SPRINT: A Unified Toolkit for Evaluating and Demystifying Zero-shot Neural Sparse Retrieval 19 Jul 2023 · 1 repository · arXiv:2307.10488
-
Analyzing sports commentary in order to automatically recognize events and extract insights 18 Jul 2023 · 2 repositories · arXiv:2307.10303
-
Application of BERT in Wind Power Forecasting-Teletraan's Solution in Baidu KDD Cup 2022 18 Jul 2023 · 1 repository · arXiv:2307.09248
-
Automated Ableism: An Exploration of Explicit Disability Biases in Sentiment and Toxicity Analysis Models 18 Jul 2023 · 0 repositories · arXiv:2307.09209
-
Can Model Fusing Help Transformers in Long Document Classification? An Empirical Study 18 Jul 2023 · 1 repository · arXiv:2307.09532
-
KATIE: A System for Key Attributes Identification in Product Knowledge Graph Construction 18 Jul 2023 · 0 repositories
-
Cross-Lingual NER for Financial Transaction Data in Low-Resource Languages 16 Jul 2023 · 0 repositories · arXiv:2307.08714
-
Recognition of Mental Adjectives in An Efficient and Automatic Style 16 Jul 2023 · 0 repositories · arXiv:2307.11767
-
Do not Mask Randomly: Effective Domain-adaptive Pre-training by Masking In-domain Keywords 14 Jul 2023 · 0 repositories · arXiv:2307.07160
-
Improving BERT with Hybrid Pooling Network and Drop Mask 14 Jul 2023 · 0 repositories · arXiv:2307.07258
-
Sensi-BERT: Towards Sensitivity Driven Fine-Tuning for Parameter-Efficient BERT 14 Jul 2023 · 0 repositories · arXiv:2307.11764
-
Towards spoken dialect identification of Irish 14 Jul 2023 · 0 repositories · arXiv:2307.07436
-
TVPR: Text-to-Video Person Retrieval and a New Benchmark 14 Jul 2023 · 0 repositories · arXiv:2307.07184
-
Convolutional Neural Networks for Sentiment Analysis on Weibo Data: A Natural Language Processing Approach 13 Jul 2023 · 0 repositories · arXiv:2307.06540
-
SecureFalcon: Are We There Yet in Automated Software Vulnerability Detection with LLMs? 13 Jul 2023 · 0 repositories · arXiv:2307.06616
-
Tackling Fake News in Bengali: Unraveling the Impact of Summarization vs. Augmentation on Pre-trained Language Models 13 Jul 2023 · 1 repository · arXiv:2307.06979
-
Retrieval Augmented Generation using Engineering Design Knowledge 13 Jul 2023 · 2 repositories · arXiv:2307.06985
-
Detecting the Presence of COVID-19 Vaccination Hesitancy from South African Twitter Data Using Machine Learning 12 Jul 2023 · 0 repositories · arXiv:2307.15072
-
No Train No Gain: Revisiting Efficient Training Algorithms For Transformer-based Language Models 12 Jul 2023 · 1 repository · arXiv:2307.06440Syntology official (archive's flag): 9 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 3 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
Prompt Generate Train (PGT): Few-shot Domain Adaption of Retrieval Augmented Generation Models for Open Book Question-Answering 12 Jul 2023 · 0 repositories · arXiv:2307.05915
-
Vacaspati: A Diverse Corpus of Bangla Literature 11 Jul 2023 · 0 repositories · arXiv:2307.05083
-
ChatGPT for Digital Forensic Investigation: The Good, The Bad, and The Unknown 10 Jul 2023 · 1 repository · arXiv:2307.10195
-
Is ChatGPT a Good Personality Recognizer? A Preliminary Study 8 Jul 2023 · 0 repositories · arXiv:2307.03952
-
Goal-Conditioned Predictive Coding for Offline Reinforcement Learning 7 Jul 2023 · 0 repositories · arXiv:2307.03406
-
TRAQ: Trustworthy Retrieval Augmented Question Answering via Conformal Prediction 7 Jul 2023 · 1 repository · arXiv:2307.04642
-
A Novel Site-Agnostic Multimodal Deep Learning Model to Identify Pro-Eating Disorder Content on Social Media 6 Jul 2023 · 0 repositories · arXiv:2307.06775
-
Can ChatGPT's Responses Boost Traditional Natural Language Processing? 6 Jul 2023 · 1 repository · arXiv:2307.04648
-
Text Alignment Is An Efficient Unified Model for Massive NLP Tasks 6 Jul 2023 · 1 repository · arXiv:2307.02729
-
UIT-Saviors at MEDVQA-GI 2023: Improving Multimodal Learning with Image Enhancement for Gastrointestinal Visual Question Answering 6 Jul 2023 · 0 repositories · arXiv:2307.02783
-
CAME: Confidence-guided Adaptive Memory Efficient Optimization 5 Jul 2023 · 2 repositories · arXiv:2307.02047Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; the one sample that ran constructed an object rather than computing a result (of 2 harvested samples) · 1 pointer-only (licence)
-
Emoji Prediction in Tweets using BERT 5 Jul 2023 · 1 repository · arXiv:2307.02054
-
Evaluating the Effectiveness of Large Language Models in Representing Textual Descriptions of Geometry and Spatial Relations 5 Jul 2023 · 0 repositories · arXiv:2307.03678
-
Named Entity Inclusion in Abstractive Text Summarization 5 Jul 2023 · 0 repositories · arXiv:2307.02570
-
KDSTM: Neural Semi-supervised Topic Modeling with Knowledge Distillation 4 Jul 2023 · 0 repositories · arXiv:2307.01878Syntology 6 ran (of which 1 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples)
-
ALBERTI, a Multilingual Domain Specific Language Model for Poetry Analysis 3 Jul 2023 · 0 repositories · arXiv:2307.01387
-
Improving Language Plasticity via Pretraining with Active Forgetting 3 Jul 2023 · 1 repository · arXiv:2307.01163
-
Interpretability and Transparency-Driven Detection and Transformation of Textual Adversarial Examples (IT-DT) 3 Jul 2023 · 0 repositories · arXiv:2307.01225
-
How far is Language Model from 100% Few-shot Named Entity Recognition in Medical Domain 1 Jul 2023 · 1 repository · arXiv:2307.00186
-
Ticket-BERT: Labeling Incident Management Tickets with Language Models 30 Jun 2023 · 0 repositories · arXiv:2307.00108