Methods › Natural Language Processing › Subword Segmentation › WordPiece › Papers, page 24
WordPiece
Papers archive 2025-07-28
archive papers tagged: 7,063 · with a code link: 2,910 · where Syntology ran a sample: 650 (529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (650 of 7,063 tagged: 529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument)
Page 24 of 71: papers 2,301 to 2,400 of 7,063, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
DACBERT: Leveraging Dependency Agreement for Cost-Efficient Bert Pretraining 8 Nov 2023 · 0 repositories · arXiv:2311.04799
-
Deep Learning Brasil at ABSAPT 2022: Portuguese Transformer Ensemble Approaches 8 Nov 2023 · 1 repository · arXiv:2311.05051
-
DeepLearningBrasil@LT-EDI-2023: Exploring Deep Learning Techniques for Detecting Depression in Social Media Text 8 Nov 2023 · 1 repository · arXiv:2311.05047
-
Determination of toxic comments and unintended model bias minimization using Deep learning approach 8 Nov 2023 · 1 repository · arXiv:2311.04789
-
Pre-training LLMs using human-like development data corpus 8 Nov 2023 · 0 repositories · arXiv:2311.04666
-
Enhancing LLM Intelligence with ARM-RAG: Auxiliary Rationale Memory for Retrieval Augmented Generation 7 Nov 2023 · 0 repositories · arXiv:2311.04177
-
Modelling Sentiment Analysis: LLMs and data augmentation techniques 7 Nov 2023 · 1 repository · arXiv:2311.04139
-
Personality Style Recognition via Machine Learning: Identifying Anaclitic and Introjective Personality Styles from Patients' Speech 7 Nov 2023 · 0 repositories · arXiv:2311.04088
-
Accumulating Word Representations in Multi-level Context Integration for ERC Task 6 Nov 2023 · 1 repository
-
Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch 6 Nov 2023 · 3 repositories · arXiv:2311.03099Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
AI-TA: Towards an Intelligent Question-Answer Teaching Assistant using Open-Source LLMs 5 Nov 2023 · 1 repository · arXiv:2311.02775
-
You Only Forward Once: Prediction and Rationalization in A Single Forward Pass 4 Nov 2023 · 0 repositories · arXiv:2311.02344
-
Data-Free Distillation of Language Model by Text-to-Text Transfer 3 Nov 2023 · 0 repositories · arXiv:2311.01689
-
Simplifying Transformer Blocks 3 Nov 2023 · 1 repository · arXiv:2311.01906
-
Measuring Five Accountable Talk Moves to Improve Instruction at Scale 2 Nov 2023 · 0 repositories · arXiv:2311.10749
-
An Improved Transformer-based Model for Detecting Phishing, Spam, and Ham: A Large Language Model Approach 1 Nov 2023 · 0 repositories · arXiv:2311.04913
-
Entity Alignment Method of Science and Technology Patent based on Graph Convolution Network and Information Fusion 1 Nov 2023 · 0 repositories · arXiv:2311.00300
-
Syntactic Inductive Bias in Transformer Language Models: Especially Helpful for Low-Resource Languages? 1 Nov 2023 · 1 repository · arXiv:2311.00268
-
BERTwich: Extending BERT's Capabilities to Model Dialectal and Noisy Text 31 Oct 2023 · 0 repositories · arXiv:2311.00116
-
Breaking the Token Barrier: Chunking and Convolution for Efficient Long Text Classification with BERT 31 Oct 2023 · 0 repositories · arXiv:2310.20558
-
EELBERT: Tiny Models through Dynamic Embeddings 31 Oct 2023 · 0 repositories · arXiv:2310.20144
-
FA Team at the NTCIR-17 UFO Task 31 Oct 2023 · 0 repositories · arXiv:2310.20322
-
GAR-meets-RAG Paradigm for Zero-Shot Information Retrieval 31 Oct 2023 · 0 repositories · arXiv:2310.20158
-
Increasing The Performance of Cognitively Inspired Data-Efficient Language Models via Implicit Structure Building 31 Oct 2023 · 1 repository · arXiv:2310.20589
-
BTRec: BERT-Based Trajectory Recommendation for Personalized Tours 30 Oct 2023 · 1 repository · arXiv:2310.19886
-
Jina Embeddings 2: 8192-Token General-Purpose Text Embeddings for Long Documents 30 Oct 2023 · 2 repositories · arXiv:2310.19923Syntology 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Partial Tensorized Transformers for Natural Language Processing 30 Oct 2023 · 0 repositories · arXiv:2310.20077
-
Split-NER: Named Entity Recognition via Two Question-Answering-based Classifications 30 Oct 2023 · 1 repository · arXiv:2310.19942Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
From Chatbots to PhishBots? -- Preventing Phishing scams created using ChatGPT, Google Bard and Claude 29 Oct 2023 · 0 repositories · arXiv:2310.19181
-
Prompt-Engineering and Transformer-based Question Generation and Evaluation 29 Oct 2023 · 0 repositories · arXiv:2310.18867
-
Retrofitting Light-weight Language Models for Emotions using Supervised Contrastive Learning 29 Oct 2023 · 0 repositories · arXiv:2310.18930
-
Style Description based Text-to-Speech with Conditional Prosodic Layer Normalization based Diffusion GAN 27 Oct 2023 · 0 repositories · arXiv:2310.18169
-
Arabic Fine-Grained Entity Recognition 26 Oct 2023 · 0 repositories · arXiv:2310.17333
-
FedPEAT: Convergence of Federated Learning, Parameter-Efficient Fine Tuning, and Emulator Assisted Tuning for Artificial Intelligence Foundation Models with Mobile Edge Computing 26 Oct 2023 · 0 repositories · arXiv:2310.17491
-
Harnessing GPT-3.5-turbo for Rhetorical Role Prediction in Legal Cases 26 Oct 2023 · 0 repositories · arXiv:2310.17413
-
Sliceformer: Make Multi-head Attention as Simple as Sorting in Discriminative Tasks 26 Oct 2023 · 1 repository · arXiv:2310.17683
-
torchdistill Meets Hugging Face Libraries for Reproducible, Coding-Free Deep Learning Studies: A Case Study on NLP 26 Oct 2023 · 1 repository · arXiv:2310.17644
-
ZeroQuant-HERO: Hardware-Enhanced Robust Optimized Post-Training Quantization Framework for W8A8 Transformers 26 Oct 2023 · 0 repositories · arXiv:2310.17723
-
Enhancing Document Information Analysis with Multi-Task Pre-training: A Robust Approach for Information Extraction in Visually-Rich Documents 25 Oct 2023 · 0 repositories · arXiv:2310.16527
-
LLM-FP4: 4-Bit Floating-Point Quantized Transformers 25 Oct 2023 · 1 repository · arXiv:2310.16836Syntology official (archive's flag): 3 ran · 3 ran (of which 3 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; every one of the 3 samples that ran constructed an object rather than computing a result (of 3 harvested samples)
-
𝕍𝔻-𝔾ℝ: Boosting 𝕍isual 𝔻ialog with Cascaded Spatial-Temporal Multi-Modal 𝔾ℝaphs 25 Oct 2023 · 0 repositories · arXiv:2310.16590
-
URL-BERT: Training Webpage Representations via Social Media Engagements 25 Oct 2023 · 0 repositories · arXiv:2310.16303
-
Attention-Enhancing Backdoor Attacks Against BERT-based Models 23 Oct 2023 · 0 repositories · arXiv:2310.14480
-
Health Disparities through Generative AI Models: A Comparison Study Using A Domain Specific large language model 23 Oct 2023 · 0 repositories · arXiv:2310.18355
-
Unleashing the potential of prompt engineering for large language models 23 Oct 2023 · 0 repositories · arXiv:2310.14735
-
ITEm: Unsupervised Image-Text Embedding Learning for eCommerce 22 Oct 2023 · 0 repositories · arXiv:2311.02084
-
Towards Harmful Erotic Content Detection through Coreference-Driven Contextual Analysis 22 Oct 2023 · 0 repositories · arXiv:2310.14325
-
COVIDFakeExplainer: An Explainable Machine Learning based Web Application for Detecting COVID-19 Fake News 21 Oct 2023 · 0 repositories · arXiv:2310.13890
-
LLM-Prop: Predicting Physical And Electronic Properties Of Crystalline Solids From Their Text Descriptions 21 Oct 2023 · 1 repository · arXiv:2310.14029Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples)
-
Anomaly Detection of Command Shell Sessions based on DistilBERT: Unsupervised and Supervised Approaches 20 Oct 2023 · 0 repositories · arXiv:2310.13247
-
Exploring the Impact of Corpus Diversity on Financial Pretrained Language Models 20 Oct 2023 · 1 repository · arXiv:2310.13312
-
FABULA: Intelligence Report Generation Using Retrieval-Augmented Narrative Construction 20 Oct 2023 · 0 repositories · arXiv:2310.13848
-
Multi-level Contrastive Learning for Script-based Character Understanding 20 Oct 2023 · 1 repository · arXiv:2310.13231
-
MedAI Dialog Corpus (MEDIC): Zero-Shot Classification of Doctor and AI Responses in Health Consultations 19 Oct 2023 · 0 repositories · arXiv:2310.12489
-
ExtractGPT: Exploring the Potential of Large Language Models for Product Attribute Value Extraction 19 Oct 2023 · 1 repository · arXiv:2310.12537
-
Towards Robust Pruning: An Adaptive Knowledge-Retention Pruning Strategy for Language Models 19 Oct 2023 · 0 repositories · arXiv:2310.13191
-
Transformer-based Entity Legal Form Classification 19 Oct 2023 · 1 repository · arXiv:2310.12766
-
Field-testing items using artificial intelligence: Natural language processing with transformers 18 Oct 2023 · 0 repositories · arXiv:2310.11655
-
Improving Long Document Topic Segmentation Models With Enhanced Coherence Modeling 18 Oct 2023 · 1 repository · arXiv:2310.11772
-
Disentangling the Linguistic Competence of Privacy-Preserving BERT 17 Oct 2023 · 0 repositories · arXiv:2310.11363
-
Entity Matching using Large Language Models 17 Oct 2023 · 1 repository · arXiv:2310.11244
-
MASON-NLP at eRisk 2023: Deep Learning-Based Detection of Depression Symptoms from Social Media Texts 17 Oct 2023 · 0 repositories · arXiv:2310.10941
-
Neural Attention: Enhancing QKV Calculation in Self-Attention Mechanism with Neural Networks 17 Oct 2023 · 1 repository · arXiv:2310.11398
-
Fine-tuning ChatGPT for Automatic Scoring 16 Oct 2023 · 0 repositories · arXiv:2310.10072
-
Investigating Bias in Multilingual Language Models: Cross-Lingual Transfer of Debiasing Techniques 16 Oct 2023 · 1 repository · arXiv:2310.10310
-
Learning to Rank Context for Named Entity Recognition Using a Synthetic Dataset 16 Oct 2023 · 1 repository · arXiv:2310.10118
-
Personalization of CTC-based End-to-End Speech Recognition Using Pronunciation-Driven Subword Tokenization 16 Oct 2023 · 0 repositories · arXiv:2310.09988
-
Domain-Specific Language Model Post-Training for Indonesian Financial NLP 15 Oct 2023 · 1 repository · arXiv:2310.09736
-
DPZero: Private Fine-Tuning of Language Models without Backpropagation 14 Oct 2023 · 1 repository · arXiv:2310.09639Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 2 where Syntology's instrument failed) · 9 unverified (of 17 harvested samples) · 3 pointer-only (licence)
-
Leveraging Generative AI: Improving Software Metadata Classification with Generated Code-Comment Pairs 14 Oct 2023 · 0 repositories · arXiv:2311.03365
-
Enhancing BERT-Based Visual Question Answering through Keyword-Driven Sentence Selection 13 Oct 2023 · 0 repositories · arXiv:2310.09432
-
Analyzing Textual Data for Fatality Classification in Afghanistan's Armed Conflicts: A BERT Approach 12 Oct 2023 · 0 repositories · arXiv:2310.08653
-
Detection and prediction of clopidogrel treatment failures using longitudinal structured electronic health records 12 Oct 2023 · 0 repositories · arXiv:2310.08757
-
Evaluating The Effectiveness of Capsule Neural Network in Toxic Comment Classification using Pre-trained BERT Embeddings 12 Oct 2023 · 1 repository
-
LEMON: Lossless model expansion 12 Oct 2023 · 1 repository · arXiv:2310.07999Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
LLM-augmented Preference Learning from Natural Language 12 Oct 2023 · 0 repositories · arXiv:2310.08523
-
The Uncertainty-based Retrieval Framework for Ancient Chinese CWS and POS 12 Oct 2023 · 1 repository · arXiv:2310.08496
-
Fast-ELECTRA for Efficient Pre-training 11 Oct 2023 · 0 repositories · arXiv:2310.07347
-
Jaeger: A Concatenation-Based Multi-Transformer VQA Model 11 Oct 2023 · 0 repositories · arXiv:2310.07091
-
Toward Understanding BERT-Like Pre-Training for DNA Foundation Models 11 Oct 2023 · 0 repositories · arXiv:2310.07644
-
A Comparative Study of Transformer-based Neural Text Representation Techniques on Bug Triaging 10 Oct 2023 · 0 repositories · arXiv:2310.06913
-
GeoLLM: Extracting Geospatial Knowledge from Large Language Models 10 Oct 2023 · 1 repository · arXiv:2310.06213Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
GPT-4 as an Agronomist Assistant? Answering Agriculture Exams Using Large Language Models 10 Oct 2023 · 0 repositories · arXiv:2310.06225
-
Large Language Models for Propaganda Detection 10 Oct 2023 · 2 repositories · arXiv:2310.06422
-
Auditing Gender Analyzers on Text Data 9 Oct 2023 · 0 repositories · arXiv:2310.06061
-
Cabbage Sweeter than Cake? Analysing the Potential of Large Language Models for Learning Conceptual Spaces 9 Oct 2023 · 0 repositories · arXiv:2310.05481
-
Foundation Models Meet Visualizations: Challenges and Opportunities 9 Oct 2023 · 0 repositories · arXiv:2310.05771
-
Transformer Fusion with Optimal Transport 9 Oct 2023 · 1 repository · arXiv:2310.05719Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 7 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Breaking Down Word Semantics from Pre-trained Language Models through Layer-wise Dimension Selection 8 Oct 2023 · 0 repositories · arXiv:2310.05115
-
Enhancing Pre-Trained Language Models with Sentence Position Embeddings for Rhetorical Roles Recognition in Legal Opinions 8 Oct 2023 · 0 repositories · arXiv:2310.05276
-
LLM4VV: Developing LLM-Driven Testsuite for Compiler Validation 8 Oct 2023 · 1 repository · arXiv:2310.04963
-
RAC-BERT: Character Radical Enhanced BERT for Ancient Chinese 8 Oct 2023 · 0 repositories
-
A Process for Topic Modelling Via Word Embeddings 6 Oct 2023 · 0 repositories · arXiv:2312.03705
-
Automatic Aspect Extraction from Scientific Texts 6 Oct 2023 · 1 repository · arXiv:2310.04074
-
Quantized Transformer Language Model Implementations on Edge Devices 6 Oct 2023 · 0 repositories · arXiv:2310.03971
-
Segmented Harmonic Loss: Handling Class-Imbalanced Multi-Label Clinical Data for Medical Coding with Large Language Models 6 Oct 2023 · 0 repositories · arXiv:2310.04595
-
COVID-19 South African Vaccine Hesitancy Models Show Boost in Performance Upon Fine-Tuning on M-pox Tweets 4 Oct 2023 · 0 repositories · arXiv:2310.04453
-
Memoria: Resolving Fateful Forgetting Problem through Human-Inspired Memory Architecture 4 Oct 2023 · 1 repository · arXiv:2310.03052Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 9 unverified (of 10 harvested samples)
-
Retrieval-augmented Generation to Improve Math Question-Answering: Trade-offs Between Groundedness and Human Preference 4 Oct 2023 · 2 repositories · arXiv:2310.03184Syntology official (archive's flag): 7 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 8 unverified (of 20 harvested samples)
-
Harnessing Pre-Trained Sentence Transformers for Offensive Language Detection in Indian Languages 3 Oct 2023 · 0 repositories · arXiv:2310.02249