Methods › Natural Language Processing › Subword Segmentation › WordPiece › Papers, page 23
WordPiece
Papers archive 2025-07-28
archive papers tagged: 7,063 · with a code link: 2,910 · where Syntology ran a sample: 650 (529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (650 of 7,063 tagged: 529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument)
Page 23 of 71: papers 2,201 to 2,300 of 7,063, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Exploring Automatic Text Simplification of German Narrative Documents 15 Dec 2023 · 1 repository · arXiv:2312.09907
-
No-Skim: Towards Efficiency Robustness Evaluation on Skimming-based Language Models 15 Dec 2023 · 0 repositories · arXiv:2312.09494
-
Dissecting vocabulary biases datasets through statistical testing and automated data augmentation for artifact mitigation in Natural Language Inference 14 Dec 2023 · 1 repository · arXiv:2312.08747
-
Mono3DVG: 3D Visual Grounding in Monocular Images 13 Dec 2023 · 1 repository · arXiv:2312.08022Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Prompt Engineering-assisted Malware Dynamic Analysis Using GPT-4 13 Dec 2023 · 1 repository · arXiv:2312.08317
-
Harnessing Retrieval-Augmented Generation (RAG) for Uncovering Knowledge Gaps 12 Dec 2023 · 1 repository · arXiv:2312.07796
-
SEOpinion: Summarization and Exploration Opinion of E-Commerce Websites 12 Dec 2023 · 0 repositories · arXiv:2312.14171
-
Towards Equipping Transformer with the Ability of Systematic Compositionality 12 Dec 2023 · 1 repository · arXiv:2312.07280
-
A Natural Language Processing-Based Classification and Mode-Based Ranking of Musculoskeletal Disorder Risk Factors 12 Dec 2023 · 0 repositories · arXiv:2312.11517
-
Contrastive News and Social Media Linking using BERT for Articles and Tweets across Dual Platforms 11 Dec 2023 · 0 repositories · arXiv:2312.07599
-
Revisiting the Role of Label Smoothing in Enhanced Text Sentiment Classification 11 Dec 2023 · 0 repositories · arXiv:2312.06522
-
Survey on Foundation Models for Prognostics and Health Management in Industrial Cyber-Physical Systems 11 Dec 2023 · 0 repositories · arXiv:2312.06261
-
Textual Prompt Guided Image Restoration 11 Dec 2023 · 1 repository · arXiv:2312.06162
-
Where exactly does contextualization in a PLM happen? 11 Dec 2023 · 0 repositories · arXiv:2312.06514
-
Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs 10 Dec 2023 · 0 repositories · arXiv:2312.05934
-
FP8-BERT: Post-Training Quantization for Transformer 10 Dec 2023 · 0 repositories · arXiv:2312.05725
-
GAMC: An Unsupervised Method for Fake News Detection using Graph Autoencoder with Masking 10 Dec 2023 · 1 repository · arXiv:2312.05739
-
A Review of Hybrid and Ensemble in Deep Learning for Natural Language Processing 9 Dec 2023 · 0 repositories · arXiv:2312.05589
-
Context Tuning for Retrieval Augmented Generation 9 Dec 2023 · 0 repositories · arXiv:2312.05708
-
Enhanced E-Commerce Attribute Extraction: Innovating with Decorative Relation Correction and LLAMA 2.0-Based Annotation 9 Dec 2023 · 0 repositories · arXiv:2312.06684
-
Enhancing Medical Specialty Assignment to Patients using NLP Techniques 9 Dec 2023 · 0 repositories · arXiv:2312.05585
-
Hate Speech and Offensive Content Detection in Indo-Aryan Languages: A Battle of LSTM and Transformers 9 Dec 2023 · 0 repositories · arXiv:2312.05671
-
Labrador: Exploring the Limits of Masked Language Modeling for Laboratory Data 9 Dec 2023 · 1 repository · arXiv:2312.11502Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Sim-GPT: Text Similarity via GPT Annotated Data 9 Dec 2023 · 1 repository · arXiv:2312.05603
-
Teamwork Dimensions Classification Using BERT 9 Dec 2023 · 0 repositories · arXiv:2312.05483
-
Cross-BERT for Point Cloud Pretraining 8 Dec 2023 · 0 repositories · arXiv:2312.04891
-
Illicit Darkweb Classification via Natural-language Processing: Classifying Illicit Content of Webpages based on Textual Information 8 Dec 2023 · 0 repositories · arXiv:2312.04944
-
INSPECT: Intrinsic and Systematic Probing Evaluation for Code Transformers 8 Dec 2023 · 1 repository · arXiv:2312.05092
-
PaperQA: Retrieval-Augmented Generative Agent for Scientific Research 8 Dec 2023 · 0 repositories · arXiv:2312.07559
-
Fortify the Shortest Stave in Attention: Enhancing Context Awareness of Large Language Models for Effective Tool Use 7 Dec 2023 · 1 repository · arXiv:2312.04455Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Language Model Knowledge Distillation for Efficient Question Answering in Spanish 7 Dec 2023 · 1 repository · arXiv:2312.04193Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
A Text-to-Text Model for Multilingual Offensive Language Identification 6 Dec 2023 · 0 repositories · arXiv:2312.03379
-
Automatic Transcription of Handwritten Old Occitan Language 6 Dec 2023 · 1 repository
-
Corporate Bankruptcy Prediction with Domain-Adapted BERT 6 Dec 2023 · 0 repositories · arXiv:2312.03194
-
Exploring the Reversal Curse and Other Deductive Logical Reasoning in BERT and GPT-Based Large Language Models 6 Dec 2023 · 1 repository · arXiv:2312.03633
-
A Survey on Large Language Model (LLM) Security and Privacy: The Good, the Bad, and the Ugly 4 Dec 2023 · 0 repositories · arXiv:2312.02003
-
Expand BERT Representation with Visual Information via Grounded Language Learning with Multimodal Partial Alignment 4 Dec 2023 · 0 repositories · arXiv:2312.01592
-
Prompting Disentangled Embeddings for Knowledge Graph Completion with Pre-trained Language Model 4 Dec 2023 · 1 repository · arXiv:2312.01837
-
ArabIcros: AI-Powered Arabic Crossword Puzzle Generation for Educational Applications 3 Dec 2023 · 0 repositories · arXiv:2312.01339
-
NLEBench+NorGLM: A Comprehensive Empirical Analysis and Benchmark Dataset for Generative Language Models in Norwegian 3 Dec 2023 · 1 repository · arXiv:2312.01314Syntology official: harvested, nothing ran · 0 ran · 4 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
On Significance of Subword tokenization for Low Resource and Efficient Named Entity Recognition: A case study in Marathi 3 Dec 2023 · 0 repositories · arXiv:2312.01306
-
A ripple in time: a discontinuity in American history 2 Dec 2023 · 1 repository · arXiv:2312.01185
-
Knowledge Graph Enhanced Aspect-Level Sentiment Analysis 2 Dec 2023 · 0 repositories · arXiv:2312.10048
-
Automatic Scoring of Students' Science Writing Using Hybrid Neural Network 2 Dec 2023 · 0 repositories · arXiv:2312.03752
-
IAG: Induction-Augmented Generation Framework for Answering Reasoning Questions 30 Nov 2023 · 0 repositories · arXiv:2311.18397
-
LLVMs4Protest: Harnessing the Power of Large Language and Vision Models for Deciphering Protests in the News 30 Nov 2023 · 1 repository · arXiv:2311.18241
-
Transfer Learning across Different Chemical Domains: Virtual Screening of Organic Materials with Deep Learning Models Pretrained on Small Molecule and Chemical Reaction Data 30 Nov 2023 · 0 repositories · arXiv:2311.18377
-
A natural language processing-based approach: mapping human perception by understanding deep semantic features in street view images 29 Nov 2023 · 0 repositories · arXiv:2311.17354
-
Biomedical knowledge graph-optimized prompt generation for large language models 29 Nov 2023 · 1 repository · arXiv:2311.17330Syntology official: harvested, nothing ran · 0 ran · 3 unverified (of 3 harvested samples)
-
Enhancing Answer Selection in Community Question Answering with Pre-trained and Large Language Models 29 Nov 2023 · 0 repositories · arXiv:2311.17502
-
Improving the Robustness of Transformer-based Large Language Models with Dynamic Attention 29 Nov 2023 · 0 repositories · arXiv:2311.17400
-
RoKEPG: RoBERTa and Knowledge Enhancement for Prescription Generation of Traditional Chinese Medicine 29 Nov 2023 · 0 repositories · arXiv:2311.17307
-
TARGET: Template-Transferable Backdoor Attack Against Prompt-based NLP Models via GPT4 29 Nov 2023 · 0 repositories · arXiv:2311.17429
-
TimelyGPT: Extrapolatable Transformer Pre-training for Long-term Time-Series Forecasting in Healthcare 29 Nov 2023 · 0 repositories · arXiv:2312.00817
-
TurkishBERTweet: Fast and Reliable Large Language Model for Social Media Analysis 29 Nov 2023 · 2 repositories · arXiv:2311.18063
-
Natural Language Processing Through Transfer Learning: A Case Study on Sentiment Analysis 28 Nov 2023 · 0 repositories · arXiv:2311.16965
-
Syntax-Informed Interactive Model for Comprehensive Aspect-Based Sentiment Analysis 28 Nov 2023 · 0 repositories · arXiv:2312.03739
-
A Social-aware Gaussian Pre-trained Model for Effective Cold-start Recommendation 27 Nov 2023 · 0 repositories · arXiv:2311.15790
-
BERT Goes Off-Topic: Investigating the Domain Transfer Challenge using Genre Classification 27 Nov 2023 · 1 repository · arXiv:2311.16083
-
Leveraging deep active learning to identify low-resource mobility functioning information in public clinical notes 27 Nov 2023 · 0 repositories · arXiv:2311.15946
-
SSIN: Self-Supervised Learning for Rainfall Spatial Interpolation 27 Nov 2023 · 1 repository · arXiv:2311.15530
-
Uncertainty-aware Language Modeling for Selective Question Answering 26 Nov 2023 · 0 repositories · arXiv:2311.15451
-
SwiftLearn: A Data-Efficient Training Method of Deep Learning Models using Importance Sampling 25 Nov 2023 · 0 repositories · arXiv:2311.15134
-
CMed-GPT: Prompt Tuning for Entity-Aware Chinese Medical Dialogue Generation 24 Nov 2023 · 0 repositories · arXiv:2311.14539
-
A Multi-solution Study on GDPR AI-enabled Completeness Checking of DPAs 23 Nov 2023 · 0 repositories · arXiv:2311.13881
-
Annotation Sensitivity: Training Data Collection Methods Affect Model Performance 23 Nov 2023 · 1 repository · arXiv:2311.14212
-
Minimizing Factual Inconsistency and Hallucination in Large Language Models 23 Nov 2023 · 0 repositories · arXiv:2311.13878
-
Current Topological and Machine Learning Applications for Bias Detection in Text 22 Nov 2023 · 0 repositories · arXiv:2311.13495
-
Detecting out-of-distribution text using topological features of transformer-based language models 22 Nov 2023 · 1 repository · arXiv:2311.13102
-
White-Box Transformers via Sparse Rate Reduction: Compression Is All There Is? 22 Nov 2023 · 1 repository · arXiv:2311.13110Syntology 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
LowResource at BLP-2023 Task 2: Leveraging BanglaBert for Low Resource Sentiment Analysis of Bangla Language 21 Nov 2023 · 1 repository · arXiv:2311.12735
-
Utilizing Language Models for Tour Itinerary Recommendation 21 Nov 2023 · 0 repositories · arXiv:2311.12355
-
LogLead -- Fast and Integrated Log Loader, Enhancer, and Anomaly Detector 20 Nov 2023 · 1 repository · arXiv:2311.11809
-
LQ-LoRA: Low-rank Plus Quantized Matrix Decomposition for Efficient Language Model Finetuning 20 Nov 2023 · 1 repository · arXiv:2311.12023Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 12 harvested samples)
-
Tensor-Aware Energy Accounting 19 Nov 2023 · 1 repository · arXiv:2311.11424
-
Compositional Fusion of Signals in Data Embedding 18 Nov 2023 · 0 repositories · arXiv:2311.11085
-
Bias A-head? Analyzing Bias in Transformer-Based Language Model Attention Heads 17 Nov 2023 · 0 repositories · arXiv:2311.10395
-
Extracting periodontitis diagnosis in clinical notes with RoBERTa and regular expression 17 Nov 2023 · 0 repositories · arXiv:2311.10809
-
Hashing it Out: Predicting Unhealthy Conversations on Twitter 17 Nov 2023 · 1 repository · arXiv:2311.10596
-
Use GPT-J Prompt Generation with RoBERTa for NER Models on Diagnosis Extraction of Periodontal Diagnosis from Electronic Dental Records 17 Nov 2023 · 0 repositories · arXiv:2311.10810
-
ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems 16 Nov 2023 · 1 repository · arXiv:2311.09476Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 15 harvested samples)
-
Generative AI for Hate Speech Detection: Evaluation and Findings 16 Nov 2023 · 0 repositories · arXiv:2311.09993
-
LifeTox: Unveiling Implicit Toxicity in Life Advice 16 Nov 2023 · 1 repository · arXiv:2311.09585
-
Sequencing Matters: A Generate-Retrieve-Generate Model for Building Conversational Agents 16 Nov 2023 · 0 repositories · arXiv:2311.09513
-
Chemist-X: Large Language Model-empowered Agent for Reaction Condition Recommendation in Chemical Synthesis 16 Nov 2023 · 0 repositories · arXiv:2311.10776
-
An Eye on Clinical BERT: Investigating Language Model Generalization for Diabetic Eye Disease Phenotyping 15 Nov 2023 · 1 repository · arXiv:2311.08687
-
Exponentially Faster Language Modelling 15 Nov 2023 · 3 repositories · arXiv:2311.10770
-
German FinBERT: A German Pre-trained Language Model 15 Nov 2023 · 0 repositories · arXiv:2311.08793
-
Toucan: Token-Aware Character Level Language Modeling 15 Nov 2023 · 0 repositories · arXiv:2311.08620
-
"We Demand Justice!": Towards Social Context Grounding of Political Texts 15 Nov 2023 · 1 repository · arXiv:2311.09106
-
XplainLLM: A Knowledge-Augmented Dataset for Reliable Grounded Explanations in LLMs 15 Nov 2023 · 1 repository · arXiv:2311.08614
-
AI-generated text boundary detection with RoFT 14 Nov 2023 · 1 repository · arXiv:2311.08349Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Exploring Semi-supervised Hierarchical Stacked Encoder for Legal Judgement Prediction 14 Nov 2023 · 1 repository · arXiv:2311.08103
-
Investigating the Encoding of Words in BERT's Neurons using Feature Textualization 14 Nov 2023 · 0 repositories · arXiv:2311.08240
-
Memory-efficient Stochastic methods for Memory-based Transformers 14 Nov 2023 · 1 repository · arXiv:2311.08123
-
Teach me with a Whisper: Enhancing Large Language Models for Analyzing Spoken Transcripts using Speech Embeddings 13 Nov 2023 · 0 repositories · arXiv:2311.07014
-
Retrieval and Generative Approaches for a Pregnancy Chatbot in Nepali with Stemmed and Non-Stemmed Data : A Comparative Study 12 Nov 2023 · 0 repositories · arXiv:2311.06898
-
Establishing Performance Baselines in Fine-Tuning, Retrieval-Augmented Generation and Soft-Prompting for Non-Specialist LLM Users 10 Nov 2023 · 0 repositories · arXiv:2311.05903
-
Deep Natural Language Feature Learning for Interpretable Prediction 9 Nov 2023 · 1 repository · arXiv:2311.05754
-
LogShield: A Transformer-based APT Detection System Leveraging Self-Attention 9 Nov 2023 · 0 repositories · arXiv:2311.05733