Methods › Natural Language Processing › Subword Segmentation › WordPiece › Papers, page 19
WordPiece
Papers archive 2025-07-28
archive papers tagged: 7,063 · with a code link: 2,910 · where Syntology ran a sample: 650 (529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (650 of 7,063 tagged: 529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument)
Page 19 of 71: papers 1,801 to 1,900 of 7,063, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Exploiting ChatGPT for Diagnosing Autism-Associated Language Disorders and Identifying Distinct Features 3 May 2024 · 1 repository · arXiv:2405.01799
-
Attribution in Scientific Literature: New Benchmark and Methods 3 May 2024 · 0 repositories · arXiv:2405.02228
-
Structural Pruning of Pre-trained Language Models via Neural Architecture Search 3 May 2024 · 1 repository · arXiv:2405.02267
-
Early Transformers: A study on Efficient Training of Transformer Models through Early-Bird Lottery Tickets 2 May 2024 · 0 repositories · arXiv:2405.02353
-
Investigating Wit, Creativity, and Detectability of Large Language Models in Domain-Specific Writing Style Adaptation of Reddit's Showerthoughts 2 May 2024 · 1 repository · arXiv:2405.01660
-
A Named Entity Recognition and Topic Modeling-based Solution for Locating and Better Assessment of Natural Disasters in Social Media 1 May 2024 · 0 repositories · arXiv:2405.00903
-
Integrating A.I. in Higher Education: Protocol for a Pilot Study with 'SAMCares: An Adaptive Learning Hub' 1 May 2024 · 1 repository · arXiv:2405.00330
-
Opinion Mining Using Pre-Trained Large Language Models: Identifying the Type, Polarity, Intensity, Expression, and Source of Private States 1 May 2024 · 1 repository
-
Graph Neural Network Approach to Semantic Type Detection in Tables 30 Apr 2024 · 1 repository · arXiv:2405.00123
-
Towards a Search Engine for Machines: Unified Ranking for Multiple Retrieval-Augmented Large Language Models 30 Apr 2024 · 1 repository · arXiv:2405.00175
-
A cost minimization approach to fix the vocabulary size in a tokenizer for an End-to-End ASR system 29 Apr 2024 · 0 repositories · arXiv:2406.02563
-
FeDeRA:Efficient Fine-tuning of Language Models in Federated Learning Leveraging Weight Decomposition 29 Apr 2024 · 0 repositories · arXiv:2404.18848
-
Can Perplexity Predict Fine-Tuning Performance? An Investigation of Tokenization Effects on Sequential Language Models for Nepali 28 Apr 2024 · 0 repositories · arXiv:2404.18071
-
L3Cube-MahaNews: News-based Short Text and Long Document Classification Datasets in Marathi 28 Apr 2024 · 1 repository · arXiv:2404.18216
-
Tabular Embedding Model (TEM): Finetuning Embedding Models For Tabular RAG Applications 28 Apr 2024 · 0 repositories · arXiv:2405.01585
-
Enhancing Pre-Trained Generative Language Models with Question Attended Span Extraction on Machine Reading Comprehension 27 Apr 2024 · 0 repositories · arXiv:2404.17991
-
Tool Calling: Enhancing Medication Consultation via Retrieval-Augmented Large Language Models 27 Apr 2024 · 0 repositories · arXiv:2404.17897
-
Enhancing Legal Compliance and Regulation Analysis with Large Language Models 26 Apr 2024 · 0 repositories · arXiv:2404.17522
-
Human-Imperceptible Retrieval Poisoning Attacks in LLM-Powered Applications 26 Apr 2024 · 0 repositories · arXiv:2404.17196
-
Retrieval-Augmented Generation with Knowledge Graphs for Customer Service Question Answering 26 Apr 2024 · 0 repositories · arXiv:2404.17723
-
A Short Survey of Human Mobility Prediction in Epidemic Modeling from Transformers to LLMs 25 Apr 2024 · 0 repositories · arXiv:2404.16921
-
Análise de ambiguidade linguística em modelos de linguagem de grande escala (LLMs) 25 Apr 2024 · 0 repositories · arXiv:2404.16653
-
Evaluating Consistency and Reasoning Capabilities of Large Language Models 25 Apr 2024 · 0 repositories · arXiv:2404.16478
-
Exploring Internal Numeracy in Language Models: A Case Study on ALBERT 25 Apr 2024 · 0 repositories · arXiv:2404.16574
-
Incorporating Lexical and Syntactic Knowledge for Unsupervised Cross-Lingual Transfer 25 Apr 2024 · 1 repository · arXiv:2404.16627
-
A Comprehensive Survey on Evaluating Large Language Model Applications in the Medical Industry 24 Apr 2024 · 0 repositories · arXiv:2404.15777
-
BERT vs GPT for financial engineering 24 Apr 2024 · 0 repositories · arXiv:2405.12990
-
Detecting Conceptual Abstraction in LLMs 24 Apr 2024 · 0 repositories · arXiv:2404.15848
-
From Local to Global: A Graph RAG Approach to Query-Focused Summarization 24 Apr 2024 · 3 repositories · arXiv:2404.16130Syntology 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Prompt Leakage effect and defense strategies for multi-turn LLM interactions 24 Apr 2024 · 0 repositories · arXiv:2404.16251
-
Studying Large Language Model Behaviors Under Context-Memory Conflicts With Real Documents 24 Apr 2024 · 1 repository · arXiv:2404.16032Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Telco-RAG: Navigating the Challenges of Retrieval-Augmented Language Models for Telecommunications 24 Apr 2024 · 1 repository · arXiv:2404.15939
-
AI and Machine Learning for Next Generation Science Assessments 23 Apr 2024 · 0 repositories · arXiv:2405.06660
-
IryoNLP at MEDIQA-CORR 2024: Tackling the Medical Error Detection & Correction Task On the Shoulders of Medical Agents 23 Apr 2024 · 0 repositories · arXiv:2404.15488
-
Pre-Calc: Learning to Use the Calculator Improves Numeracy in Language Models 22 Apr 2024 · 1 repository · arXiv:2404.14355Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
LLMs Know What They Need: Leveraging a Missing Information Guided Framework to Empower Retrieval-Augmented Generation 22 Apr 2024 · 1 repository · arXiv:2404.14043
-
Marking: Visual Grading with Highlighting Errors and Annotating Missing Bits 22 Apr 2024 · 0 repositories · arXiv:2404.14301
-
Typos that Broke the RAG's Back: Genetic Attack on RAG Pipeline by Simulating Documents in the Wild via Low-level Perturbations 22 Apr 2024 · 1 repository · arXiv:2404.13948Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
What do Transformers Know about Government? 22 Apr 2024 · 1 repository · arXiv:2404.14270
-
Zero-shot Cross-lingual Stance Detection via Adversarial Language Adaptation 22 Apr 2024 · 1 repository · arXiv:2404.14339
-
Automated Text Mining of Experimental Methodologies from Biomedical Literature 21 Apr 2024 · 0 repositories · arXiv:2404.13779
-
Evaluating Retrieval Quality in Retrieval-Augmented Generation 21 Apr 2024 · 1 repository · arXiv:2404.13781Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
BERT: Accelerating Vital Signs Measurement for Bioradar with An Efficient Recursive Technique 20 Apr 2024 · 0 repositories · arXiv:2404.13315
-
Do "English" Named Entity Recognizers Work Well on Global Englishes? 20 Apr 2024 · 2 repositories · arXiv:2404.13465
-
Evaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge 20 Apr 2024 · 1 repository · arXiv:2404.13292
-
Dubo-SQL: Diverse Retrieval-Augmented Generation and Fine Tuning for Text-to-SQL 19 Apr 2024 · 1 repository · arXiv:2404.12560
-
Enabling Natural Zero-Shot Prompting on Encoder Models via Statement-Tuning 19 Apr 2024 · 0 repositories · arXiv:2404.12897
-
Multi Class Depression Detection Through Tweets using Artificial Intelligence 19 Apr 2024 · 1 repository · arXiv:2404.13104
-
Unlocking Multi-View Insights in Knowledge-Dense Retrieval-Augmented Generation 19 Apr 2024 · 0 repositories · arXiv:2404.12879
-
Augmenting emotion features in irony detection with Large language modeling 18 Apr 2024 · 0 repositories · arXiv:2404.12291
-
emrQA-msquad: A Medical Dataset Structured with the SQuAD V2.0 Framework, Enriched with emrQA Medical Information 18 Apr 2024 · 0 repositories · arXiv:2404.12050
-
iRAG: Advancing RAG for Videos with an Incremental Approach 18 Apr 2024 · 0 repositories · arXiv:2404.12309
-
LongEmbed: Extending Embedding Models for Long Context Retrieval 18 Apr 2024 · 1 repository · arXiv:2404.12096Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
RAGAR, Your Falsehood Radar: RAG-Augmented Reasoning for Political Fact-Checking using Multimodal Large Language Models 18 Apr 2024 · 0 repositories · arXiv:2404.12065
-
RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation 18 Apr 2024 · 0 repositories · arXiv:2404.12457
-
RAM: Towards an Ever-Improving Memory System by Learning from Communications 18 Apr 2024 · 0 repositories · arXiv:2404.12045
-
Stance Detection on Social Media with Fine-Tuned Large Language Models 18 Apr 2024 · 0 repositories · arXiv:2404.12171
-
A Survey on Retrieval-Augmented Text Generation for Large Language Models 17 Apr 2024 · 0 repositories · arXiv:2404.10981
-
Demystifying Legalese: An Automated Approach for Summarizing and Analyzing Overlaps in Privacy Policies and Terms of Service 17 Apr 2024 · 0 repositories · arXiv:2404.13087
-
Enhancing Q&A with Domain-Specific Fine-Tuning and Iterative Reasoning: A Comparative Study 17 Apr 2024 · 0 repositories · arXiv:2404.11792
-
Improvement in Semantic Address Matching using Natural Language Processing 17 Apr 2024 · 0 repositories · arXiv:2404.11691
-
A Sentiment Analysis of Medical Text Based on Deep Learning 16 Apr 2024 · 0 repositories · arXiv:2404.10503
-
BayesJudge: Bayesian Kernel Language Modelling with Confidence Uncertainty in Legal Judgment Prediction 16 Apr 2024 · 0 repositories · arXiv:2404.10481
-
Empowering Interdisciplinary Research with BERT-Based Models: An Approach Through SciBERT-CNN with Topic Modeling 16 Apr 2024 · 0 repositories · arXiv:2404.13078
-
LaDiC: Are Diffusion Models Really Inferior to Autoregressive Counterparts for Image-to-Text Generation? 16 Apr 2024 · 1 repository · arXiv:2404.10763
-
Relational Graph Convolutional Networks for Sentiment Analysis 16 Apr 2024 · 0 repositories · arXiv:2404.13079
-
Spiral of Silence: How is Large Language Model Killing Information Retrieval? -- A Case Study on Open Domain Question Answering 16 Apr 2024 · 1 repository · arXiv:2404.10496
-
Detecting AI Generated Text Based on NLP and Machine Learning Approaches 15 Apr 2024 · 0 repositories · arXiv:2404.10032
-
BERT-LSH: Reducing Absolute Compute For Attention 12 Apr 2024 · 1 repository · arXiv:2404.08836
-
Reducing hallucination in structured outputs via Retrieval-Augmented Generation 12 Apr 2024 · 0 repositories · arXiv:2404.08189
-
Generative Information Retrieval Evaluation 11 Apr 2024 · 0 repositories · arXiv:2404.08137
-
Emotion-cause pair extraction method based on multi-granularity information and multi-module interaction 10 Apr 2024 · 0 repositories · arXiv:2404.06812
-
Llama-VITS: Enhancing TTS Synthesis with Semantic Awareness 10 Apr 2024 · 1 repository · arXiv:2404.06714Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Not All Contexts Are Equal: Teaching LLMs Credibility-aware Generation 10 Apr 2024 · 1 repository · arXiv:2404.06809
-
Simpler becomes Harder: Do LLMs Exhibit a Coherent Behavior on Simplified Corpora? 10 Apr 2024 · 1 repository · arXiv:2404.06838
-
Superposition Prompting: Improving and Accelerating Retrieval-Augmented Generation 10 Apr 2024 · 1 repository · arXiv:2404.06910
-
Comprehensive Study on German Language Models for Clinical and Biomedical Text Understanding 8 Apr 2024 · 0 repositories · arXiv:2404.05694
-
MedExpQA: Multilingual Benchmarking of Large Language Models for Medical Question Answering 8 Apr 2024 · 0 repositories · arXiv:2404.05590
-
Semantic Stealth: Adversarial Text Attacks on NLP Using Several Methods 8 Apr 2024 · 0 repositories · arXiv:2404.05159
-
A Morphology-Based Investigation of Positional Encodings 6 Apr 2024 · 0 repositories · arXiv:2404.04530
-
Deciphering Political Entity Sentiment in News with Large Language Models: Zero-Shot and Few-Shot Strategies 5 Apr 2024 · 1 repository · arXiv:2404.04361
-
Investigating the Robustness of Modelling Decisions for Few-Shot Cross-Topic Stance Detection: A Preregistered Study 5 Apr 2024 · 1 repository · arXiv:2404.03987
-
BanglaAutoKG: Automatic Bangla Knowledge Graph Construction with Semantic Neural Graph Filtering 4 Apr 2024 · 1 repository · arXiv:2404.03528Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
CBR-RAG: Case-Based Reasoning for Retrieval Augmented Generation in LLMs for Legal Question Answering 4 Apr 2024 · 1 repository · arXiv:2404.04302
-
CONFLARE: CONFormal LArge language model REtrieval 4 Apr 2024 · 1 repository · arXiv:2404.04287
-
Outlier-Efficient Hopfield Layers for Large Transformer-Based Models 4 Apr 2024 · 1 repository · arXiv:2404.03828Syntology official (archive's flag): 10 ran · 10 ran (of which 5 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 5 where Syntology's instrument failed) · 3 unverified (of 13 harvested samples)
-
The Death of Feature Engineering? BERT with Linguistic Features on SQuAD 2.0 4 Apr 2024 · 0 repositories · arXiv:2404.03184
-
BCAmirs at SemEval-2024 Task 4: Beyond Words: A Multimodal and Multilingual Exploration of Persuasion in Memes 3 Apr 2024 · 1 repository · arXiv:2404.03022
-
CLAPNQ: Cohesive Long-form Answers from Passages in Natural Questions for RAG systems 2 Apr 2024 · 1 repository · arXiv:2404.02103
-
GINopic: Topic Modeling with Graph Isomorphism Network 2 Apr 2024 · 1 repository · arXiv:2404.02115Syntology official: no sample here; runs from other or unrecorded repositories · 10 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 12 harvested samples) · 4 pointer-only (licence)
-
Symbolic Prompt Program Search: A Structure-Aware Approach to Efficient Compile-Time Prompt Optimization 2 Apr 2024 · 1 repository · arXiv:2404.02319
-
Advancing AI with Integrity: Ethical Challenges and Solutions in Neural Machine Translation 1 Apr 2024 · 0 repositories · arXiv:2404.01070
-
ARAGOG: Advanced RAG Output Grading 1 Apr 2024 · 1 repository · arXiv:2404.01037
-
BERT-Enhanced Retrieval Tool for Homework Plagiarism Detection System 1 Apr 2024 · 0 repositories · arXiv:2404.01582
-
Securing Social Spaces: Harnessing Deep Learning to Eradicate Cyberbullying 1 Apr 2024 · 0 repositories · arXiv:2404.03686
-
Observations on Building RAG Systems for Technical Documents 31 Mar 2024 · 0 repositories · arXiv:2404.00657
-
RQ-RAG: Learning to Refine Queries for Retrieval Augmented Generation 31 Mar 2024 · 1 repository · arXiv:2404.00610
-
Leveraging Pre-trained and Transformer-derived Embeddings from EHRs to Characterize Heterogeneity Across Alzheimer's Disease and Related Dementias 30 Mar 2024 · 0 repositories · arXiv:2404.00464
-
Explainable Deep Learning: A Visual Analytics Approach with Transition Matrices 29 Mar 2024 · 1 repository
-
LayerNorm: A key component in parameter-efficient fine-tuning 29 Mar 2024 · 0 repositories · arXiv:2403.20284