Methods › Natural Language Processing › Subword Segmentation › WordPiece › Papers, page 4
WordPiece
Papers archive 2025-07-28
archive papers tagged: 7,063 · with a code link: 2,910 · where Syntology ran a sample: 650 (529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (650 of 7,063 tagged: 529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument)
Page 4 of 71: papers 301 to 400 of 7,063, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Pneuma: Leveraging LLMs for Tabular Data Representation and Retrieval in an End-to-End System 12 Apr 2025 · 1 repository · arXiv:2504.09207
-
Semantic Commit: Helping Users Update Intent Specifications for AI Memory at Scale 12 Apr 2025 · 0 repositories · arXiv:2504.09283
-
Adopting Large Language Models to Automated System Integration 11 Apr 2025 · 0 repositories · arXiv:2504.08490
-
HyperCore: The Core Framework for Building Hyperbolic Foundation Models with Comprehensive Modules 11 Apr 2025 · 1 repository · arXiv:2504.08912Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Integrated ensemble of BERT- and features-based models for authorship attribution in Japanese literary works 11 Apr 2025 · 0 repositories · arXiv:2504.08527
-
ModernBERT or DeBERTaV3? Examining Architecture and Data Influence on Transformer Encoder Models Performance 11 Apr 2025 · 0 repositories · arXiv:2504.08716
-
Out of Style: RAG's Fragility to Linguistic Variation 11 Apr 2025 · 1 repository · arXiv:2504.08231Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 14 harvested samples)
-
PCA-RAG: Principal Component Analysis for Efficient Retrieval-Augmented Generation 11 Apr 2025 · 0 repositories · arXiv:2504.08386
-
RTLRepoCoder: Repository-Level RTL Code Completion through the Combination of Fine-Tuning and Retrieval Augmentation 11 Apr 2025 · 0 repositories · arXiv:2504.08862
-
The Other Side of the Coin: Exploring Fairness in Retrieval-Augmented Generation 11 Apr 2025 · 1 repository · arXiv:2504.12323
-
A System for Comprehensive Assessment of RAG Frameworks 10 Apr 2025 · 1 repository · arXiv:2504.07803
-
AgentAda: Skill-Adaptive Data Analytics for Tailored Insight Discovery 10 Apr 2025 · 1 repository · arXiv:2504.07421
-
APSQ: Additive Partial Sum Quantization with Algorithm-Hardware Co-Design 10 Apr 2025 · 0 repositories · arXiv:2505.03748
-
Beyond LLMs: A Linguistic Approach to Causal Graph Generation from Narrative Texts 10 Apr 2025 · 0 repositories · arXiv:2504.07459
-
ConceptFormer: Towards Efficient Use of Knowledge-Graph Embeddings in Large Language Models 10 Apr 2025 · 0 repositories · arXiv:2504.07624
-
MRD-RAG: Enhancing Medical Diagnosis with Multi-Round Retrieval-Augmented Generation 10 Apr 2025 · 1 repository · arXiv:2504.07724
-
A new training approach for text classification in Mental Health: LatentGLoss 9 Apr 2025 · 1 repository · arXiv:2504.07245
-
Evaluating Retrieval Augmented Generative Models for Document Queries in Transportation Safety 9 Apr 2025 · 0 repositories · arXiv:2504.07022
-
Poly-Vector Retrieval: Reference and Content Embeddings for Legal Documents 9 Apr 2025 · 0 repositories · arXiv:2504.10508
-
Graph-based Approaches and Functionalities in Retrieval-Augmented Generation: A Comprehensive Survey 8 Apr 2025 · 0 repositories · arXiv:2504.10499
-
PathGPT: Leveraging Large Language Models for Personalized Route Generation 8 Apr 2025 · 0 repositories · arXiv:2504.05846
-
Retrieval Augmented Generation with Collaborative Filtering for Personalized Text Generation 8 Apr 2025 · 1 repository · arXiv:2504.05731
-
AI for Climate Finance: Agentic Retrieval and Multi-Step Reasoning for Early Warning System Investments 7 Apr 2025 · 0 repositories · arXiv:2504.05104
-
Collab-RAG: Boosting Retrieval-Augmented Generation for Complex Question Answering via White-Box and Black-Box LLM Collaboration 7 Apr 2025 · 1 repository · arXiv:2504.04915Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Leveraging LLMs for Utility-Focused Annotation: Reducing Manual Effort for Retrieval and RAG 7 Apr 2025 · 0 repositories · arXiv:2504.05220
-
Driving-RAG: Driving Scenarios Embedding, Search, and RAG Applications 6 Apr 2025 · 0 repositories · arXiv:2504.04419
-
Exploring Generative AI Techniques in Government: A Case Study 6 Apr 2025 · 0 repositories · arXiv:2504.10497
-
Hierarchical Planning for Complex Tasks with Knowledge Graph-RAG and Symbolic Verification 6 Apr 2025 · 0 repositories · arXiv:2504.04578
-
Psychological Health Knowledge-Enhanced LLM-based Social Network Crisis Intervention Text Transfer Recognition Method 5 Apr 2025 · 0 repositories · arXiv:2504.07983
-
QE-RAG: A Robust Retrieval-Augmented Generation Benchmark for Query Entry Errors 5 Apr 2025 · 0 repositories · arXiv:2504.04062
-
Sigma: A dataset for text-to-code semantic parsing with statistical analysis 5 Apr 2025 · 1 repository · arXiv:2504.04301
-
Efficient Dynamic Clustering-Based Document Compression for Retrieval-Augmented-Generation 4 Apr 2025 · 1 repository · arXiv:2504.03165
-
Generative AI Enhanced Financial Risk Management Information Retrieval 4 Apr 2025 · 1 repository · arXiv:2504.06293
-
Learning Natural Language Constraints for Safe Reinforcement Learning of Language Agents 4 Apr 2025 · 0 repositories · arXiv:2504.03185
-
Multilingual Retrieval-Augmented Generation for Knowledge-Intensive Task 4 Apr 2025 · 0 repositories · arXiv:2504.03616
-
Practical Poisoning Attacks against Retrieval-Augmented Generation 4 Apr 2025 · 0 repositories · arXiv:2504.03957
-
Rotation Invariance in Floor Plan Digitization using Zernike Moments 4 Apr 2025 · 0 repositories · arXiv:2504.03241
-
Structured Extraction of Process Structure Properties Relationships in Materials Science 4 Apr 2025 · 0 repositories · arXiv:2504.03979
-
AD-GPT: Large Language Models in Alzheimer's Disease 3 Apr 2025 · 0 repositories · arXiv:2504.03071
-
Adapting Large Language Models for Multi-Domain Retrieval-Augmented-Generation 3 Apr 2025 · 0 repositories · arXiv:2504.02411
-
HyperRAG: Enhancing Quality-Efficiency Tradeoffs in Retrieval-Augmented Generation with Reranker KV-Cache Reuse 3 Apr 2025 · 0 repositories · arXiv:2504.02921
-
A thorough benchmark of automatic text classification: From traditional approaches to large language models 2 Apr 2025 · 1 repository · arXiv:2504.01930
-
Biomedical Question Answering via Multi-Level Summarization on a Local Knowledge Graph 2 Apr 2025 · 0 repositories · arXiv:2504.01309
-
Breaking BERT: Gradient Attack on Twitter Sentiment Analysis for Targeted Misclassification 2 Apr 2025 · 1 repository · arXiv:2504.01345
-
Context-Aware Toxicity Detection in Multiplayer Games: Integrating Domain-Adaptive Pretraining and Match Metadata 2 Apr 2025 · 1 repository · arXiv:2504.01534
-
CoRAG: Collaborative Retrieval-Augmented Generation 2 Apr 2025 · 0 repositories · arXiv:2504.01883
-
GeoRAG: A Question-Answering Approach from a Geographical Perspective 2 Apr 2025 · 0 repositories · arXiv:2504.01458
-
LARGE: Legal Retrieval Augmented Generation Evaluation Tool 2 Apr 2025 · 1 repository · arXiv:2504.01840
-
OnRL-RAG: Real-Time Personalized Mental Health Dialogue System 2 Apr 2025 · 0 repositories · arXiv:2504.02894
-
Scaling Test-Time Inference with Policy-Optimized, Dynamic Retrieval-Augmented Generation via KV Caching and Decoding 2 Apr 2025 · 0 repositories · arXiv:2504.01281
-
A Unified Virtual Mixture-of-Experts Framework:Enhanced Inference and Hallucination Mitigation in Single-Model System 1 Apr 2025 · 0 repositories · arXiv:2504.03739
-
Accelerating Causal Network Discovery of Alzheimer Disease Biomarkers via Scientific Literature-based Retrieval Augmented Generation 1 Apr 2025 · 0 repositories · arXiv:2504.08768
-
GS_DravidianLangTech@2025: Women Targeted Abusive Texts Detection on Social Media 1 Apr 2025 · 0 repositories · arXiv:2504.02863
-
LLM-Assisted Proactive Threat Intelligence for Automated Reasoning 1 Apr 2025 · 0 repositories · arXiv:2504.00428
-
Transformer-Based Named Entity Recognition for Automated Server Provisioning 1 Apr 2025 · 1 repository
-
WikiVideo: Article Generation from Multiple Videos 1 Apr 2025 · 1 repository · arXiv:2504.00939
-
A Systematic Evaluation of LLM Strategies for Mental Health Text Analysis: Fine-tuning vs. Prompt Engineering vs. RAG 31 Mar 2025 · 0 repositories · arXiv:2503.24307
-
Dynamic Parametric Retrieval Augmented Generation for Test-time Knowledge Enhancement 31 Mar 2025 · 1 repository · arXiv:2503.23895Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
CrossFormer: Cross-Segment Semantic Fusion for Document Segmentation 31 Mar 2025 · 0 repositories · arXiv:2503.23671
-
Enhancing Large Language Models (LLMs) for Telecommunications using Knowledge Graphs and Retrieval-Augmented Generation 31 Mar 2025 · 0 repositories · arXiv:2503.24245
-
Synthetic News Generation for Fake News Classification 31 Mar 2025 · 0 repositories · arXiv:2503.24206
-
UltraRAG: A Modular and Automated Toolkit for Adaptive Retrieval-Augmented Generation 31 Mar 2025 · 1 repository · arXiv:2504.08761
-
Hyper-RAG: Combating LLM Hallucinations using Hypergraph-Driven Retrieval-Augmented Generation 30 Mar 2025 · 0 repositories · arXiv:2504.08758
-
Measuring Online Hate on 4chan using Pre-trained Deep Learning Models 30 Mar 2025 · 0 repositories · arXiv:2504.00045
-
Multi-Stakeholder Disaster Insights from Social Media Using Large Language Models 30 Mar 2025 · 0 repositories · arXiv:2504.00046
-
Enhancing Knowledge Graph Completion with Entity Neighborhood and Relation Context 29 Mar 2025 · 0 repositories · arXiv:2503.23205
-
MHTS: Multi-Hop Tree Structure Framework for Generating Difficulty-Controllable QA Datasets for RAG Evaluation 29 Mar 2025 · 0 repositories · arXiv:2504.08756
-
Historical Ink: Exploring Large Language Models for Irony Detection in 19th-Century Spanish 28 Mar 2025 · 1 repository · arXiv:2503.22585
-
As easy as PIE: understanding when pruning causes language models to disagree 27 Mar 2025 · 1 repository · arXiv:2503.21714
-
Hybrid Emotion Recognition: Enhancing Customer Interactions Through Acoustic and Textual Analysis 27 Mar 2025 · 0 repositories · arXiv:2503.21927
-
HyperGraphRAG: Retrieval-Augmented Generation with Hypergraph-Structured Knowledge Representation 27 Mar 2025 · 1 repository · arXiv:2503.21322Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
MemInsight: Autonomous Memory Augmentation for LLM Agents 27 Mar 2025 · 0 repositories · arXiv:2503.21760
-
Real-Time Evaluation Models for RAG: Who Detects Hallucinations Best? 27 Mar 2025 · 0 repositories · arXiv:2503.21157
-
ReaRAG: Knowledge-guided Reasoning Enhances Factuality of Large Reasoning Models with Iterative Retrieval Augmented Generation 27 Mar 2025 · 1 repository · arXiv:2503.21729
-
A Survey of Multimodal Retrieval-Augmented Generation 26 Mar 2025 · 0 repositories · arXiv:2504.08748
-
Advancements in Natural Language Processing: Exploring Transformer-Based Architectures for Text Understanding 26 Mar 2025 · 0 repositories · arXiv:2503.20227
-
Advancing Vulnerability Classification with BERT: A Multi-Objective Learning Model 26 Mar 2025 · 0 repositories · arXiv:2503.20831
-
MCTS-RAG: Enhancing Retrieval-Augmented Generation with Monte Carlo Tree Search 26 Mar 2025 · 1 repository · arXiv:2503.20757
-
RALLRec+: Retrieval Augmented Large Language Model Recommendation with Reasoning 26 Mar 2025 · 1 repository · arXiv:2503.20430
-
RGL: A Graph-Centric, Modular Framework for Efficient Retrieval-Augmented Generation on Graphs 25 Mar 2025 · 1 repository · arXiv:2503.19314
-
CausalRAG: Integrating Causal Graphs into Retrieval-Augmented Generation 25 Mar 2025 · 0 repositories · arXiv:2503.19878
-
Machine-assisted writing evaluation: Exploring pre-trained language models in analyzing argumentative moves 25 Mar 2025 · 0 repositories · arXiv:2503.19279
-
Taxonomy Inference for Tabular Data Using Large Language Models 25 Mar 2025 · 0 repositories · arXiv:2503.21810
-
Analyzing Islamophobic Discourse Using Semi-Coded Terms and LLMs 24 Mar 2025 · 0 repositories · arXiv:2503.18273
-
Construction Identification and Disambiguation Using BERT: A Case Study of NPN 24 Mar 2025 · 0 repositories · arXiv:2503.18751
-
Enhancing Recommender Systems Using Textual Embeddings from Pre-trained Language Models 24 Mar 2025 · 0 repositories · arXiv:2504.08746
-
Improving RAG for Personalization with Author Features and Contrastive Examples 24 Mar 2025 · 1 repository · arXiv:2504.08745
-
Synthetic Function Demonstrations Improve Generation in Low-Resource Programming Languages 24 Mar 2025 · 0 repositories · arXiv:2503.18760
-
ExpertRAG: Efficient RAG with Mixture of Experts -- Optimizing Context Retrieval for Adaptive LLM Responses 23 Mar 2025 · 0 repositories · arXiv:2504.08744
-
Investigating Recent Large Language Models for Vietnamese Machine Reading Comprehension 23 Mar 2025 · 0 repositories · arXiv:2503.18062
-
LakotaBERT: A Transformer-based Model for Low Resource Lakota Language 23 Mar 2025 · 0 repositories · arXiv:2503.18212
-
Retrieval Augmented Generation and Understanding in Vision: A Survey and New Outlook 23 Mar 2025 · 1 repository · arXiv:2503.18016
-
Enhancing Arabic Automated Essay Scoring with Synthetic Data and Error Injection 22 Mar 2025 · 0 repositories · arXiv:2503.17739
-
Autonomous Radiotherapy Treatment Planning Using DOLA: A Privacy-Preserving, LLM-Based Optimization Agent 21 Mar 2025 · 0 repositories · arXiv:2503.17553
-
Meme Similarity and Emotion Detection using Multimodal Analysis 21 Mar 2025 · 0 repositories · arXiv:2503.17493
-
Design and Implementation of an FPGA-Based Hardware Accelerator for Transformer 20 Mar 2025 · 1 repository · arXiv:2503.16731
-
Don't Fight Hallucinations, Use Them: Estimating Image Realism using NLI over Atomic Facts 20 Mar 2025 · 1 repository · arXiv:2503.15948
-
Financial Analysis: Intelligent Financial Data Analysis System Based on LLM-RAG 20 Mar 2025 · 0 repositories · arXiv:2504.06279
-
InhibiDistilbert: Knowledge Distillation for a ReLU and Addition-based Transformer 20 Mar 2025 · 0 repositories · arXiv:2503.15983
-
Parameters vs. Context: Fine-Grained Control of Knowledge Reliance in Language Models 20 Mar 2025 · 1 repository · arXiv:2503.15888