Methods › Natural Language Processing › Subword Segmentation › WordPiece › Papers, page 9
WordPiece
Papers archive 2025-07-28
archive papers tagged: 7,063 · with a code link: 2,910 · where Syntology ran a sample: 650 (529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (650 of 7,063 tagged: 529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument)
Page 9 of 71: papers 801 to 900 of 7,063, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
A Reality Check on Context Utilisation for Retrieval-Augmented Generation 22 Dec 2024 · 1 repository · arXiv:2412.17031
-
AlzheimerRAG: Multimodal Retrieval Augmented Generation for PubMed articles 21 Dec 2024 · 0 repositories · arXiv:2412.16701
-
Distilling Large Language Models for Efficient Clinical Information Extraction 21 Dec 2024 · 0 repositories · arXiv:2501.00031
-
Formal Language Knowledge Corpus for Retrieval Augmented Generation 21 Dec 2024 · 0 repositories · arXiv:2412.16689
-
Identifying Cyberbullying Roles in Social Media 21 Dec 2024 · 0 repositories · arXiv:2412.16417
-
Quantum-Like Contextuality in Large Language Models 21 Dec 2024 · 1 repository · arXiv:2412.16806
-
Research on Violent Text Detection System Based on BERT-fasttext Model 21 Dec 2024 · 0 repositories · arXiv:2412.16455
-
TimeRAG: BOOSTING LLM Time Series Forecasting via Retrieval-Augmented Generation 21 Dec 2024 · 0 repositories · arXiv:2412.16643
-
Towards More Robust Retrieval-Augmented Generation: Evaluating RAG Under Adversarial Poisoning Attacks 21 Dec 2024 · 1 repository · arXiv:2412.16708
-
Adversarial Robustness through Dynamic Ensemble Learning 20 Dec 2024 · 0 repositories · arXiv:2412.16254
-
Decoding Linguistic Nuances in Mental Health Text Classification Using Expressive Narrative Stories 20 Dec 2024 · 0 repositories · arXiv:2412.16302
-
Don't Do RAG: When Cache-Augmented Generation is All You Need for Knowledge Tasks 20 Dec 2024 · 1 repository · arXiv:2412.15605Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
HybGRAG: Hybrid Retrieval-Augmented Generation on Textual and Relational Knowledge Bases 20 Dec 2024 · 0 repositories · arXiv:2412.16311
-
Towards Interpretable Radiology Report Generation via Concept Bottlenecks using a Multi-Agentic RAG 20 Dec 2024 · 1 repository · arXiv:2412.16086
-
XRAG: eXamining the Core -- Benchmarking Foundational Components in Advanced Retrieval-Augmented Generation 20 Dec 2024 · 1 repository · arXiv:2412.15529Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples)
-
Analysis and Visualization of Linguistic Structures in Large Language Models: Neural Representations of Verb-Particle Constructions in BERT 19 Dec 2024 · 0 repositories · arXiv:2412.14670
-
CORD: Balancing COnsistency and Rank Distillation for Robust Retrieval-Augmented Generation 19 Dec 2024 · 0 repositories · arXiv:2412.14581
-
Decade of Natural Language Processing in Chronic Pain: A Systematic Review 19 Dec 2024 · 0 repositories · arXiv:2412.15360
-
Dehallucinating Parallel Context Extension for Retrieval-Augmented Generation 19 Dec 2024 · 0 repositories · arXiv:2412.14905
-
DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs 19 Dec 2024 · 0 repositories · arXiv:2412.14838
-
Graph-Convolutional Networks: Named Entity Recognition and Large Language Model Embedding in Document Clustering 19 Dec 2024 · 0 repositories · arXiv:2412.14867
-
Knowledge Injection via Prompt Distillation 19 Dec 2024 · 0 repositories · arXiv:2412.14964
-
PA-RAG: RAG Alignment via Multi-Perspective Preference Optimization 19 Dec 2024 · 1 repository · arXiv:2412.14510
-
Query pipeline optimization for cancer patient question answering systems 19 Dec 2024 · 0 repositories · arXiv:2412.14751
-
Review-Then-Refine: A Dynamic Framework for Multi-Hop Question Answering with Temporal Adaptability 19 Dec 2024 · 0 repositories · arXiv:2412.15101
-
SKETCH: Structured Knowledge Enhanced Text Comprehension for Holistic Retrieval 19 Dec 2024 · 0 repositories · arXiv:2412.15443
-
Till the Layers Collapse: Compressing a Deep Neural Network through the Lenses of Batch Normalization Layers 19 Dec 2024 · 1 repository · arXiv:2412.15077
-
VISA: Retrieval Augmented Generation with Visual Source Attribution 19 Dec 2024 · 0 repositories · arXiv:2412.14457
-
Enhancing Rhetorical Figure Annotation: An Ontology-Based Web Application with RAG Integration 18 Dec 2024 · 1 repository · arXiv:2412.13799
-
EvoWiki: Evaluating LLMs on Evolving Knowledge 18 Dec 2024 · 0 repositories · arXiv:2412.13582
-
FarExStance: Explainable Stance Detection for Farsi 18 Dec 2024 · 2 repositories · arXiv:2412.14008
-
Federated Learning and RAG Integration: A Scalable Approach for Medical Large Language Models 18 Dec 2024 · 0 repositories · arXiv:2412.13720
-
RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment 18 Dec 2024 · 1 repository · arXiv:2412.13746
-
Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference 18 Dec 2024 · 2 repositories · arXiv:2412.13663Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
A MapReduce Approach to Effectively Utilize Long Context Information in Retrieval Augmented Language Models 17 Dec 2024 · 0 repositories · arXiv:2412.15271
-
Adaptations of AI models for querying the LandMatrix database in natural language 17 Dec 2024 · 1 repository · arXiv:2412.12961
-
C-FedRAG: A Confidential Federated Retrieval-Augmented Generation System 17 Dec 2024 · 0 repositories · arXiv:2412.13163
-
Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language Models 17 Dec 2024 · 0 repositories · arXiv:2412.15265
-
EXIT: Context-Aware Extractive Compression for Enhancing Retrieval-Augmented Generation 17 Dec 2024 · 1 repository · arXiv:2412.12559
-
LLMs are Also Effective Embedding Models: An In-depth Overview 17 Dec 2024 · 0 repositories · arXiv:2412.12591
-
OmniEval: An Omnidirectional and Automatic RAG Evaluation Benchmark in Financial Domain 17 Dec 2024 · 1 repository · arXiv:2412.13018
-
PERC: Plan-As-Query Example Retrieval for Underrepresented Code Generation 17 Dec 2024 · 0 repositories · arXiv:2412.12447
-
RAG-Star: Enhancing Deliberative Reasoning with Retrieval Augmented Verification and Refinement 17 Dec 2024 · 0 repositories · arXiv:2412.12881
-
RemoteRAG: A Privacy-Preserving LLM Cloud RAG Service 17 Dec 2024 · 0 repositories · arXiv:2412.12775
-
SimGRAG: Leveraging Similar Subgraphs for Knowledge Graphs Driven Retrieval-Augmented Generation 17 Dec 2024 · 1 repository · arXiv:2412.15272
-
What External Knowledge is Preferred by LLMs? Characterizing and Exploring Chain of Evidence in Imperfect Context 17 Dec 2024 · 0 repositories · arXiv:2412.12632
-
A Benchmark and Robustness Study of In-Context-Learning with Large Language Models in Music Entity Detection 16 Dec 2024 · 1 repository · arXiv:2412.11851
-
Investigating Mixture of Experts in Dense Retrieval 16 Dec 2024 · 0 repositories · arXiv:2412.11864
-
Look Ahead Text Understanding and LLM Stitching 16 Dec 2024 · 1 repository · arXiv:2412.17836
-
Optimized Quran Passage Retrieval Using an Expanded QA Dataset and Fine-Tuned Language Models 16 Dec 2024 · 0 repositories · arXiv:2412.11431
-
RAG Playground: A Framework for Systematic Evaluation of Retrieval Strategies and Prompt Engineering in RAG Systems 16 Dec 2024 · 1 repository · arXiv:2412.12322
-
Unanswerability Evaluation for Retrieval Augmented Generation 16 Dec 2024 · 0 repositories · arXiv:2412.12300
-
A Contextualized BERT model for Knowledge Graph Completion 15 Dec 2024 · 0 repositories · arXiv:2412.11016
-
Accelerating Retrieval-Augmented Generation 14 Dec 2024 · 0 repositories · arXiv:2412.15246
-
Inference Scaling for Bridging Retrieval and Augmented Generation 14 Dec 2024 · 0 repositories · arXiv:2412.10684
-
Tokens, the oft-overlooked appetizer: Large language models, the distributional hypothesis, and meaning 14 Dec 2024 · 0 repositories · arXiv:2412.10924
-
VisDoM: Multi-Document QA with Visually Rich Elements Using Multimodal Retrieval-Augmented Generation 14 Dec 2024 · 0 repositories · arXiv:2412.10704
-
Evidence Contextualization and Counterfactual Attribution for Conversational QA over Heterogeneous Data with RAG Systems 13 Dec 2024 · 0 repositories · arXiv:2412.10571
-
RAGServe: Fast Quality-Aware RAG Systems with Configuration Adaptation 13 Dec 2024 · 0 repositories · arXiv:2412.10543
-
VLR-Bench: Multilingual Benchmark Dataset for Vision-Language Retrieval Augmented Generation 13 Dec 2024 · 0 repositories · arXiv:2412.10151
-
Assessing the Robustness of Retrieval-Augmented Generation Systems in K-12 Educational Question Answering with Knowledge Discrepancies 12 Dec 2024 · 0 repositories · arXiv:2412.08985
-
Context Canvas: Enhancing Text-to-Image Diffusion Models with Knowledge Graph-Based RAG 12 Dec 2024 · 0 repositories · arXiv:2412.09614
-
OG-RAG: Ontology-Grounded Retrieval-Augmented Generation For Large Language Models 12 Dec 2024 · 0 repositories · arXiv:2412.15235
-
Accurate Medical Named Entity Recognition Through Specialized NLP Models 11 Dec 2024 · 0 repositories · arXiv:2412.08255
-
Advancing Single and Multi-task Text Classification through Large Language Model Fine-tuning 11 Dec 2024 · 0 repositories · arXiv:2412.08587
-
Leveraging Graph-RAG and Prompt Engineering to Enhance LLM-Based Automated Requirement Traceability and Compliance Checks 11 Dec 2024 · 0 repositories · arXiv:2412.08593
-
NLPineers@ NLU of Devanagari Script Languages 2025: Hate Speech Detection using Ensembling of BERT-based models 11 Dec 2024 · 2 repositories · arXiv:2412.08163
-
Adapting to Non-Stationary Environments: Multi-Armed Bandit Enhanced Retrieval-Augmented Generation on Knowledge Graphs 10 Dec 2024 · 1 repository · arXiv:2412.07618Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 1 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Bumblebee: Foundation Model for Particle Physics Discovery 10 Dec 2024 · 0 repositories · arXiv:2412.07867
-
Can linguists better understand DNA? 10 Dec 2024 · 1 repository · arXiv:2412.07678
-
Generating Knowledge Graphs from Large Language Models: A Comparative Study of GPT-4, LLaMA 2, and BERT 10 Dec 2024 · 0 repositories · arXiv:2412.07412
-
LLM as HPC Expert: Extending RAG Architecture for HPC Data 9 Dec 2024 · 0 repositories · arXiv:2501.14733
-
Optimizing Multi-Task Learning for Enhanced Performance in Large Language Models 9 Dec 2024 · 0 repositories · arXiv:2412.06249
-
SiReRAG: Indexing Similar and Related Information for Multihop Reasoning 9 Dec 2024 · 0 repositories · arXiv:2412.06206Syntology 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 1 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
The Rosetta Paradox: Domain-Specific Performance Inversions in Large Language Models 9 Dec 2024 · 0 repositories · arXiv:2412.17821
-
Unseen Attack Detection in Software-Defined Networking Using a BERT-Based Large Language Model 9 Dec 2024 · 0 repositories · arXiv:2412.06239
-
A Collaborative Multi-Agent Approach to Retrieval-Augmented Generation Across Diverse Data 8 Dec 2024 · 0 repositories · arXiv:2412.05838
-
Mixture-of-PageRanks: Replacing Long-Context with Real-Time, Sparse GraphRAG 8 Dec 2024 · 0 repositories · arXiv:2412.06078
-
BERTCaps: BERT Capsule for Persian Multi-Domain Sentiment Analysis 7 Dec 2024 · 0 repositories · arXiv:2412.05591
-
KG-Retriever: Efficient Knowledge Indexing for Retrieval-Augmented Large Language Models 7 Dec 2024 · 1 repository · arXiv:2412.05547Syntology official: harvested, nothing ran · 0 ran · 3 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Shifting NER into High Gear: The Auto-AdvER Approach 7 Dec 2024 · 0 repositories · arXiv:2412.05655
-
SLA Management in Reconfigurable Multi-Agent RAG: A Systems Approach to Question Answering 7 Dec 2024 · 0 repositories · arXiv:2412.06832
-
TOBUGraph: Knowledge Graph-Based Retrieval for Enhanced LLM Performance Beyond RAG 6 Dec 2024 · 0 repositories · arXiv:2412.05447
-
Enhancing Cross-Language Code Translation via Task-Specific Embedding Alignment in Retrieval-Augmented Generation 6 Dec 2024 · 0 repositories · arXiv:2412.05159
-
NLP-ADBench: NLP Anomaly Detection Benchmark 6 Dec 2024 · 1 repository · arXiv:2412.04784
-
Privacy-Preserving Retrieval-Augmented Generation with Differential Privacy 6 Dec 2024 · 0 repositories · arXiv:2412.04697
-
QueEn: A Large Language Model for Quechua-English Translation 6 Dec 2024 · 0 repositories · arXiv:2412.05184
-
Addressing Hallucinations with RAG and NMISS in Italian Healthcare LLM Chatbots 5 Dec 2024 · 0 repositories · arXiv:2412.04235
-
Comprehensive Audio Query Handling System with Integrated Expert Models and Contextual Understanding 5 Dec 2024 · 0 repositories · arXiv:2412.03980
-
Exploring AI Text Generation, Retrieval-Augmented Generation, and Detection Technologies: a Comprehensive Overview 5 Dec 2024 · 0 repositories · arXiv:2412.03933
-
HEAL: Hierarchical Embedding Alignment Loss for Improved Retrieval and Representation Learning 5 Dec 2024 · 1 repository · arXiv:2412.04661
-
Uniform Discretized Integrated Gradients: An effective attribution based method for explaining large language models 5 Dec 2024 · 0 repositories · arXiv:2412.03886
-
Advancing Conversational Psychotherapy: Integrating Privacy, Dual-Memory, and Domain Expertise with Large Language Models 4 Dec 2024 · 0 repositories · arXiv:2412.02987
-
FANAL -- Financial Activity News Alerting Language Modeling Framework 4 Dec 2024 · 0 repositories · arXiv:2412.03527
-
Multimodal Sentiment Analysis Based on BERT and ResNet 4 Dec 2024 · 0 repositories · arXiv:2412.03625
-
Achieving Semantic Consistency: Contextualized Word Representations for Political Text Analysis 3 Dec 2024 · 0 repositories · arXiv:2412.04505
-
CAISSON: Concept-Augmented Inference Suite of Self-Organizing Neural Networks 3 Dec 2024 · 0 repositories · arXiv:2412.02835
-
CPTQuant -- A Novel Mixed Precision Post-Training Quantization Techniques for Large Language Models 3 Dec 2024 · 0 repositories · arXiv:2412.03599
-
Gracefully Filtering Backdoor Samples for Generative Large Language Models without Retraining 3 Dec 2024 · 1 repository · arXiv:2412.02454
-
OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation 3 Dec 2024 · 1 repository · arXiv:2412.02592