Methods › Natural Language Processing › Subword Segmentation › WordPiece › Papers, page 3
WordPiece
Papers archive 2025-07-28
archive papers tagged: 7,063 · with a code link: 2,910 · where Syntology ran a sample: 650 (529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (650 of 7,063 tagged: 529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument)
Page 3 of 71: papers 201 to 300 of 7,063, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Osiris: A Lightweight Open-Source Hallucination Detection System 7 May 2025 · 0 repositories · arXiv:2505.04844
-
Personalized Risks and Regulatory Strategies of Large Language Models in Digital Advertising 7 May 2025 · 0 repositories · arXiv:2505.04665
-
Retrieval Augmented Generation Evaluation for Health Documents 7 May 2025 · 0 repositories · arXiv:2505.04680
-
A Reasoning-Focused Legal Retrieval Benchmark 6 May 2025 · 0 repositories · arXiv:2505.03970
-
An Analysis of Hyper-Parameter Optimization Methods for Retrieval Augmented Generation 6 May 2025 · 0 repositories · arXiv:2505.03452
-
Hesitation is defeat? Connecting Linguistic and Predictive Uncertainty 6 May 2025 · 0 repositories · arXiv:2505.03910
-
IndicSQuAD: A Comprehensive Multilingual Question Answering Dataset for Indic Languages 6 May 2025 · 1 repository · arXiv:2505.03688
-
Transformers for Learning on Noisy and Task-Level Manifolds: Approximation and Generalization Insights 6 May 2025 · 0 repositories · arXiv:2505.03205
-
Advancing Email Spam Detection: Leveraging Zero-Shot Learning and Large Language Models 5 May 2025 · 1 repository · arXiv:2505.02362
-
Automatic Proficiency Assessment in L2 English Learners 5 May 2025 · 0 repositories · arXiv:2505.02615
-
Direct Retrieval-augmented Optimization: Synergizing Knowledge Selection and Language Models 5 May 2025 · 1 repository · arXiv:2505.03075
-
Knowing You Don't Know: Learning When to Continue Search in Multi-round RAG through Self-Practicing 5 May 2025 · 1 repository · arXiv:2505.02811
-
Less is More: Efficient Weight Farcasting with 1-Layer Neural Network 5 May 2025 · 0 repositories · arXiv:2505.02714
-
Rapid yet accurate Tile-circuit and device modeling for Analog In-Memory Computing 5 May 2025 · 0 repositories · arXiv:2506.00004
-
SymbioticRAG: Enhancing Document Intelligence Through Human-LLM Symbiotic Collaboration 5 May 2025 · 0 repositories · arXiv:2505.02418
-
A New HOPE: Domain-agnostic Automatic Evaluation of Text Chunking 4 May 2025 · 0 repositories · arXiv:2505.02171
-
Adversarial Cooperative Rationalization: The Risk of Spurious Correlations in Even Clean Datasets 4 May 2025 · 1 repository · arXiv:2505.02118Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample) · 1 pointer-only (licence)
-
Exploring new Approaches for Information Retrieval through Natural Language Processing 4 May 2025 · 0 repositories · arXiv:2505.02199
-
Real-time Spatial Retrieval Augmented Generation for Urban Environments 4 May 2025 · 0 repositories · arXiv:2505.02271
-
Retrieval-augmented in-context learning for multimodal large language models in disease classification 4 May 2025 · 0 repositories · arXiv:2505.02087
-
Positional Attention for Efficient BERT-Based Named Entity Recognition 3 May 2025 · 0 repositories · arXiv:2505.01868
-
CHORUS: Zero-shot Hierarchical Retrieval and Orchestration for Generating Linear Programming Code 2 May 2025 · 0 repositories · arXiv:2505.01485
-
Retrieval-Augmented Generation in Biomedicine: A Survey of Technologies, Datasets, and Clinical Applications 2 May 2025 · 0 repositories · arXiv:2505.01146
-
CSE-SFP: Enabling Unsupervised Sentence Representation Learning via a Single Forward Pass 1 May 2025 · 1 repository · arXiv:2505.00389
-
EnronQA: Towards Personalized RAG over Private Documents 1 May 2025 · 0 repositories · arXiv:2505.00263
-
Patchwork: A Unified Framework for RAG Serving 1 May 2025 · 0 repositories · arXiv:2505.07833
-
LLM-Empowered Embodied Agent for Memory-Augmented Task Planning in Household Robotics 30 Apr 2025 · 1 repository · arXiv:2504.21716Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Talk Before You Retrieve: Agent-Led Discussions for Better RAG in Medical QA 30 Apr 2025 · 1 repository · arXiv:2504.21252
-
Traceback of Poisoning Attacks to Retrieval-Augmented Generation 30 Apr 2025 · 0 repositories · arXiv:2504.21668
-
ARCS: Agentic Retrieval-Augmented Code Synthesis with Iterative Refinement 29 Apr 2025 · 0 repositories · arXiv:2504.20434
-
BrightCookies at SemEval-2025 Task 9: Exploring Data Augmentation for Food Hazard Classification 29 Apr 2025 · 1 repository · arXiv:2504.20703
-
CBM-RAG: Demonstrating Enhanced Interpretability in Radiology Report Generation with Multi-Agent RAG and Concept Bottleneck Models 29 Apr 2025 · 1 repository · arXiv:2504.20898
-
Graph RAG for Legal Norms: A Hierarchical and Temporal Approach 29 Apr 2025 · 0 repositories · arXiv:2505.00039
-
ReasonIR: Training Retrievers for Reasoning Tasks 29 Apr 2025 · 1 repository · arXiv:2504.20595Syntology official: no sample here; runs from other or unrecorded repositories · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
UniversalRAG: Retrieval-Augmented Generation over Corpora of Diverse Modalities and Granularities 29 Apr 2025 · 1 repository · arXiv:2504.20734
-
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets 28 Apr 2025 · 0 repositories · arXiv:2504.20119
-
Chatbot Arena Meets Nuggets: Towards Explanations and Diagnostics in the Evaluation of LLM Responses 28 Apr 2025 · 0 repositories · arXiv:2504.20006
-
Reconstructing Context: Evaluating Advanced Chunking Strategies for Retrieval-Augmented Generation 28 Apr 2025 · 1 repository · arXiv:2504.19754
-
Security Bug Report Prediction Within and Across Projects: A Comparative Study of BERT and Random Forest 28 Apr 2025 · 0 repositories · arXiv:2504.21037
-
TreeHop: Generate and Filter Next Query Embeddings Efficiently for Multi-hop Question Answering 28 Apr 2025 · 1 repository · arXiv:2504.20114
-
Enhancing Speech-to-Speech Dialogue Modeling with End-to-End Retrieval-Augmented Generation 27 Apr 2025 · 1 repository · arXiv:2505.00028
-
The Influence of Text Variation on User Engagement in Cross-Platform Content Sharing 26 Apr 2025 · 0 repositories · arXiv:2505.03769
-
A model and package for German ColBERT 25 Apr 2025 · 0 repositories · arXiv:2504.20083
-
SMARTFinRAG: Interactive Modularized Financial RAG Benchmark 25 Apr 2025 · 1 repository · arXiv:2504.18024
-
A RAG-Based Multi-Agent LLM System for Natural Hazard Resilience and Adaptation 24 Apr 2025 · 1 repository · arXiv:2504.17200
-
FinBERT-QA: Financial Question Answering with pre-trained BERT Language Models 24 Apr 2025 · 1 repository · arXiv:2505.00725
-
A Novel Graph Transformer Framework for Gene Regulatory Network Inference 23 Apr 2025 · 0 repositories · arXiv:2504.16961
-
How Effective are Generative Large Language Models in Performing Requirements Classification? 23 Apr 2025 · 0 repositories · arXiv:2504.16768
-
Automated Bug Report Prioritization in Large Open-Source Projects 22 Apr 2025 · 1 repository · arXiv:2504.15912
-
CiteFix: Enhancing RAG Accuracy Through Post-Processing Citation Correction 22 Apr 2025 · 0 repositories · arXiv:2504.15629
-
FinDER: Financial Dataset for Question Answering and Evaluating Retrieval-Augmented Generation 22 Apr 2025 · 0 repositories · arXiv:2504.15800
-
Grounded in Context: Retrieval-Based Method for Hallucination Detection 22 Apr 2025 · 0 repositories · arXiv:2504.15771
-
Sentiment Analysis in Software Engineering: Evaluating Generative Pre-trained Transformers 22 Apr 2025 · 0 repositories · arXiv:2505.14692
-
Synergizing RAG and Reasoning: A Systematic Review 22 Apr 2025 · 0 repositories · arXiv:2504.15909
-
The Viability of Crowdsourcing for RAG Evaluation 22 Apr 2025 · 1 repository · arXiv:2504.15689
-
AlignRAG: Leveraging Critique Learning for Evidence-Sensitive Retrieval-Augmented Reasoning 21 Apr 2025 · 1 repository · arXiv:2504.14858Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Leveraging Language Models for Automated Patient Record Linkage 21 Apr 2025 · 0 repositories · arXiv:2504.15261
-
POLYRAG: Integrating Polyviews into Retrieval-Augmented Generation for Medical Applications 21 Apr 2025 · 0 repositories · arXiv:2504.14917
-
Retrieval Augmented Generation Evaluation in the Era of Large Language Models: A Comprehensive Survey 21 Apr 2025 · 1 repository · arXiv:2504.14891
-
Support Evaluation for the TREC 2024 RAG Track: Comparing Human versus LLM Judges 21 Apr 2025 · 0 repositories · arXiv:2504.15205
-
The Great Nugget Recall: Automating Fact Extraction and RAG Evaluation with Large Language Models 21 Apr 2025 · 0 repositories · arXiv:2504.15068
-
The Synthetic Imputation Approach: Generating Optimal Synthetic Texts For Underrepresented Categories In Supervised Classification Tasks 21 Apr 2025 · 0 repositories · arXiv:2504.15160
-
Disentangling Linguistic Features with Dimension-Wise Analysis of Vector Embeddings 20 Apr 2025 · 0 repositories · arXiv:2504.14766
-
FinSage: A Multi-aspect RAG System for Financial Filings Question Answering 20 Apr 2025 · 0 repositories · arXiv:2504.14493
-
Assessing AI-Generated Questions' Alignment with Cognitive Frameworks in Educational Assessment 19 Apr 2025 · 0 repositories · arXiv:2504.14232
-
LegalRAG: A Hybrid RAG System for Multilingual Legal Information Retrieval 19 Apr 2025 · 0 repositories · arXiv:2504.16121
-
CacheFormer: High Attention-Based Segment Caching 18 Apr 2025 · 0 repositories · arXiv:2504.13981
-
Contextual Embedding-based Clustering to Identify Topics for Healthcare Service Improvement 18 Apr 2025 · 0 repositories · arXiv:2504.14068
-
CoT-RAG: Integrating Chain of Thought and Retrieval-Augmented Generation to Enhance Reasoning in Large Language Models 18 Apr 2025 · 0 repositories · arXiv:2504.13534
-
RAG Without the Lag: Interactive Debugging for Retrieval-Augmented Generation Pipelines 18 Apr 2025 · 0 repositories · arXiv:2504.13587
-
Secure Multifaceted-RAG for Enterprise: Hybrid Knowledge Retrieval with Security Filtering 18 Apr 2025 · 0 repositories · arXiv:2504.13425
-
Word Embedding Techniques for Classification of Star Ratings 18 Apr 2025 · 0 repositories · arXiv:2504.13653
-
Accuracy is Not Agreement: Expert-Aligned Evaluation of Crash Narrative Classification Models 17 Apr 2025 · 0 repositories · arXiv:2504.13068
-
CDF-RAG: Causal Dynamic Feedback for Adaptive Retrieval-Augmented Generation 17 Apr 2025 · 1 repository · arXiv:2504.12560
-
Estimating Optimal Context Length for Hybrid Retrieval-augmented Multi-document Summarization 17 Apr 2025 · 1 repository · arXiv:2504.12972
-
FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents 17 Apr 2025 · 0 repositories · arXiv:2504.13128
-
InstructRAG: Leveraging Retrieval-Augmented Generation on Instruction Graphs for LLM-Based Task Planning 17 Apr 2025 · 0 repositories · arXiv:2504.13032
-
Retrieval-Augmented Generation with Conflicting Evidence 17 Apr 2025 · 1 repository · arXiv:2504.13079Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
A Visual RAG Pipeline for Few-Shot Fine-Grained Product Classification 16 Apr 2025 · 0 repositories · arXiv:2504.11838
-
ARCeR: an Agentic RAG for the Automated Definition of Cyber Ranges 16 Apr 2025 · 0 repositories · arXiv:2504.12143
-
Mapping Controversies Using Artificial Intelligence: An Analysis of the Hamas-Israel Conflict on YouTube 16 Apr 2025 · 0 repositories · arXiv:2504.12177
-
On the Feasibility of Using MultiModal LLMs to Execute AR Social Engineering Attacks 16 Apr 2025 · 0 repositories · arXiv:2504.13209
-
CSPLADE: Learned Sparse Retrieval with Causal Language Models 15 Apr 2025 · 0 repositories · arXiv:2504.10816
-
Efficient Distributed Retrieval-Augmented Generation for Enhancing Language Model Performance 15 Apr 2025 · 0 repositories · arXiv:2504.11197
-
Exploring the Role of Knowledge Graph-Based RAG in Japanese Medical Question Answering with Small-Scale LLMs 15 Apr 2025 · 0 repositories · arXiv:2504.10982
-
Intraoperative perfusion assessment by continuous, low-latency hyperspectral light-field imaging: development, methodology, and clinical application 15 Apr 2025 · 0 repositories · arXiv:2504.10953
-
LayoutCoT: Unleashing the Deep Reasoning Potential of Large Language Models for Layout Generation 15 Apr 2025 · 0 repositories · arXiv:2504.10829
-
Towards Automated Safety Requirements Derivation Using Agent-based RAG 15 Apr 2025 · 0 repositories · arXiv:2504.11243
-
A Survey of Personalization: From RAG to Agent 14 Apr 2025 · 1 repository · arXiv:2504.10147
-
DioR: Adaptive Cognitive Detection and Contextual Retrieval Optimization for Dynamic Retrieval-Augmented Generation 14 Apr 2025 · 0 repositories · arXiv:2504.10198
-
Hallucination Detection in LLMs via Topological Divergence on Attention Graphs 14 Apr 2025 · 0 repositories · arXiv:2504.10063
-
MMKB-RAG: A Multi-Modal Knowledge-Based Retrieval-Augmented Generation Framework 14 Apr 2025 · 0 repositories · arXiv:2504.10074
-
RAKG:Document-level Retrieval Augmented Knowledge Graph Construction 14 Apr 2025 · 1 repository · arXiv:2504.09823
-
Understanding and Optimizing Multi-Stage AI Inference Pipelines 14 Apr 2025 · 0 repositories · arXiv:2504.09775
-
VDocRAG: Retrieval-Augmented Generation over Visually-Rich Documents 14 Apr 2025 · 0 repositories · arXiv:2504.09795
-
XY-Cut++: Advanced Layout Ordering via Hierarchical Mask Mechanism on a Novel Benchmark 14 Apr 2025 · 1 repository · arXiv:2504.10258
-
ControlNET: A Firewall for RAG-based LLM System 13 Apr 2025 · 0 repositories · arXiv:2504.09593
-
HD-RAG: Retrieval-Augmented Generation for Hybrid Documents Containing Text and Hierarchical Tables 13 Apr 2025 · 0 repositories · arXiv:2504.09554
-
HM-RAG: Hierarchical Multi-Agent Multimodal Retrieval Augmented Generation 13 Apr 2025 · 1 repository · arXiv:2504.12330
-
HeteRAG: A Heterogeneous Retrieval-augmented Generation Framework with Decoupled Knowledge Representations 12 Apr 2025 · 0 repositories · arXiv:2504.10529