Methods › Natural Language Processing › Subword Segmentation › WordPiece › Papers, page 2
WordPiece
Papers archive 2025-07-28
archive papers tagged: 7,063 · with a code link: 2,910 · where Syntology ran a sample: 650 (529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (650 of 7,063 tagged: 529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument)
Page 2 of 71: papers 101 to 200 of 7,063, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Augmenting LLM Reasoning with Dynamic Notes Writing for Complex QA 22 May 2025 · 0 repositories · arXiv:2505.16293
-
Chain-of-Thought Poisoning Attacks against R1-based Retrieval-Augmented Generation Systems 22 May 2025 · 0 repositories · arXiv:2505.16367
-
Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering Tasks 22 May 2025 · 0 repositories · arXiv:2505.16901
-
Comparative analysis of subword tokenization approaches for Indian languages 22 May 2025 · 0 repositories · arXiv:2505.16868
-
DailyQA: A Benchmark to Evaluate Web Retrieval Augmented LLMs Based on Capturing Real-World Changes 22 May 2025 · 0 repositories · arXiv:2505.17162
-
Multimodal Generative AI for Story Point Estimation in Software Development 22 May 2025 · 0 repositories · arXiv:2505.16290
-
Personalizing Student-Agent Interactions Using Log-Contextualized Retrieval Augmented Generation (RAG) 22 May 2025 · 0 repositories · arXiv:2505.17238
-
R1-Searcher++: Incentivizing the Dynamic Knowledge Acquisition of LLMs via Reinforcement Learning 22 May 2025 · 3 repositories · arXiv:2505.17005Syntology community repositories only · 9 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 10 harvested samples)
-
Search Wisely: Mitigating Sub-optimal Agentic Searches By Reducing Uncertainty 22 May 2025 · 0 repositories · arXiv:2505.17281
-
VoxRAG: A Step Toward Transcription-Free RAG Systems in Spoken Question Answering 22 May 2025 · 0 repositories · arXiv:2505.17326
-
Walk&Retrieve: Simple Yet Effective Zero-shot Retrieval-Augmented Generation via Knowledge Graph Walks 22 May 2025 · 1 repository · arXiv:2505.16849
-
AdUE: Improving uncertainty estimation head for LoRA adapters in LLMs 21 May 2025 · 0 repositories · arXiv:2505.15443
-
BR-TaxQA-R: A Dataset for Question Answering with References for Brazilian Personal Income Tax Law, including case law 21 May 2025 · 0 repositories · arXiv:2505.15916
-
Diffusion vs. Autoregressive Language Models: A Text Embedding Perspective 21 May 2025 · 0 repositories · arXiv:2505.15045
-
HDLxGraph: Bridging Large Language Models and HDL Repositories via HDL Graph Databases 21 May 2025 · 1 repository · arXiv:2505.15701
-
InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation 21 May 2025 · 0 repositories · arXiv:2505.15872
-
Leveraging Foundation Models for Multimodal Graph-Based Action Recognition 21 May 2025 · 0 repositories · arXiv:2505.15192
-
MaxPoolBERT: Enhancing BERT Classification via Layer- and Token-Wise Aggregation 21 May 2025 · 0 repositories · arXiv:2505.15696
-
Ranking Free RAG: Replacing Re-ranking with Selection in RAG for Sensitive Domains 21 May 2025 · 0 repositories · arXiv:2505.16014
-
Reranking with Compressed Document Representation 21 May 2025 · 0 repositories · arXiv:2505.15394
-
Silent Leaks: Implicit Knowledge Extraction Attack on RAG Systems through Benign Queries 21 May 2025 · 1 repository · arXiv:2505.15420
-
Single LLM, Multiple Roles: A Unified Retrieval-Augmented Generation Framework Using Role-Specific Token Optimization 21 May 2025 · 0 repositories · arXiv:2505.15444
-
Automatic Dataset Generation for Knowledge Intensive Question Answering Tasks 20 May 2025 · 0 repositories · arXiv:2505.14212
-
Beyond Text: Unveiling Privacy Vulnerabilities in Multi-modal Retrieval-Augmented Generation 20 May 2025 · 0 repositories · arXiv:2505.13957
-
Divide by Question, Conquer by Agent: SPLIT-RAG with Question-Driven Graph Partitioning 20 May 2025 · 0 repositories · arXiv:2505.13994
-
Enhancing Abstractive Summarization of Scientific Papers Using Structure Information 20 May 2025 · 1 repository · arXiv:2505.14179
-
Informatics for Food Processing 20 May 2025 · 1 repository · arXiv:2505.17087
-
Multimodal RAG-driven Anomaly Detection and Classification in Laser Powder Bed Fusion using Large Language Models 20 May 2025 · 0 repositories · arXiv:2505.13828
-
Probing BERT for German Compound Semantics 20 May 2025 · 0 repositories · arXiv:2505.14130
-
Process vs. Outcome Reward: Which is Better for Agentic RAG Reinforcement Learning 20 May 2025 · 1 repository · arXiv:2505.14069Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
s3: You Don't Need That Much Data to Train a Search Agent via RL 20 May 2025 · 1 repository · arXiv:2505.14146Syntology official (archive's flag): 7 ran · 7 ran (of which 3 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 4 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples)
-
SCAN: Semantic Document Layout Analysis for Textual and Visual Retrieval-Augmented Generation 20 May 2025 · 0 repositories · arXiv:2505.14381
-
Accelerating Adaptive Retrieval Augmented Generation via Instruction-Driven Representation Reduction of Retrieval Overlaps 19 May 2025 · 0 repositories · arXiv:2505.12731
-
AMAQA: A Metadata-based QA Dataset for RAG Systems 19 May 2025 · 0 repositories · arXiv:2505.13557
-
Climate Research Domain BERTs: Pretraining, Adaptation, and Evaluation 19 May 2025 · 0 repositories
-
EAVIT: Efficient and Accurate Human Value Identification from Text data via LLMs 19 May 2025 · 0 repositories · arXiv:2505.12792
-
Effective and Transparent RAG: Adaptive-Reward Reinforcement Learning for Decision Traceability 19 May 2025 · 1 repository · arXiv:2505.13258
-
Evaluating the Performance of RAG Methods for Conversational AI in the Airport Domain 19 May 2025 · 0 repositories · arXiv:2505.13006
-
Know Or Not: a library for evaluating out-of-knowledge base robustness 19 May 2025 · 1 repository · arXiv:2505.13545
-
Know3-RAG: A Knowledge-aware RAG Framework with Adaptive Retrieval, Generation, and Filtering 19 May 2025 · 1 repository · arXiv:2505.12662
-
Optimizing Retrieval Augmented Generation for Object Constraint Language 19 May 2025 · 0 repositories · arXiv:2505.13129
-
RAR: Setting Knowledge Tripwires for Retrieval Augmented Rejection 19 May 2025 · 0 repositories · arXiv:2505.13581
-
Suicide Risk Assessment Using Multimodal Speech Features: A Study on the SW1 Challenge Dataset 19 May 2025 · 0 repositories · arXiv:2505.13069
-
KGAlign: Joint Semantic-Structural Knowledge Encoding for Multimodal Fake News Detection 18 May 2025 · 1 repository · arXiv:2505.14714
-
PoisonArena: Uncovering Competing Poisoning Attacks in Retrieval-Augmented Generation 18 May 2025 · 1 repository · arXiv:2505.12574
-
RAGXplain: From Explainable Evaluation to Actionable Guidance of RAG Pipelines 18 May 2025 · 0 repositories · arXiv:2505.13538
-
ELITE: Embedding-Less retrieval with Iterative Text Exploration 17 May 2025 · 1 repository · arXiv:2505.11908
-
Let's have a chat with the EU AI Act 17 May 2025 · 0 repositories · arXiv:2505.11946
-
Neuro-Symbolic Query Compiler 17 May 2025 · 1 repository · arXiv:2505.11932
-
Telco-oRAG: Optimizing Retrieval-augmented Generation for Telecom Queries via Hybrid Retrieval and Neural Routing 17 May 2025 · 0 repositories · arXiv:2505.11856
-
Unveiling Knowledge Utilization Mechanisms in LLM-based Retrieval-Augmented Generation 17 May 2025 · 0 repositories · arXiv:2505.11995
-
SubGCache: Accelerating Graph-based RAG with Subgraph-level KV Cache 16 May 2025 · 0 repositories · arXiv:2505.10951
-
RAGSynth: Synthetic Data for Robust and Faithful RAG Component Optimization 16 May 2025 · 1 repository · arXiv:2505.10989
-
mmRAG: A Modular Benchmark for Retrieval-Augmented Generation over Text, Tables, and Knowledge Graphs 16 May 2025 · 1 repository · arXiv:2505.11180
-
Towards Cultural Bridge by Bahnaric-Vietnamese Translation Using Transfer Learning of Sequence-To-Sequence Pre-training Language Model 16 May 2025 · 0 repositories · arXiv:2505.11421
-
EcoSafeRAG: Efficient Security through Context Analysis in Retrieval-Augmented Generation 16 May 2025 · 0 repositories · arXiv:2505.13506
-
Finetune-RAG: Fine-Tuning Language Models to Resist Hallucination in Retrieval-Augmented Generation 16 May 2025 · 1 repository · arXiv:2505.10792
-
THELMA: Task Based Holistic Evaluation of Large Language Model Applications-RAG Question Answering 16 May 2025 · 0 repositories · arXiv:2505.11626
-
MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly 15 May 2025 · 1 repository · arXiv:2505.10610Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Artificial Intelligence Bias on English Language Learners in Automatic Scoring 15 May 2025 · 0 repositories · arXiv:2505.10643
-
AI Agents vs. Agentic AI: A Conceptual Taxonomy, Applications and Challenges 15 May 2025 · 0 repositories · arXiv:2505.10468
-
CAFE: Retrieval Head-based Coarse-to-Fine Information Seeking to Enhance Multi-Document QA Capability 15 May 2025 · 0 repositories · arXiv:2505.10063
-
CL-RAG: Bridging the Gap in Retrieval-Augmented Generation with Curriculum Learning 15 May 2025 · 0 repositories · arXiv:2505.10493
-
Hierarchical Document Refinement for Long-context Retrieval-augmented Generation 15 May 2025 · 1 repository · arXiv:2505.10413
-
Leveraging Graph Retrieval-Augmented Generation to Support Learners' Understanding of Knowledge Concepts in MOOCs 15 May 2025 · 0 repositories · arXiv:2505.10074
-
One Shot Dominance: Knowledge Poisoning Attack on Retrieval-Augmented Generation Systems 15 May 2025 · 0 repositories · arXiv:2505.11548
-
CXMArena: Unified Dataset to benchmark performance in realistic CXM Scenarios 14 May 2025 · 1 repository · arXiv:2505.09436
-
Multilingual Machine Translation with Quantum Encoder Decoder Attention-based Convolutional Variational Circuits 14 May 2025 · 0 repositories · arXiv:2505.09407
-
Enhancing Thyroid Cytology Diagnosis with RAG-Optimized LLMs and Pa-thology Foundation Models 13 May 2025 · 0 repositories · arXiv:2505.08590
-
Hakim: Farsi Text Embedding Model 13 May 2025 · 0 repositories · arXiv:2505.08435
-
IterKey: Iterative Keyword Generation with LLMs for Enhanced Retrieval Augmented Generation 13 May 2025 · 0 repositories · arXiv:2505.08450
-
Optimizing Retrieval-Augmented Generation: Analysis of Hyperparameter Impact on Performance and Efficiency 13 May 2025 · 0 repositories · arXiv:2505.08445
-
Scaling Context, Not Parameters: Training a Compact 7B Language Model for Efficient Long-Context Processing 13 May 2025 · 0 repositories · arXiv:2505.08651
-
Securing RAG: A Risk Assessment and Mitigation Framework 13 May 2025 · 0 repositories · arXiv:2505.08728
-
WixQA: A Multi-Dataset Benchmark for Enterprise Retrieval-Augmented Generation 13 May 2025 · 0 repositories · arXiv:2505.08643
-
Benchmarking Retrieval-Augmented Generation for Chemistry 12 May 2025 · 0 repositories · arXiv:2505.07671
-
Comparative sentiment analysis of public perception: Monkeypox vs. COVID-19 behavioral insights 12 May 2025 · 0 repositories · arXiv:2505.07430
-
DynamicRAG: Leveraging Outputs of Large Language Model as Feedback for Dynamic Reranking in Retrieval-Augmented Generation 12 May 2025 · 1 repository · arXiv:2505.07233
-
Efficient and Reproducible Biomedical Question Answering using Retrieval Augmented Generation 12 May 2025 · 1 repository · arXiv:2505.07917
-
HAMLET: Healthcare-focused Adaptive Multilingual Learning Embedding-based Topic Modeling 12 May 2025 · 0 repositories · arXiv:2505.07157
-
KAQG: A Knowledge-Graph-Enhanced RAG for Difficulty-Controlled Question Generation 12 May 2025 · 0 repositories · arXiv:2505.07618
-
Pre-training vs. Fine-tuning: A Reproducibility Study on Dense Retrieval Knowledge Acquisition 12 May 2025 · 1 repository · arXiv:2505.07166
-
SEReDeEP: Hallucination Detection in Retrieval-Augmented Models via Semantic Entropy and Context-Parameter Fusion 12 May 2025 · 0 repositories · arXiv:2505.07528
-
Towards Requirements Engineering for RAG Systems 12 May 2025 · 0 repositories · arXiv:2505.07553
-
Why Uncertainty Estimation Methods Fall Short in RAG: An Axiomatic Analysis 12 May 2025 · 0 repositories · arXiv:2505.07459
-
IM-BERT: Enhancing Robustness of BERT through the Implicit Euler Method 11 May 2025 · 0 repositories · arXiv:2505.06889
-
The Distracting Effect: Understanding Irrelevant Passages in RAG 11 May 2025 · 0 repositories · arXiv:2505.06914
-
MacRAG: Compress, Slice, and Scale-up for Multi-Scale Adaptive Context RAG 10 May 2025 · 1 repository · arXiv:2505.06569
-
OMGM: Orchestrate Multiple Granularities and Modalities for Efficient Multimodal Retrieval 10 May 2025 · 0 repositories · arXiv:2505.07879Syntology 4 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; every one of the 4 samples that ran constructed an object rather than computing a result (of 5 harvested samples) · 5 pointer-only (licence)
-
The Sound of Populism: Distinct Linguistic Features Across Populist Variants 10 May 2025 · 0 repositories · arXiv:2505.07874
-
Attention on Multiword Expressions: A Multilingual Study of BERT-based Models with Regard to Idiomaticity and Microsyntax 9 May 2025 · 1 repository · arXiv:2505.06062
-
Multimodal Integrated Knowledge Transfer to Large Language Models through Preference Optimization with Biomedical Applications 9 May 2025 · 1 repository · arXiv:2505.05736
-
AI Approaches to Qualitative and Quantitative News Analytics on NATO Unity 8 May 2025 · 0 repositories · arXiv:2505.06313
-
LiteLMGuard: Seamless and Lightweight On-Device Prompt Filtering for Safeguarding Small Language Models against Quantization-induced Risks and Vulnerabilities 8 May 2025 · 2 repositories · arXiv:2505.05619
-
Lost in OCR Translation? Vision-Based Approaches to Robust Document Retrieval 8 May 2025 · 0 repositories · arXiv:2505.05666
-
QualBench: Benchmarking Chinese LLMs with Localized Professional Qualifications for Vertical Domain Evaluation 8 May 2025 · 0 repositories · arXiv:2505.05225
-
Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards 7 May 2025 · 1 repository · arXiv:2505.04847Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample) · 1 pointer-only (licence)
-
Fine-Tuning Large Language Models and Evaluating Retrieval Methods for Improved Question Answering on Building Codes 7 May 2025 · 0 repositories · arXiv:2505.04666
-
Flower Across Time and Media: Sentiment Analysis of Tang Song Poetry and Visual Correspondence 7 May 2025 · 0 repositories · arXiv:2505.04785
-
HiPerRAG: High-Performance Retrieval Augmented Generation for Scientific Insights 7 May 2025 · 0 repositories · arXiv:2505.04846