Methods › Natural Language Processing › Subword Segmentation › WordPiece › Papers, page 5
WordPiece
Papers archive 2025-07-28
archive papers tagged: 7,063 · with a code link: 2,910 · where Syntology ran a sample: 650 (529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (650 of 7,063 tagged: 529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument)
Page 5 of 71: papers 401 to 500 of 7,063, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Towards Lighter and Robust Evaluation for Retrieval Augmented Generation 20 Mar 2025 · 1 repository · arXiv:2503.16161Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Tuning LLMs by RAG Principles: Towards LLM-native Memory 20 Mar 2025 · 1 repository · arXiv:2503.16071
-
Typed-RAG: Type-aware Multi-Aspect Decomposition for Non-Factoid Question Answering 20 Mar 2025 · 1 repository · arXiv:2503.15879
-
Enhancing Pancreatic Cancer Staging with Large Language Models: The Role of Retrieval-Augmented Generation 19 Mar 2025 · 0 repositories · arXiv:2503.15664
-
Bias Evaluation and Mitigation in Retrieval-Augmented Medical Question-Answering Systems 19 Mar 2025 · 0 repositories · arXiv:2503.15454
-
Optimizing Retrieval Strategies for Financial Question Answering Documents in Retrieval-Augmented Generation Systems 19 Mar 2025 · 1 repository · arXiv:2503.15191Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
RAG-based User Profiling for Precision Planning in Mixed-precision Over-the-Air Federated Learning 19 Mar 2025 · 0 repositories · arXiv:2503.15569
-
Shushing! Let's Imagine an Authentic Speech from the Silent Video 19 Mar 2025 · 0 repositories · arXiv:2503.14928
-
Enhancing LLM Generation with Knowledge Hypergraph for Evidence-Based Medicine 18 Mar 2025 · 0 repositories · arXiv:2503.16530
-
Good/Evil Reputation Judgment of Celebrities by LLMs via Retrieval Augmented Generation 18 Mar 2025 · 0 repositories · arXiv:2503.14382
-
JuDGE: Benchmarking Judgment Document Generation for Chinese Legal System 18 Mar 2025 · 1 repository · arXiv:2503.14258
-
Beyond Single Pass, Looping Through Time: KG-IRAG with Iterative Knowledge Retrieval 18 Mar 2025 · 0 repositories · arXiv:2503.14234
-
MDocAgent: A Multi-Modal Multi-Agent Framework for Document Understanding 18 Mar 2025 · 1 repository · arXiv:2503.13964Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
MoK-RAG: Mixture of Knowledge Paths Enhanced Retrieval-Augmented Generation for Embodied AI Environments 18 Mar 2025 · 0 repositories · arXiv:2503.13882
-
Predicting Human Choice Between Textually Described Lotteries 18 Mar 2025 · 0 repositories · arXiv:2503.14004
-
RAGO: Systematic Performance Optimization for Retrieval-Augmented Generation Serving 18 Mar 2025 · 1 repository · arXiv:2503.14649
-
MES-RAG: Bringing Multi-modal, Entity-Storage, and Secure Enhancements to RAG 17 Mar 2025 · 1 repository · arXiv:2503.13563
-
OSCAR: Online Soft Compression And Reranking 17 Mar 2025 · 0 repositories · arXiv:2504.07109
-
PAUSE: Low-Latency and Privacy-Aware Active User Selection for Federated Learning 17 Mar 2025 · 1 repository · arXiv:2503.13173
-
Privacy-Aware RAG: Secure and Isolated Knowledge Retrieval 17 Mar 2025 · 0 repositories · arXiv:2503.15548
-
GraphEval: A Lightweight Graph-Based LLM Framework for Idea Evaluation 16 Mar 2025 · 0 repositories · arXiv:2503.12600
-
Semantic Matters: Multimodal Features for Affective Analysis 16 Mar 2025 · 0 repositories · arXiv:2504.11460
-
Integrating Chain-of-Thought and Retrieval Augmented Generation Enhances Rare Disease Diagnosis from Clinical Notes 15 Mar 2025 · 0 repositories · arXiv:2503.12286
-
Language Models for Automated Classification of Brain MRI Reports and Growth Chart Generation 15 Mar 2025 · 0 repositories · arXiv:2503.12143
-
RAG-KG-IL: A Multi-Agent Hybrid Framework for Reducing Hallucinations and Enhancing LLM Reasoning through RAG and Incremental Knowledge Graph Learning Integration 14 Mar 2025 · 0 repositories · arXiv:2503.13514
-
Relevance Isn't All You Need: Scaling RAG Systems With Inference-Time Compute Via Multi-Criteria Reranking 14 Mar 2025 · 2 repositories · arXiv:2504.07104
-
Semantic and Contextual Modeling for Malicious Comment Detection with BERT-BiLSTM 14 Mar 2025 · 0 repositories · arXiv:2503.11084
-
ARLED: Leveraging LED-based ARMAN Model for Abstractive Summarization of Persian Long Documents 13 Mar 2025 · 0 repositories · arXiv:2503.10233
-
AttentionRAG: Attention-Guided Context Pruning in Retrieval-Augmented Generation 13 Mar 2025 · 0 repositories · arXiv:2503.10720
-
Cognitive-Mental-LLM: Evaluating Reasoning in Large Language Models for Mental Health Prediction via Online Text 13 Mar 2025 · 1 repository · arXiv:2503.10095
-
FG-RAG: Enhancing Query-Focused Summarization with Context-Aware Fine-Grained Graph RAG 13 Mar 2025 · 1 repository · arXiv:2504.07103
-
Predicting Stock Movement with BERTweet and Transformers 13 Mar 2025 · 0 repositories · arXiv:2503.10957
-
Retrieval-Augmented Generation with Hierarchical Knowledge 13 Mar 2025 · 1 repository · arXiv:2503.10150Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples)
-
Taxonomic Reasoning for Rare Arthropods: Combining Dense Image Captioning and RAG for Interpretable Classification 13 Mar 2025 · 0 repositories · arXiv:2503.10886
-
An Evaluation of LLMs for Detecting Harmful Computing Terms 12 Mar 2025 · 0 repositories · arXiv:2503.09341
-
ClaimTrust: Propagation Trust Scoring for RAG Systems 12 Mar 2025 · 0 repositories · arXiv:2503.10702
-
Conversational Gold: Evaluating Personalized Conversational Search System using Gold Nuggets 12 Mar 2025 · 1 repository · arXiv:2503.09902
-
A Global Dataset Mapping the AI Innovation from Academic Research to Industrial Patents 12 Mar 2025 · 0 repositories · arXiv:2503.09257
-
How to Protect Yourself from 5G Radiation? Investigating LLM Responses to Implicit Misinformation 12 Mar 2025 · 1 repository · arXiv:2503.09598
-
Memory-enhanced Retrieval Augmentation for Long Video Understanding 12 Mar 2025 · 0 repositories · arXiv:2503.09149
-
MoC: Mixtures of Text Chunking Learners for Retrieval-Augmented Generation System 12 Mar 2025 · 1 repository · arXiv:2503.09600
-
Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning 12 Mar 2025 · 3 repositories · arXiv:2503.09516Syntology official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 1 pointer-only (licence)
-
A Grey-box Text Attack Framework using Explainable AI 11 Mar 2025 · 0 repositories · arXiv:2503.08226
-
A Survey on Knowledge-Oriented Retrieval-Augmented Generation 11 Mar 2025 · 0 repositories · arXiv:2503.10677
-
ESNLIR: A Spanish Multi-Genre Dataset with Causal Relationships 11 Mar 2025 · 0 repositories · arXiv:2503.08803
-
LLM-based Corroborating and Refuting Evidence Retrieval for Scientific Claim Verification 11 Mar 2025 · 0 repositories · arXiv:2503.07937
-
OpenRAG: Optimizing RAG End-to-End via In-Context Retrieval Learning 11 Mar 2025 · 0 repositories · arXiv:2503.08398
-
A LongFormer-Based Framework for Accurate and Efficient Medical Text Summarization 10 Mar 2025 · 0 repositories · arXiv:2503.06888
-
CtrlRAG: Black-box Adversarial Attacks Based on Masked Language Models in Retrieval-Augmented Language Generation 10 Mar 2025 · 0 repositories · arXiv:2503.06950
-
Talking to GDELT Through Knowledge Graphs 10 Mar 2025 · 0 repositories · arXiv:2503.07584
-
Human Cognition Inspired RAG with Knowledge Graph for Complex Problem Solving 9 Mar 2025 · 0 repositories · arXiv:2503.06567
-
Multimodal Emotion Recognition and Sentiment Analysis in Multi-Party Conversation Contexts 9 Mar 2025 · 0 repositories · arXiv:2503.06805
-
Seeing Delta Parameters as JPEG Images: Data-Free Delta Compression with Discrete Cosine Transform 9 Mar 2025 · 0 repositories · arXiv:2503.06676
-
Breaking Free from MMI: A New Frontier in Rationalization by Probing Input Utilization 8 Mar 2025 · 1 repository · arXiv:2503.06202Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 1 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Constructions are Revealed in Word Distributions 8 Mar 2025 · 1 repository · arXiv:2503.06048
-
Poisoned-MRAG: Knowledge Poisoning Attacks to Multimodal Retrieval Augmented Generation 8 Mar 2025 · 0 repositories · arXiv:2503.06254
-
Evaluating open-source Large Language Models for automated fact-checking 7 Mar 2025 · 0 repositories · arXiv:2503.05565
-
FMT:A Multimodal Pneumonia Detection Model Based on Stacking MOE Framework 7 Mar 2025 · 0 repositories · arXiv:2503.05626
-
Leveraging Approximate Caching for Faster Retrieval-Augmented Generation 7 Mar 2025 · 0 repositories · arXiv:2503.05530
-
Leveraging Semantic Type Dependencies for Clinical Named Entity Recognition 7 Mar 2025 · 0 repositories · arXiv:2503.05373
-
Quantifying the Robustness of Retrieval-Augmented Language Models Against Spurious Features in Grounding Data 7 Mar 2025 · 0 repositories · arXiv:2503.05587
-
R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning 7 Mar 2025 · 5 repositories · arXiv:2503.05592Syntology 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Beyond RAG: Task-Aware KV Cache Compression for Comprehensive Knowledge Reasoning 6 Mar 2025 · 0 repositories · arXiv:2503.04973
-
Collapse of Dense Retrievers: Short, Early, and Literal Biases Outranking Factual Evidence 6 Mar 2025 · 0 repositories · arXiv:2503.05037
-
In-depth Analysis of Graph-based RAG in a Unified Framework 6 Mar 2025 · 0 repositories · arXiv:2503.04338
-
Incentivizing Multi-Tenant Split Federated Learning for Foundation Models at the Network Edge 6 Mar 2025 · 0 repositories · arXiv:2503.04971
-
Can Frontier LLMs Replace Annotators in Biomedical Text Mining? Analyzing Challenges and Exploring Solutions 5 Mar 2025 · 1 repository · arXiv:2503.03261
-
Intermediate-Task Transfer Learning: Leveraging Sarcasm Detection for Stance Detection 5 Mar 2025 · 0 repositories · arXiv:2503.03172
-
Large language models in finance : what is financial sentiment? 5 Mar 2025 · 0 repositories · arXiv:2503.03612
-
Personalized Federated Fine-tuning for Heterogeneous Data: An Automatic Rank Learning Approach via Two-Level LoRA 5 Mar 2025 · 0 repositories · arXiv:2503.03920
-
Sarcasm Detection as a Catalyst: Improving Stance Detection with Cross-Target Capabilities 5 Mar 2025 · 0 repositories · arXiv:2503.03787
-
The Box is in the Pen: Evaluating Commonsense Reasoning in Neural Machine Translation 5 Mar 2025 · 1 repository · arXiv:2503.03308
-
LLaVE: Large Language and Vision Embedding Models with Hardness-Weighted Contrastive Learning 4 Mar 2025 · 0 repositories · arXiv:2503.04812
-
Optimizing open-domain question answering with graph-based retrieval augmented generation 4 Mar 2025 · 0 repositories · arXiv:2503.02922
-
PennyLang: Pioneering LLM-Based Quantum Code Generation with a Novel PennyLane-Centric Dataset 4 Mar 2025 · 0 repositories · arXiv:2503.02497
-
Wikipedia in the Era of LLMs: Evolution and Risks 4 Mar 2025 · 1 repository · arXiv:2503.02879
-
Enhancing Social Media Rumor Detection: A Semantic and Graph Neural Network Approach for the 2024 Global Election 3 Mar 2025 · 0 repositories · arXiv:2503.01394
-
SAGE: A Framework of Precise Retrieval for RAG 3 Mar 2025 · 0 repositories · arXiv:2503.01713
-
Boolean-aware Attention for Dense Retrieval 3 Mar 2025 · 0 repositories · arXiv:2503.01753
-
Efficient or Powerful? Trade-offs Between Machine Learning and Deep Learning for Mental Illness Detection on Social Media 3 Mar 2025 · 0 repositories · arXiv:2503.01082
-
HoH: A Dynamic Benchmark for Evaluating the Impact of Outdated Information on Retrieval-Augmented Generation 3 Mar 2025 · 0 repositories · arXiv:2503.04800Syntology 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Retrieval-Augmented Perception: High-Resolution Image Perception Meets Visual RAG 3 Mar 2025 · 1 repository · arXiv:2503.01222
-
SePer: Measure Retrieval Utility Through The Lens Of Semantic Perplexity Reduction 3 Mar 2025 · 1 repository · arXiv:2503.01478Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples)
-
SRAG: Structured Retrieval-Augmented Generation for Multi-Entity Question Answering over Wikipedia Graph 3 Mar 2025 · 0 repositories · arXiv:2503.01346
-
ER-RAG: Enhance RAG with ER-Based Unified Modeling of Heterogeneous Data Sources 2 Mar 2025 · 0 repositories · arXiv:2504.06271
-
Optimizing Multi-Hop Document Retrieval Through Intermediate Representations 2 Mar 2025 · 0 repositories · arXiv:2503.04796
-
Pseudo-Knowledge Graph: Meta-Path Guided Retrieval and In-Graph Text for RAG-Equipped LLM 1 Mar 2025 · 0 repositories · arXiv:2503.00309
-
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge 1 Mar 2025 · 0 repositories · arXiv:2503.00596
-
Hierarchical Multi-Stage BERT Fusion Framework with Dual Attention for Enhanced Cyberbullying Detection in Social Media 1 Mar 2025 · 0 repositories · arXiv:2503.00342
-
U-NIAH: Unified RAG and LLM Evaluation for Long Context Needle-In-A-Haystack 1 Mar 2025 · 1 repository · arXiv:2503.00353
-
SafeAuto: Knowledge-Enhanced Safe Autonomous Driving with Multimodal Foundation Models 28 Feb 2025 · 1 repository · arXiv:2503.00211
-
Ext2Gen: Alignment through Unified Extraction and Generation for Robust Retrieval-Augmented Generation 28 Feb 2025 · 0 repositories · arXiv:2503.04789
-
LexRAG: Benchmarking Retrieval-Augmented Generation in Multi-Turn Legal Consultation Conversation 28 Feb 2025 · 1 repository · arXiv:2502.20640
-
Retrieval Augmented Generation for Topic Modeling in Organizational Research: An Introduction with Empirical Demonstration 28 Feb 2025 · 0 repositories · arXiv:2502.20963
-
RuCCoD: Towards Automated ICD Coding in Russian 28 Feb 2025 · 1 repository · arXiv:2502.21263
-
SuperRAG: Beyond RAG with Layout-Aware Graph Modeling 28 Feb 2025 · 0 repositories · arXiv:2503.04790
-
TeleRAG: Efficient Retrieval-Augmented Generation Inference with Lookahead Retrieval 28 Feb 2025 · 0 repositories · arXiv:2502.20969
-
The RAG Paradox: A Black-Box Attack Exploiting Unintentional Vulnerabilities in Retrieval-Augmented Generation Systems 28 Feb 2025 · 0 repositories · arXiv:2502.20995
-
Advanced Deep Learning Techniques for Analyzing Earnings Call Transcripts: Methodologies and Applications 27 Feb 2025 · 0 repositories · arXiv:2503.01886
-
An exploration of features to improve the generalisability of fake news detection models 27 Feb 2025 · 0 repositories · arXiv:2502.20299