Methods › Natural Language Processing › Subword Segmentation › WordPiece › Papers, page 6
WordPiece
Papers archive 2025-07-28
archive papers tagged: 7,063 · with a code link: 2,910 · where Syntology ran a sample: 650 (529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (650 of 7,063 tagged: 529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument)
Page 6 of 71: papers 501 to 600 of 7,063, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Bridging Legal Knowledge and AI: Retrieval-Augmented Generation with Vector Stores, Knowledge Graphs, and Hierarchical Non-negative Matrix Factorization 27 Feb 2025 · 1 repository · arXiv:2502.20364
-
How Much is Enough? The Diminishing Returns of Tokenization Training Data 27 Feb 2025 · 0 repositories · arXiv:2502.20273
-
Long-Context Inference with Retrieval-Augmented Speculative Decoding 27 Feb 2025 · 1 repository · arXiv:2502.20330
-
Efficient Federated Search for Retrieval-Augmented Generation 26 Feb 2025 · 0 repositories · arXiv:2502.19280
-
MEBench: Benchmarking Large Language Models for Cross-Document Multi-Entity Question Answering 26 Feb 2025 · 0 repositories · arXiv:2502.18993
-
NeoBERT: A Next-Generation BERT 26 Feb 2025 · 1 repository · arXiv:2502.19587
-
Broadening Discovery through Structural Models: Multimodal Combination of Local and Structural Properties for Predicting Chemical Features 25 Feb 2025 · 0 repositories · arXiv:2502.17986
-
Detecting Knowledge Boundary of Vision Large Language Models by Sampling-Based Inference 25 Feb 2025 · 0 repositories · arXiv:2502.18023
-
Enhancing Text Classification with a Novel Multi-Agent Collaboration Framework Leveraging BERT 25 Feb 2025 · 0 repositories · arXiv:2502.18653
-
Faster, Cheaper, Better: Multi-Objective Hyperparameter Optimization for LLM and RAG Systems 25 Feb 2025 · 0 repositories · arXiv:2502.18635
-
LevelRAG: Enhancing Retrieval-Augmented Generation with Multi-hop Logic Planning over Rewriting Augmented Searchers 25 Feb 2025 · 1 repository · arXiv:2502.18139
-
MM-PoisonRAG: Disrupting Multimodal RAG with Local and Global Poisoning Attacks 25 Feb 2025 · 1 repository · arXiv:2502.17832
-
MuCoS: Efficient Drug-Target Prediction through Multi-Context-Aware Sampling 25 Feb 2025 · 0 repositories · arXiv:2502.17784
-
Say Less, Mean More: Leveraging Pragmatics in Retrieval-Augmented Generation 25 Feb 2025 · 0 repositories · arXiv:2502.17839
-
ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents 25 Feb 2025 · 1 repository · arXiv:2502.18017Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
"Actionable Help" in Crises: A Novel Dataset and Resource-Efficient Models for Identifying Request and Offer Social Media Posts 24 Feb 2025 · 0 repositories · arXiv:2502.16839
-
Adversarial Training for Defense Against Label Poisoning Attacks 24 Feb 2025 · 1 repository · arXiv:2502.17121Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Applying LLMs to Active Learning: Towards Cost-Efficient Cross-Task Text Classification without Manually Labeled Data 24 Feb 2025 · 0 repositories · arXiv:2502.16892
-
Benchmarking Retrieval-Augmented Generation in Multi-Modal Contexts 24 Feb 2025 · 2 repositories · arXiv:2502.17297
-
Evaluating the Effect of Retrieval Augmentation on Social Biases 24 Feb 2025 · 0 repositories · arXiv:2502.17611
-
LettuceDetect: A Hallucination Detection Framework for RAG Applications 24 Feb 2025 · 2 repositories · arXiv:2502.17125
-
MEMERAG: A Multilingual End-to-End Meta-Evaluation Benchmark for Retrieval Augmented Generation 24 Feb 2025 · 1 repository · arXiv:2502.17163
-
Mitigating Bias in RAG: Controlling the Embedder 24 Feb 2025 · 1 repository · arXiv:2502.17390
-
Mutual Reinforcement of LLM Dialogue Synthesis and Summarization Capabilities for Few-Shot Dialogue Summarization 24 Feb 2025 · 0 repositories · arXiv:2502.17328
-
Towards Typologically Aware Rescoring to Mitigate Unfaithfulness in Lower-Resource Languages 24 Feb 2025 · 0 repositories · arXiv:2502.17664
-
D2S-FLOW: Automated Parameter Extraction from Datasheets for SPICE Model Generation Using Large Language Models 23 Feb 2025 · 0 repositories · arXiv:2502.16540
-
Code Summarization Beyond Function Level 23 Feb 2025 · 1 repository · arXiv:2502.16704
-
Layer-Wise Evolution of Representations in Fine-Tuned Transformers: Insights from Sparse AutoEncoders 23 Feb 2025 · 0 repositories · arXiv:2502.16722
-
Optimizing Retrieval-Augmented Generation of Medical Content for Spaced Repetition Learning 23 Feb 2025 · 0 repositories · arXiv:2503.01859
-
Retrieval-Augmented Visual Question Answering via Built-in Autoregressive Search Engines 23 Feb 2025 · 0 repositories · arXiv:2502.16641
-
Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries 23 Feb 2025 · 1 repository · arXiv:2502.16636
-
An End-to-End Homomorphically Encrypted Neural Network 22 Feb 2025 · 0 repositories · arXiv:2502.16176
-
Iterative Auto-Annotation for Scientific Named Entity Recognition Using BERT-Based Models 22 Feb 2025 · 0 repositories · arXiv:2502.16312
-
RAG-Enhanced Collaborative LLM Agents for Drug Discovery 22 Feb 2025 · 0 repositories · arXiv:2502.17506
-
Worse than Zero-shot? A Fact-Checking Dataset for Evaluating the Robustness of RAG Against Misleading Retrievals 22 Feb 2025 · 0 repositories · arXiv:2502.16101
-
A Close Look at Decomposition-based XAI-Methods for Transformer Language Models 21 Feb 2025 · 2 repositories · arXiv:2502.15886
-
Chain-of-Rank: Enhancing Large Language Models for Domain-Specific RAG in Edge Device 21 Feb 2025 · 0 repositories · arXiv:2502.15134
-
Cross-Format Retrieval-Augmented Generation in XR with LLMs for Context-Aware Maintenance Assistance 21 Feb 2025 · 0 repositories · arXiv:2502.15604
-
Enhancing Domain-Specific Retrieval-Augmented Generation: Synthetic Data Generation and Evaluation using Reasoning Models 21 Feb 2025 · 1 repository · arXiv:2502.15854
-
Extraction multi-étiquettes de relations en utilisant des couches de Transformer 21 Feb 2025 · 0 repositories · arXiv:2502.15619
-
Retrieval-Augmented Speech Recognition Approach for Domain Challenges 21 Feb 2025 · 0 repositories · arXiv:2502.15264
-
Robust Bias Detection in MLMs and its Application to Human Trait Ratings 21 Feb 2025 · 1 repository · arXiv:2502.15600
-
Tokenization is Sensitive to Language Variation 21 Feb 2025 · 0 repositories · arXiv:2502.15343
-
A Socratic RAG Approach to Connect Natural Language Queries on Research Topics with Knowledge Organization Systems 20 Feb 2025 · 0 repositories · arXiv:2502.15005
-
FIND: Fine-grained Information Density Guided Adaptive Retrieval-Augmented Generation for Disease Diagnosis 20 Feb 2025 · 0 repositories · arXiv:2502.14614
-
From RAG to Memory: Non-Parametric Continual Learning for Large Language Models 20 Feb 2025 · 1 repository · arXiv:2502.14802Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
Is Relevance Propagated from Retriever to Generator in RAG? 20 Feb 2025 · 0 repositories · arXiv:2502.15025
-
On the Influence of Context Size and Model Choice in Retrieval-Augmented Generation Systems 20 Feb 2025 · 1 repository · arXiv:2502.14759
-
PaperHelper: Knowledge-Based LLM QA Paper Reading Assistant 20 Feb 2025 · 0 repositories · arXiv:2502.14271
-
QUAD-LLM-MLTC: Large Language Models Ensemble Learning for Healthcare Text Multi-Label Classification 20 Feb 2025 · 0 repositories · arXiv:2502.14189
-
WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models 20 Feb 2025 · 0 repositories · arXiv:2502.14727
-
Are Large Language Models In-Context Graph Learners? 19 Feb 2025 · 0 repositories · arXiv:2502.13562
-
DH-RAG: A Dynamic Historical Context-Powered Retrieval-Augmented Generation Method for Multi-Turn Dialogue 19 Feb 2025 · 0 repositories · arXiv:2502.13847
-
Extracting Social Connections from Finnish Karelian Refugee Interviews Using LLMs 19 Feb 2025 · 0 repositories · arXiv:2502.13566
-
HawkBench: Investigating Resilience of RAG Methods on Stratified Information-Seeking Tasks 19 Feb 2025 · 0 repositories · arXiv:2502.13465
-
In-Place Updates of a Graph Index for Streaming Approximate Nearest Neighbor Search 19 Feb 2025 · 0 repositories · arXiv:2502.13826
-
RAG-Gym: Optimizing Reasoning and Search Agents with Process Supervision 19 Feb 2025 · 0 repositories · arXiv:2502.13957
-
RGAR: Recurrence Generation-augmented Retrieval for Factual-aware Medical Question Answering 19 Feb 2025 · 0 repositories · arXiv:2502.13361
-
Universal Semantic Embeddings of Chemical Elements for Enhanced Materials Inference and Discovery 19 Feb 2025 · 0 repositories · arXiv:2502.14912
-
What are Models Thinking about? Understanding Large Language Model Hallucinations "Psychology" through Model Inner State Analysis 19 Feb 2025 · 0 repositories · arXiv:2502.13490
-
HopRAG: Multi-Hop Reasoning for Logic-Aware Retrieval-Augmented Generation 18 Feb 2025 · 0 repositories · arXiv:2502.12442
-
Improving Clinical Question Answering with Multi-Task Learning: A Joint Approach for Answer Extraction and Medical Categorization 18 Feb 2025 · 0 repositories · arXiv:2502.13108
-
Oreo: A Plug-in Context Reconstructor to Enhance Retrieval-Augmented Generation 18 Feb 2025 · 0 repositories · arXiv:2502.13019
-
PathRAG: Pruning Graph-based Retrieval Augmented Generation with Relational Paths 18 Feb 2025 · 1 repository · arXiv:2502.14902Syntology official: no sample here; runs from other or unrecorded repositories · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
AdaSplash: Adaptive Sparse Flash Attention 17 Feb 2025 · 1 repository · arXiv:2502.12082Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Can LLMs Simulate Social Media Engagement? A Study on Action-Guided Response Generation 17 Feb 2025 · 0 repositories · arXiv:2502.12073
-
Does RAG Really Perform Bad For Long-Context Processing? 17 Feb 2025 · 0 repositories · arXiv:2502.11444
-
Fast or Better? Balancing Accuracy and Cost in Retrieval-Augmented Generation with Flexible User Control 17 Feb 2025 · 1 repository · arXiv:2502.12145
-
FineFilter: A Fine-grained Noise Filtering Mechanism for Retrieval-Augmented Large Language Models 17 Feb 2025 · 0 repositories · arXiv:2502.11811
-
RAG vs. GraphRAG: A Systematic Evaluation and Key Insights 17 Feb 2025 · 0 repositories · arXiv:2502.11371
-
REAL-MM-RAG: A Real-World Multi-Modal Retrieval Benchmark 17 Feb 2025 · 0 repositories · arXiv:2502.12342
-
Market-Derived Financial Sentiment Analysis: Context-Aware Language Models for Crypto Forecasting 17 Feb 2025 · 1 repository · arXiv:2502.14897
-
Revisiting Robust RAG: Do We Still Need Complex Robust Training in the Era of Powerful LLMs? 17 Feb 2025 · 0 repositories · arXiv:2502.11400
-
The geometry of BERT 17 Feb 2025 · 0 repositories · arXiv:2502.12033
-
Bridging the Gap: Enabling Natural Language Queries for NoSQL Databases through Text-to-NoSQL Translation 16 Feb 2025 · 0 repositories · arXiv:2502.11201
-
Integrating Language Models for Enhanced Network State Monitoring in DRL-Based SFC Provisioning 16 Feb 2025 · 0 repositories · arXiv:2502.11298
-
Investigating Language Preference of Multilingual RAG Systems 16 Feb 2025 · 0 repositories · arXiv:2502.11175
-
Leveraging Conditional Mutual Information to Improve Large Language Model Fine-Tuning For Classification 16 Feb 2025 · 0 repositories · arXiv:2502.11258
-
MultiTEND: A Multilingual Benchmark for Natural Language to NoSQL Query Translation 16 Feb 2025 · 0 repositories · arXiv:2502.11022
-
QuOTE: Question-Oriented Text Embeddings 16 Feb 2025 · 0 repositories · arXiv:2502.10976
-
RoseRAG: Robust Retrieval-augmented Generation with Small-scale LLMs via Margin-aware Preference Optimization 16 Feb 2025 · 0 repositories · arXiv:2502.10993
-
SpeechT-RAG: Reliable Depression Detection in LLMs with Retrieval-Augmented Generation Using Speech Timing Information 16 Feb 2025 · 0 repositories · arXiv:2502.10950
-
Vendi-RAG: Adaptively Trading-Off Diversity And Quality Significantly Improves Retrieval Augmented Generation With LLMs 16 Feb 2025 · 0 repositories · arXiv:2502.11228
-
CiteCheck: Towards Accurate Citation Faithfulness Detection 15 Feb 2025 · 1 repository · arXiv:2502.10881
-
Dataset Protection via Watermarked Canaries in Retrieval-Augmented LLMs 15 Feb 2025 · 0 repositories · arXiv:2502.10673
-
Evolving Hate Speech Online: An Adaptive Framework for Detection and Mitigation 15 Feb 2025 · 0 repositories · arXiv:2502.10921
-
NitiBench: A Comprehensive Studies of LLM Frameworks Capabilities for Thai Legal Question Answering 15 Feb 2025 · 1 repository · arXiv:2502.10868
-
ArchRAG: Attributed Community-based Hierarchical Retrieval-Augmented Generation 14 Feb 2025 · 0 repositories · arXiv:2502.09891
-
EmbBERT-Q: Breaking Memory Barriers in Embedded NLP 14 Feb 2025 · 0 repositories · arXiv:2502.10001
-
Hallucinations and Truth: A Comprehensive Accuracy Evaluation of RAG, LoRA and DoRA 14 Feb 2025 · 0 repositories · arXiv:2502.10497
-
LaRA: Benchmarking Retrieval-Augmented Generation and Long-Context LLMs - No Silver Bullet for LC or RAG Routing 14 Feb 2025 · 0 repositories · arXiv:2502.09977
-
Post-training an LLM for RAG? Train on Self-Generated Demonstrations 14 Feb 2025 · 0 repositories · arXiv:2502.10596
-
Enhancing RAG with Active Learning on Conversation Records: Reject Incapables and Answer Capables 13 Feb 2025 · 0 repositories · arXiv:2502.09073
-
KIMAs: A Configurable Knowledge Integrated Multi-Agent System 13 Feb 2025 · 0 repositories · arXiv:2502.09596
-
Predicting Cognitive Decline: A Multimodal AI Approach to Dementia Screening from Speech 13 Feb 2025 · 0 repositories · arXiv:2502.08862
-
Utilizing Pre-trained and Large Language Models for 10-K Items Segmentation 13 Feb 2025 · 0 repositories · arXiv:2502.08875
-
Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation 12 Feb 2025 · 1 repository · arXiv:2502.08826
-
ParetoRAG: Leveraging Sentence-Context Attention for Robust and Efficient Retrieval-Augmented Generation 12 Feb 2025 · 0 repositories · arXiv:2502.08178
-
Systematic Knowledge Injection into Large Language Models via Diverse Augmentation for Domain-Specific RAG 12 Feb 2025 · 1 repository · arXiv:2502.08356Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples)
-
An Advanced NLP Framework for Automated Medical Diagnosis with DeBERTa and Dynamic Contextual Positional Gating 11 Feb 2025 · 0 repositories · arXiv:2502.07755