Methods › Natural Language Processing › Subword Segmentation › WordPiece › Papers, page 13
WordPiece
Papers archive 2025-07-28
archive papers tagged: 7,063 · with a code link: 2,910 · where Syntology ran a sample: 650 (529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (650 of 7,063 tagged: 529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument)
Page 13 of 71: papers 1,201 to 1,300 of 7,063, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation 4 Oct 2024 · 2 repositories · arXiv:2410.03960
-
Towards Linguistically-Aware and Language-Independent Tokenization for Large Language Models (LLMs) 4 Oct 2024 · 0 repositories · arXiv:2410.03568
-
Variational Language Concepts for Interpreting Foundation Language Models 4 Oct 2024 · 1 repository · arXiv:2410.03964
-
Vulnerability Detection via Topological Analysis of Attention Maps 4 Oct 2024 · 1 repository · arXiv:2410.03470
-
Ward: Provable RAG Dataset Inference via LLM Watermarks 4 Oct 2024 · 0 repositories · arXiv:2410.03537Syntology 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
A Comprehensive Survey of Retrieval-Augmented Generation (RAG): Evolution, Current Landscape and Future Directions 3 Oct 2024 · 0 repositories · arXiv:2410.12837
-
Adversarial Decoding: Generating Readable Documents for Adversarial Objectives 3 Oct 2024 · 1 repository · arXiv:2410.02163
-
Domain-Specific Retrieval-Augmented Generation Using Vector Stores, Knowledge Graphs, and Tensor Factorization 3 Oct 2024 · 0 repositories · arXiv:2410.02721
-
HELMET: How to Evaluate Long-Context Language Models Effectively and Thoroughly 3 Oct 2024 · 1 repository · arXiv:2410.02694Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples) · 2 pointer-only (licence)
-
How Much Can RAG Help the Reasoning of LLM? 3 Oct 2024 · 0 repositories · arXiv:2410.02338
-
IndicSentEval: How Effectively do Multilingual Transformer Models encode Linguistic Properties for Indic Languages? 3 Oct 2024 · 0 repositories · arXiv:2410.02611
-
Intrinsic Evaluation of RAG Systems for Deep-Logic Questions 3 Oct 2024 · 0 repositories · arXiv:2410.02932
-
L-CiteEval: Do Long-Context Models Truly Leverage Context for Responding? 3 Oct 2024 · 2 repositories · arXiv:2410.02115
-
Morphological evaluation of subwords vocabulary used by BETO language model 3 Oct 2024 · 0 repositories · arXiv:2410.02283
-
Reward-RAG: Enhancing RAG with Reward Driven Supervision 3 Oct 2024 · 0 repositories · arXiv:2410.03780
-
UncertaintyRAG: Span-Level Uncertainty Enhanced Long-Context Modeling for Retrieval-Augmented Generation 3 Oct 2024 · 0 repositories · arXiv:2410.02719
-
BordIRlines: A Dataset for Evaluating Cross-lingual Retrieval-Augmented Generation 2 Oct 2024 · 1 repository · arXiv:2410.01171Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 16 harvested samples)
-
Emotion-Aware Embedding Fusion in LLMs (Flan-T5, LLAMA 2, DeepSeek-R1, and ChatGPT 4) for Intelligent Response Generation 2 Oct 2024 · 0 repositories · arXiv:2410.01306
-
Enhancing Retrieval in QA Systems with Derived Feature Association 2 Oct 2024 · 1 repository · arXiv:2410.03754
-
Financial Sentiment Analysis on News and Reports Using Large Language Models and FinBERT 2 Oct 2024 · 0 repositories · arXiv:2410.01987
-
Open-RAG: Enhanced Retrieval-Augmented Reasoning with Open-Source Large Language Models 2 Oct 2024 · 1 repository · arXiv:2410.01782
-
End-to-End Speech Recognition with Pre-trained Masked Language Model 1 Oct 2024 · 1 repository · arXiv:2410.00528
-
Optimizing and Evaluating Enterprise Retrieval-Augmented Generation (RAG): A Content Design Perspective 1 Oct 2024 · 1 repository · arXiv:2410.12812
-
Quantifying reliance on external information over parametric knowledge during Retrieval Augmented Generation (RAG) using mechanistic analysis 1 Oct 2024 · 0 repositories · arXiv:2410.00857
-
A Methodology for Explainable Large Language Models with Integrated Gradients and Linguistic Analysis in Text Classification 30 Sep 2024 · 0 repositories · arXiv:2410.00250
-
BSharedRAG: Backbone Shared Retrieval-Augmented Generation for the E-commerce Domain 30 Sep 2024 · 0 repositories · arXiv:2409.20075
-
Depression detection in social media posts using transformer-based models and auxiliary features 30 Sep 2024 · 0 repositories · arXiv:2409.20048
-
Evaluating the fairness of task-adaptive pretraining on unlabeled test data before few-shot text classification 30 Sep 2024 · 1 repository · arXiv:2410.00179
-
Ingest-And-Ground: Dispelling Hallucinations from Continually-Pretrained LLMs with RAG 30 Sep 2024 · 0 repositories · arXiv:2410.02825
-
QAEncoder: Towards Aligned Representation Learning in Question Answering System 30 Sep 2024 · 1 repository · arXiv:2409.20434
-
Does RAG Introduce Unfairness in LLMs? Evaluating Fairness in Retrieval-Augmented Generation Systems 29 Sep 2024 · 1 repository · arXiv:2409.19804Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
PEAR: Position-Embedding-Agnostic Attention Re-weighting Enhances Retrieval-Augmented Generation with Zero Inference Overhead 29 Sep 2024 · 0 repositories · arXiv:2409.19745
-
Efficient Federated Intrusion Detection in 5G ecosystem using optimized BERT-based model 28 Sep 2024 · 1 repository · arXiv:2409.19390
-
INSIGHTBUDDY-AI: Medication Extraction and Entity Linking using Large Language Models and Ensemble Learning 28 Sep 2024 · 2 repositories · arXiv:2409.19467
-
AIPatient: Simulating Patients with EHRs and LLM Powered Agentic Workflow 27 Sep 2024 · 0 repositories · arXiv:2409.18924
-
Cottention: Linear Transformers With Cosine Attention 27 Sep 2024 · 1 repository · arXiv:2409.18747Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Experimental Evaluation of Machine Learning Models for Goal-oriented Customer Service Chatbot with Pipeline Architecture 27 Sep 2024 · 0 repositories · arXiv:2409.18568
-
Meta-RTL: Reinforcement-Based Meta-Transfer Learning for Low-Resource Commonsense Reasoning 27 Sep 2024 · 0 repositories · arXiv:2409.19075
-
Multi-Source Hard and Soft Information Fusion Approach for Accurate Cryptocurrency Price Movement Prediction 27 Sep 2024 · 0 repositories · arXiv:2409.18895
-
Suicide Phenotyping from Clinical Notes in Safety-Net Psychiatric Hospital Using Multi-Label Classification with Pre-Trained Language Models 27 Sep 2024 · 0 repositories · arXiv:2409.18878
-
Comparing Unidirectional, Bidirectional, and Word2vec Models for Discovering Vulnerabilities in Compiled Lifted Code 26 Sep 2024 · 0 repositories · arXiv:2409.17513
-
Efficient In-Domain Question Answering for Resource-Constrained Environments 26 Sep 2024 · 0 repositories · arXiv:2409.17648
-
Embodied-RAG: General Non-parametric Embodied Memory for Retrieval and Generation 26 Sep 2024 · 0 repositories · arXiv:2409.18313
-
Enhancing Tourism Recommender Systems for Sustainable City Trips Using Retrieval-Augmented Generation 26 Sep 2024 · 0 repositories · arXiv:2409.18003
-
MultiClimate: Multimodal Stance Detection on Climate Change Videos 26 Sep 2024 · 1 repository · arXiv:2409.18346Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Predicting Anchored Text from Translation Memories for Machine Translation Using Deep Learning Methods 26 Sep 2024 · 0 repositories · arXiv:2409.17939
-
A Prompting-Based Representation Learning Method for Recommendation with Large Language Models 25 Sep 2024 · 0 repositories · arXiv:2409.16674
-
Deep Learning and Machine Learning, Advancing Big Data Analytics and Management: Handy Appetizer 25 Sep 2024 · 0 repositories · arXiv:2409.17120
-
LLaMa-SciQ: An Educational Chatbot for Answering Science MCQ 25 Sep 2024 · 0 repositories · arXiv:2409.16779
-
Non-stationary BERT: Exploring Augmented IMU Data For Robust Human Activity Recognition 25 Sep 2024 · 0 repositories · arXiv:2409.16730
-
Controlling Risk of Retrieval-augmented Generation: A Counterfactual Prompting Framework 24 Sep 2024 · 1 repository · arXiv:2409.16146Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
From Pixels to Words: Leveraging Explainability in Face Recognition through Interactive Natural Language Processing 24 Sep 2024 · 0 repositories · arXiv:2409.16089
-
IRSC: A Zero-shot Evaluation Benchmark for Information Retrieval through Semantic Comprehension in Retrieval-Augmented Generation Scenarios 24 Sep 2024 · 1 repository · arXiv:2409.15763
-
Lighter And Better: Towards Flexible Context Adaptation For Retrieval Augmented Generation 24 Sep 2024 · 0 repositories · arXiv:2409.15699
-
Spelling Correction through Rewriting of Non-Autoregressive ASR Lattices 24 Sep 2024 · 0 repositories · arXiv:2409.16469
-
SwiftDossier: Tailored Automatic Dossier for Drug Discovery with LLMs and Agents 24 Sep 2024 · 0 repositories · arXiv:2409.15817
-
Enhancing Scientific Reproducibility Through Automated BioCompute Object Creation Using Retrieval-Augmented Generation from Publications 23 Sep 2024 · 0 repositories · arXiv:2409.15076
-
GEM-RAG: Graphical Eigen Memories For Retrieval Augmented Generation 23 Sep 2024 · 0 repositories · arXiv:2409.15566
-
Generative AI Is Not Ready for Clinical Use in Patient Education for Lower Back Pain Patients, Even With Retrieval-Augmented Generation 23 Sep 2024 · 0 repositories · arXiv:2409.15260
-
Improving Academic Skills Assessment with NLP and Ensemble Learning 23 Sep 2024 · 0 repositories · arXiv:2409.19013
-
Learning When to Retrieve, What to Rewrite, and How to Respond in Conversational QA 23 Sep 2024 · 0 repositories · arXiv:2409.15515
-
Lessons Learned on Information Retrieval in Electronic Health Records: A Comparison of Embedding Models and Pooling Strategies 23 Sep 2024 · 0 repositories · arXiv:2409.15163
-
Retrieval Augmented Generation (RAG) and Beyond: A Comprehensive Survey on How to Make your LLMs use External Data More Wisely 23 Sep 2024 · 0 repositories · arXiv:2409.14924
-
J2N -- Nominal Adjective Identification and its Application 22 Sep 2024 · 1 repository · arXiv:2409.14374
-
SMART-RAG: Selection using Determinantal Matrices for Augmented Retrieval 21 Sep 2024 · 0 repositories · arXiv:2409.13992
-
Probing Context Localization of Polysemous Words in Pre-trained Language Model Sub-Layers 21 Sep 2024 · 0 repositories · arXiv:2409.14097
-
Towards Building Efficient Sentence BERT Models using Layer Pruning 21 Sep 2024 · 0 repositories · arXiv:2409.14168
-
QMOS: Enhancing LLMs for Telecommunication with Question Masked loss and Option Shuffling 21 Sep 2024 · 1 repository · arXiv:2409.14175
-
Drift to Remember 21 Sep 2024 · 0 repositories · arXiv:2409.13997
-
Contextual Compression in Retrieval-Augmented Generation for Large Language Models: A Survey 20 Sep 2024 · 1 repository · arXiv:2409.13385
-
Applying Pre-trained Multilingual BERT in Embeddings for Improved Malicious Prompt Injection Attacks Detection 20 Sep 2024 · 0 repositories · arXiv:2409.13331
-
Enhancing Large Language Models with Domain-specific Retrieval Augment Generation: A Case Study on Long-form Consumer Health Question Answering in Ophthalmology 20 Sep 2024 · 0 repositories · arXiv:2409.13902
-
HUT: A More Computation Efficient Fine-Tuning Method With Hadamard Updated Transformation 20 Sep 2024 · 0 repositories · arXiv:2409.13501
-
ShizishanGPT: An Agricultural Large Language Model Integrating Tools and Resources 20 Sep 2024 · 1 repository · arXiv:2409.13537
-
Enhancing E-commerce Product Title Translation with Retrieval-Augmented Generation and Large Language Models 19 Sep 2024 · 0 repositories · arXiv:2409.12880
-
Enhancing TinyBERT for Financial Sentiment Analysis Using GPT-Augmented FinBERT Distillation 19 Sep 2024 · 1 repository · arXiv:2409.18999
-
Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation 19 Sep 2024 · 2 repositories · arXiv:2409.12941
-
Incremental and Data-Efficient Concept Formation to Support Masked Word Prediction 19 Sep 2024 · 0 repositories · arXiv:2409.12440
-
Profiling Patient Transcript Using Large Language Model Reasoning Augmentation for Alzheimer's Disease Detection 19 Sep 2024 · 1 repository · arXiv:2409.12541
-
Retrieval-Augmented Test Generation: How Far Are We? 19 Sep 2024 · 0 repositories · arXiv:2409.12682
-
Should RAG Chatbots Forget Unimportant Conversations? Exploring Importance and Forgetting with Psychological Insights 19 Sep 2024 · 1 repository · arXiv:2409.12524
-
Text2Traj2Text: Learning-by-Synthesis Framework for Contextual Captioning of Human Movement Trajectories 19 Sep 2024 · 1 repository · arXiv:2409.12670
-
BERT-VBD: Vietnamese Multi-Document Summarization Framework 18 Sep 2024 · 0 repositories · arXiv:2409.12134
-
Multi-Grid Graph Neural Networks with Self-Attention for Computational Mechanics 18 Sep 2024 · 1 repository · arXiv:2409.11899Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
VERA: Validation and Enhancement for Retrieval Augmented systems 18 Sep 2024 · 0 repositories · arXiv:2409.15364
-
Cross-lingual transfer of multilingual models on low resource African Languages 17 Sep 2024 · 2 repositories · arXiv:2409.10965
-
HEARTS: A Holistic Framework for Explainable, Sustainable and Robust Text Stereotype Detection 17 Sep 2024 · 1 repository · arXiv:2409.11579
-
Learning variant product relationship and variation attributes from e-commerce website structures 17 Sep 2024 · 0 repositories · arXiv:2410.02779
-
Measuring and Enhancing Trustworthiness of LLMs in RAG through Grounded Attributions and Learning to Refuse 17 Sep 2024 · 1 repository · arXiv:2409.11242Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
P-RAG: Progressive Retrieval Augmented Generation For Planning on Embodied Everyday Task 17 Sep 2024 · 0 repositories · arXiv:2409.11279
-
THaMES: An End-to-End Tool for Hallucination Mitigation and Evaluation in Large Language Models 17 Sep 2024 · 1 repository · arXiv:2409.11353Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Towards Fair RAG: On the Impact of Fair Ranking in Retrieval-Augmented Generation 17 Sep 2024 · 1 repository · arXiv:2409.11598Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Lab-AI: Using Retrieval Augmentation to Enhance Language Models for Personalized Lab Test Interpretation in Clinical Medicine 16 Sep 2024 · 0 repositories · arXiv:2409.18986
-
SFR-RAG: Towards Contextually Faithful LLMs 16 Sep 2024 · 0 repositories · arXiv:2409.09916
-
Trustworthiness in Retrieval-Augmented Generation Systems: A Survey 16 Sep 2024 · 1 repository · arXiv:2409.10102Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Integrating AI's Carbon Footprint into Risk Management Frameworks: Strategies and Tools for Sustainable Compliance in Banking Sector 15 Sep 2024 · 0 repositories · arXiv:2410.01818
-
Language Models and Retrieval Augmented Generation for Automated Structured Data Extraction from Diagnostic Reports 15 Sep 2024 · 0 repositories · arXiv:2409.10576
-
Towards understanding evolution of science through language model series 15 Sep 2024 · 1 repository · arXiv:2409.09636
-
Active Learning to Guide Labeling Efforts for Question Difficulty Estimation 14 Sep 2024 · 1 repository · arXiv:2409.09258
-
Block-Attention for Efficient RAG 14 Sep 2024 · 1 repository · arXiv:2409.15355Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 6 unverified (of 7 harvested samples) · 7 pointer-only (licence)