Methods › Natural Language Processing › Subword Segmentation › WordPiece › Papers, page 11
WordPiece
Papers archive 2025-07-28
archive papers tagged: 7,063 · with a code link: 2,910 · where Syntology ran a sample: 650 (529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (650 of 7,063 tagged: 529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument)
Page 11 of 71: papers 1,001 to 1,100 of 7,063, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Best Practices for Distilling Large Language Models into BERT for Web Search Ranking 7 Nov 2024 · 0 repositories · arXiv:2411.04539
-
Deploying Large Language Models With Retrieval Augmented Generation 7 Nov 2024 · 1 repository · arXiv:2411.11895
-
Enhancing classroom teaching with LLMs and RAG 7 Nov 2024 · 0 repositories · arXiv:2411.04341
-
M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding 7 Nov 2024 · 0 repositories · arXiv:2411.04952
-
Selecting Between BERT and GPT for Text Classification in Political Science Research 7 Nov 2024 · 0 repositories · arXiv:2411.05050
-
Words that Move Markets- Quantifying the Impact of RBI's Monetary Policy Communications on Indian Financial Market 7 Nov 2024 · 0 repositories · arXiv:2411.04808
-
A Multilingual Sentiment Lexicon for Low-Resource Language Translation using Large Languages Models and Explainable AI 6 Nov 2024 · 0 repositories · arXiv:2411.04316
-
Advanced RAG Models with Graph Structures: Optimizing Complex Knowledge Reasoning and Text Generation 6 Nov 2024 · 0 repositories · arXiv:2411.03572
-
Fine-Grained Guidance for Retrievers: Leveraging LLMs' Feedback in Retrieval-Augmented Generation 6 Nov 2024 · 0 repositories · arXiv:2411.03957
-
From Word Vectors to Multimodal Embeddings: Techniques, Applications, and Future Directions For Large Language Models 6 Nov 2024 · 0 repositories · arXiv:2411.05036
-
RAGulator: Lightweight Out-of-Context Detectors for Grounded Text Generation 6 Nov 2024 · 0 repositories · arXiv:2411.03920
-
Understanding the Effects of Human-written Paraphrases in LLM-generated Text Detection 6 Nov 2024 · 1 repository · arXiv:2411.03806
-
HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieved Knowledge in RAG Systems 5 Nov 2024 · 1 repository · arXiv:2411.02959
-
LASER: Attention with Exponential Transformation 5 Nov 2024 · 0 repositories · arXiv:2411.03493
-
Long Context RAG Performance of Large Language Models 5 Nov 2024 · 0 repositories · arXiv:2411.03538
-
PersianRAG: A Retrieval-Augmented Generation System for Persian Language 5 Nov 2024 · 0 repositories · arXiv:2411.02832
-
A Comparative Analysis of Counterfactual Explanation Methods for Text Classifiers 4 Nov 2024 · 0 repositories · arXiv:2411.02643
-
Can Language Models Enable In-Context Database? 4 Nov 2024 · 0 repositories · arXiv:2411.01807
-
RAGViz: Diagnose and Visualize Retrieval-Augmented Generation 4 Nov 2024 · 1 repository · arXiv:2411.01751
-
TeleOracle: Fine-Tuned Retrieval-Augmented Generation with Long-Context Support for Network 4 Nov 2024 · 1 repository · arXiv:2411.02617
-
Wave Network: An Ultra-Small Language Model 4 Nov 2024 · 0 repositories · arXiv:2411.02674
-
Data Extraction Attacks in Retrieval-Augmented Generation via Backdoors 3 Nov 2024 · 0 repositories · arXiv:2411.01705
-
Enriching Tabular Data with Contextual LLM Embeddings: A Comprehensive Ablation Study for Ensemble Classifiers 3 Nov 2024 · 0 repositories · arXiv:2411.01645
-
AttackQA: Development and Adoption of a Dataset for Assisting Cybersecurity Operations using Fine-tuned and Open-Source LLMs 1 Nov 2024 · 0 repositories · arXiv:2411.01073
-
CORAG: A Cost-Constrained Retrieval Optimization System for Retrieval-Augmented Generation 1 Nov 2024 · 0 repositories · arXiv:2411.00744
-
LLM-Ref: Enhancing Reference Handling in Technical Writing with Large Language Models 1 Nov 2024 · 0 repositories · arXiv:2411.00294
-
Provenance: A Light-weight Fact-checker for Retrieval Augmented LLM Generation Output 1 Nov 2024 · 0 repositories · arXiv:2411.01022
-
Rationale-Guided Retrieval Augmented Generation for Medical Question Answering 1 Nov 2024 · 1 repository · arXiv:2411.00300Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Towards Multi-Source Retrieval-Augmented Generation via Synergizing Reasoning and Preference-Driven Retrieval 1 Nov 2024 · 0 repositories · arXiv:2411.00689
-
JudgeRank: Leveraging Large Language Models for Reasoning-Intensive Reranking 31 Oct 2024 · 0 repositories · arXiv:2411.00142
-
LEAF: Learning and Evaluation Augmented by Fact-Checking to Improve Factualness in Large Language Models 31 Oct 2024 · 0 repositories · arXiv:2410.23526
-
Responsible Retrieval Augmented Generation for Climate Decision Making from Documents 31 Oct 2024 · 0 repositories · arXiv:2410.23902
-
CORAL: Benchmarking Multi-turn Conversational Retrieval-Augmentation Generation 30 Oct 2024 · 1 repository · arXiv:2410.23090
-
Eliciting Critical Reasoning in Retrieval-Augmented Language Models via Contrastive Explanations 30 Oct 2024 · 0 repositories · arXiv:2410.22874
-
Emotional RAG: Enhancing Role-Playing Agents through Emotional Retrieval 30 Oct 2024 · 1 repository · arXiv:2410.23041
-
HijackRAG: Hijacking Attacks against Retrieval-Augmented Large Language Models 30 Oct 2024 · 0 repositories · arXiv:2410.22832
-
ProTransformer: Robustify Transformers via Plug-and-Play Paradigm 30 Oct 2024 · 1 repository · arXiv:2410.23182
-
Retrieval-Augmented Generation with Estimation of Source Reliability 30 Oct 2024 · 0 repositories · arXiv:2410.22954
-
Long²RAG: Evaluating Long-Context & Long-Form Retrieval-Augmented Generation with Key Point Recall 30 Oct 2024 · 0 repositories · arXiv:2410.23000
-
Abrupt Learning in Transformers: A Case Study on Matrix Completion 29 Oct 2024 · 0 repositories · arXiv:2410.22244Syntology 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Beyond Text: Optimizing RAG with Multimodal Inputs for Industrial Applications 29 Oct 2024 · 1 repository · arXiv:2410.21943Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Meta-Learning Adaptable Foundation Models 29 Oct 2024 · 0 repositories · arXiv:2410.22264
-
AutoRAG: Automated Framework for optimization of Retrieval Augmented Generation Pipeline 28 Oct 2024 · 2 repositories · arXiv:2410.20878
-
BanditCAT and AutoIRT: Machine Learning Approaches to Computerized Adaptive Testing and Item Calibration 28 Oct 2024 · 0 repositories · arXiv:2410.21033
-
Calibrated Decision-Making through LLM-Assisted Retrieval 28 Oct 2024 · 0 repositories · arXiv:2411.08891
-
Combining Domain-Specific Models and LLMs for Automated Disease Phenotyping from Survey Data 28 Oct 2024 · 0 repositories · arXiv:2410.20695
-
CRAT: A Multi-Agent Framework for Causality-Enhanced Reflective and Retrieval-Augmented Translation with Large Language Models 28 Oct 2024 · 0 repositories · arXiv:2410.21067
-
Deep Learning for Medical Text Processing: BERT Model Fine-Tuning and Comparative Study 28 Oct 2024 · 0 repositories · arXiv:2410.20792
-
Embedding with Large Language Models for Classification of HIPAA Safeguard Compliance Rules 28 Oct 2024 · 0 repositories · arXiv:2410.20664
-
Geo-FuB: A Method for Constructing an Operator-Function Knowledge Base for Geospatial Code Generation Tasks Using Large Language Models 28 Oct 2024 · 1 repository · arXiv:2410.20975
-
KD-LoRA: A Hybrid Approach to Efficient Fine-Tuning with LoRA and Knowledge Distillation 28 Oct 2024 · 1 repository · arXiv:2410.20777
-
LinFormer: A Linear-based Lightweight Transformer Architecture For Time-Aware MIMO Channel Prediction 28 Oct 2024 · 0 repositories · arXiv:2410.21351
-
LLMs are Biased Evaluators But Not Biased for Retrieval Augmented Generation 28 Oct 2024 · 1 repository · arXiv:2410.20833
-
MultiTok: Variable-Length Tokenization for Efficient LLMs Adapted from LZW Compression 28 Oct 2024 · 1 repository · arXiv:2410.21548
-
Plan×RAG: Planning-guided Retrieval Augmented Generation 28 Oct 2024 · 0 repositories · arXiv:2410.20753
-
Simple Is Effective: The Roles of Graphs and Large Language Models in Knowledge-Graph-Based Retrieval-Augmented Generation 28 Oct 2024 · 1 repository · arXiv:2410.20724Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified (of 15 harvested samples)
-
uOttawa at LegalLens-2024: Transformer-based Classification Experiments 28 Oct 2024 · 1 repository · arXiv:2410.21139
-
Deep Learning Based Dense Retrieval: A Comparative Study 27 Oct 2024 · 0 repositories · arXiv:2410.20315
-
LLM Robustness Against Misinformation in Biomedical Question Answering 27 Oct 2024 · 1 repository · arXiv:2410.21330
-
R^3AG: First Workshop on Refined and Reliable Retrieval Augmented Generation 27 Oct 2024 · 0 repositories · arXiv:2410.20598
-
Mask-based Membership Inference Attacks for Retrieval-Augmented Generation 26 Oct 2024 · 0 repositories · arXiv:2410.20142
-
ChunkRAG: Novel LLM-Chunk Filtering Method for RAG Systems 25 Oct 2024 · 0 repositories · arXiv:2410.19572
-
FISHNET: Financial Intelligence from Sub-querying, Harmonizing, Neural-Conditioning, Expert Swarms, and Task Planning 25 Oct 2024 · 0 repositories · arXiv:2410.19727
-
GeoLLaVA: Efficient Fine-Tuned Vision-Language Models for Temporal Change Detection in Remote Sensing 25 Oct 2024 · 1 repository · arXiv:2410.19552Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Bielik 7B v0.1: A Polish Language Model -- Development, Insights, and Evaluation 24 Oct 2024 · 0 repositories · arXiv:2410.18565
-
Difficult for Whom? A Study of Japanese Lexical Complexity 24 Oct 2024 · 1 repository · arXiv:2410.18567
-
PDL: A Declarative Prompt Programming Language 24 Oct 2024 · 1 repository · arXiv:2410.19135
-
The Nature of Mathematical Modeling and Probabilistic Optimization Engineering in Generative AI 24 Oct 2024 · 0 repositories · arXiv:2410.18441
-
Understanding Players as if They Are Talking to the Game in a Customized Language: A Pilot Study 24 Oct 2024 · 0 repositories · arXiv:2410.18605
-
An Adaptive Framework for Generating Systematic Explanatory Answer in Online Q&A Platforms 23 Oct 2024 · 1 repository · arXiv:2410.17694
-
Leveraging the Domain Adaptation of Retrieval Augmented Generation Models for Question Answering and Reducing Hallucination 23 Oct 2024 · 0 repositories · arXiv:2410.17783
-
Small Singular Values Matter: A Random Matrix Analysis of Transformer Models 23 Oct 2024 · 0 repositories · arXiv:2410.17770
-
LongRAG: A Dual-Perspective Retrieval-Augmented Generation Paradigm for Long-Context Question Answering 23 Oct 2024 · 1 repository · arXiv:2410.18050Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
MCUBERT: Memory-Efficient BERT Inference on Commodity Microcontrollers 23 Oct 2024 · 0 repositories · arXiv:2410.17957
-
SimRAG: Self-Improving Retrieval-Augmented Generation for Adapting Large Language Models to Specialized Domains 23 Oct 2024 · 0 repositories · arXiv:2410.17952
-
A Bayesian Perspective on the Maximum Score Problem 22 Oct 2024 · 0 repositories · arXiv:2410.17153
-
Dhoroni: Exploring Bengali Climate Change and Environmental Views with a Multi-Perspective News Dataset and Natural Language Processing 22 Oct 2024 · 1 repository · arXiv:2410.17225
-
Distill-SynthKG: Distilling Knowledge Graph Synthesis Workflow for Improved Coverage and Efficiency 22 Oct 2024 · 0 repositories · arXiv:2410.16597
-
DNAHLM -- DNA sequence and Human Language mixed large language Model 22 Oct 2024 · 1 repository · arXiv:2410.16917
-
SmartRAG: Jointly Learn RAG-Related Tasks From the Environment Feedback 22 Oct 2024 · 0 repositories · arXiv:2410.18141
-
Tracing the Development of the Virtual Particle Concept Using Semantic Change Detection 22 Oct 2024 · 1 repository · arXiv:2410.16855
-
Deep Learning and Data Augmentation for Detecting Self-Admitted Technical Debt 21 Oct 2024 · 1 repository · arXiv:2410.15804
-
Developing Retrieval Augmented Generation (RAG) based LLM Systems from PDFs: An Experience Report 21 Oct 2024 · 1 repository · arXiv:2410.15944
-
Exploring Pretraining via Active Forgetting for Improving Cross Lingual Transfer for Decoder Language Models 21 Oct 2024 · 0 repositories · arXiv:2410.16168
-
Leveraging Retrieval-Augmented Generation for Culturally Inclusive Hakka Chatbots: Design Insights and User Perceptions 21 Oct 2024 · 0 repositories · arXiv:2410.15572
-
LightFusionRec: Lightweight Transformers-Based Cross-Domain Recommendation Model 21 Oct 2024 · 0 repositories · arXiv:2410.15656
-
Natural GaLore: Accelerating GaLore for memory-efficient LLM Training and Fine-tuning 21 Oct 2024 · 1 repository · arXiv:2410.16029
-
RAG4ITOps: A Supervised Fine-Tunable and Comprehensive RAG Framework for IT Operations and Maintenance 21 Oct 2024 · 0 repositories · arXiv:2410.15805
-
SeisLM: a Foundation Model for Seismic Waveforms 21 Oct 2024 · 1 repository · arXiv:2410.15765
-
Who's Who: Large Language Models Meet Knowledge Conflicts in Practice 21 Oct 2024 · 1 repository · arXiv:2410.15737
-
Contextual Augmented Multi-Model Programming (CAMP): A Hybrid Local-Cloud Copilot Framework 20 Oct 2024 · 1 repository · arXiv:2410.15285
-
ConTReGen: Context-driven Tree-structured Retrieval for Open-domain Long-form Text Generation 20 Oct 2024 · 0 repositories · arXiv:2410.15511
-
Do RAG Systems Cover What Matters? Evaluating and Optimizing Responses with Sub-Question Coverage 20 Oct 2024 · 1 repository · arXiv:2410.15531
-
Evaluating Consistencies in LLM responses through a Semantic Clustering of Question Answering 20 Oct 2024 · 0 repositories · arXiv:2410.15440
-
MMDS: A Multimodal Medical Diagnosis System Integrating Image Analysis and Knowledge-based Departmental Consultation 20 Oct 2024 · 0 repositories · arXiv:2410.15403
-
Unveiling and Consulting Core Experts in Retrieval-Augmented MoE-based LLMs 20 Oct 2024 · 0 repositories · arXiv:2410.15438
-
When Machine Unlearning Meets Retrieval-Augmented Generation (RAG): Keep Secret or Forget Knowledge? 20 Oct 2024 · 0 repositories · arXiv:2410.15267
-
Evaluation Of P300 Speller Performance Using Large Language Models Along With Cross-Subject Training 19 Oct 2024 · 1 repository · arXiv:2410.15161
-
MCCoder: Streamlining Motion Control with LLM-Assisted Code Generation and Rigorous Verification 19 Oct 2024 · 1 repository · arXiv:2410.15154
-
Medical-GAT: Cancer Document Classification Leveraging Graph-Based Residual Network for Scenarios with Limited Data 19 Oct 2024 · 0 repositories · arXiv:2410.15198