Methods › Natural Language Processing › Subword Segmentation › WordPiece › Papers, page 17
WordPiece
Papers archive 2025-07-28
archive papers tagged: 7,063 · with a code link: 2,910 · where Syntology ran a sample: 650 (529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (650 of 7,063 tagged: 529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument)
Page 17 of 71: papers 1,601 to 1,700 of 7,063, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
This Paper Had the Smartest Reviewers -- Flattery Detection Utilising an Audio-Textual Transformer-Based Approach 25 Jun 2024 · 1 repository · arXiv:2406.17667
-
Unlocking Continual Learning Abilities in Language Models 25 Jun 2024 · 1 repository · arXiv:2406.17245Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples)
-
Attention Instruction: Amplifying Attention in the Middle via Prompting 24 Jun 2024 · 1 repository · arXiv:2406.17095
-
On the Role of Long-tail Knowledge in Retrieval Augmented Large Language Models 24 Jun 2024 · 0 repositories · arXiv:2406.16367
-
Panza: Design and Analysis of a Fully-Local Personalized Text Writing Assistant 24 Jun 2024 · 1 repository · arXiv:2407.10994
-
Ragnarök: A Reusable RAG Framework and Baselines for TREC 2024 Retrieval-Augmented Generation Track 24 Jun 2024 · 2 repositories · arXiv:2406.16828Syntology official (archive's flag): 9 ran · 19 ran (of which 0 constructed an object rather than computing a result; 19 with no instrument failure: 0 honoured, 0 violated, 19 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 23 harvested samples)
-
MixTex: Unambiguous Recognition Should Not Rely Solely on Real Data 24 Jun 2024 · 1 repository · arXiv:2406.17148
-
Evaluating Ensemble Methods for News Recommender Systems 23 Jun 2024 · 0 repositories · arXiv:2406.16106
-
A multi-speaker multi-lingual voice cloning system based on vits2 for limmits 2024 challenge 22 Jun 2024 · 0 repositories · arXiv:2406.17801
-
A Tale of Trust and Accuracy: Base vs. Instruct LLMs in RAG Systems 21 Jun 2024 · 1 repository · arXiv:2406.14972
-
GiusBERTo: A Legal Language Model for Personal Data De-identification in Italian Court of Auditors Decisions 21 Jun 2024 · 0 repositories · arXiv:2406.15032
-
LongRAG: Enhancing Retrieval-Augmented Generation with Long-context LLMs 21 Jun 2024 · 0 repositories · arXiv:2406.15319
-
Pistis-RAG: Enhancing Retrieval-Augmented Generation with Human Feedback 21 Jun 2024 · 0 repositories · arXiv:2407.00072
-
Unsupervised Morphological Tree Tokenizer 21 Jun 2024 · 0 repositories · arXiv:2406.15245
-
QPaug: Question and Passage Augmentation for Open-Domain Question Answering of LLMs 20 Jun 2024 · 1 repository · arXiv:2406.14277Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
CodeRAG-Bench: Can Retrieval Augment Code Generation? 20 Jun 2024 · 1 repository · arXiv:2406.14497Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
DIRAS: Efficient LLM Annotation of Document Relevance in Retrieval Augmented Generation 20 Jun 2024 · 1 repository · arXiv:2406.14162
-
Evaluating RAG-Fusion with RAGElo: an Automated Elo-based Framework 20 Jun 2024 · 1 repository · arXiv:2406.14783Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Healing Powers of BERT: How Task-Specific Fine-Tuning Recovers Corrupted Language Models 20 Jun 2024 · 0 repositories · arXiv:2406.14459
-
Relation Extraction with Fine-Tuned Large Language Models in Retrieval Augmented Generation Frameworks 20 Jun 2024 · 0 repositories · arXiv:2406.14745
-
Can Long-Context Language Models Subsume Retrieval, RAG, SQL, and More? 19 Jun 2024 · 1 repository · arXiv:2406.13121Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Fine-Tuning BERTs for Definition Extraction from Mathematical Text 19 Jun 2024 · 0 repositories · arXiv:2406.13827
-
FoRAG: Factuality-optimized Retrieval Augmented Generation for Web-enhanced Long-form Question Answering 19 Jun 2024 · 0 repositories · arXiv:2406.13779
-
InstructRAG: Instructing Retrieval-Augmented Generation via Self-Synthesized Rationales 19 Jun 2024 · 1 repository · arXiv:2406.13629Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Model Internals-based Answer Attribution for Trustworthy Retrieval-Augmented Generation 19 Jun 2024 · 1 repository · arXiv:2406.13663Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Multi-Meta-RAG: Improving RAG for Multi-Hop Queries using Database Filtering with LLM-Extracted Metadata 19 Jun 2024 · 1 repository · arXiv:2406.13213
-
R^2AG: Incorporating Retrieval Information into Retrieval Augmented Generation 19 Jun 2024 · 1 repository · arXiv:2406.13249Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples)
-
WikiContradict: A Benchmark for Evaluating LLMs on Real-World Knowledge Conflicts from Wikipedia 19 Jun 2024 · 0 repositories · arXiv:2406.13805
-
From RAGs to rich parameters: Probing how language models utilize external knowledge over parametric information for factual queries 18 Jun 2024 · 0 repositories · arXiv:2406.12824
-
Intermediate Distillation: Data-Efficient Distillation from Black-Box LLMs for Information Retrieval 18 Jun 2024 · 0 repositories · arXiv:2406.12169
-
PlanRAG: A Plan-then-Retrieval Augmented Generation for Generative Large Language Models as Decision Makers 18 Jun 2024 · 1 repository · arXiv:2406.12430Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
Retrieval-Augmented Generation for Generative Artificial Intelligence in Medicine 18 Jun 2024 · 0 repositories · arXiv:2406.12449
-
RichRAG: Crafting Rich Responses for Multi-faceted Queries in Retrieval-Augmented Generation 18 Jun 2024 · 0 repositories · arXiv:2406.12566
-
Unified Active Retrieval for Retrieval Augmented Generation 18 Jun 2024 · 1 repository · arXiv:2406.12534
-
What Makes Two Language Models Think Alike? 18 Jun 2024 · 0 repositories · arXiv:2406.12620
-
Breaking Boundaries: Investigating the Effects of Model Editing on Cross-linguistic Performance 17 Jun 2024 · 0 repositories · arXiv:2406.11139
-
CrAM: Credibility-Aware Attention Modification in LLMs for Combating Misinformation in RAG 17 Jun 2024 · 1 repository · arXiv:2406.11497
-
SeRTS: Self-Rewarding Tree Search for Biomedical Retrieval-Augmented Generation 17 Jun 2024 · 0 repositories · arXiv:2406.11258
-
Evaluating the Efficacy of Open-Source LLMs in Enterprise-Specific RAG Systems: A Comparative Study of Performance and Scalability 17 Jun 2024 · 1 repository · arXiv:2406.11424
-
Fine-Tuning or Fine-Failing? Debunking Performance Myths in Large Language Models 17 Jun 2024 · 0 repositories · arXiv:2406.11201
-
Iterative Utility Judgment Framework via LLMs Inspired by Relevance in Philosophy 17 Jun 2024 · 0 repositories · arXiv:2406.11290
-
R-Eval: A Unified Toolkit for Evaluating Domain Knowledge of Retrieval Augmented Large Language Models 17 Jun 2024 · 1 repository · arXiv:2406.11681
-
Satyrn: A Platform for Analytics Augmented Generation 17 Jun 2024 · 1 repository · arXiv:2406.12069
-
Refiner: Restructure Retrieval Content Efficiently to Advance Question-Answering Capabilities 17 Jun 2024 · 1 repository · arXiv:2406.11357Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
TRACE the Evidence: Constructing Knowledge-Grounded Reasoning Chains for Retrieval-Augmented Generation 17 Jun 2024 · 2 repositories · arXiv:2406.11460Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
WellDunn: On the Robustness and Explainability of Language Models and Large Language Models in Identifying Wellness Dimensions 17 Jun 2024 · 1 repository · arXiv:2406.12058Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Predicting the Understandability of Computational Notebooks through Code Metrics Analysis 16 Jun 2024 · 1 repository · arXiv:2406.10989
-
ShareLoRA: Parameter Efficient and Robust Large Language Model Fine-tuning via Shared Low-Rank Adaptation 16 Jun 2024 · 1 repository · arXiv:2406.10785Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
A Comprehensive Survey of Foundation Models in Medicine 15 Jun 2024 · 0 repositories · arXiv:2406.10729
-
We Care: Multimodal Depression Detection and Knowledge Infused Mental Health Therapeutic Response Generation 15 Jun 2024 · 0 repositories · arXiv:2406.10561
-
Bag of Lies: Robustness in Continuous Pre-training BERT 14 Jun 2024 · 0 repositories · arXiv:2406.09967
-
HIRO: Hierarchical Information Retrieval Optimization 14 Jun 2024 · 1 repository · arXiv:2406.09979
-
The Devil is in the Neurons: Interpreting and Mitigating Social Biases in Pre-trained Language Models 14 Jun 2024 · 1 repository · arXiv:2406.10130
-
Analyzing Gender Polarity in Short Social Media Texts with BERT: The Role of Emojis and Emoticons 13 Jun 2024 · 0 repositories · arXiv:2406.09573
-
BPE-knockout: Pruning Pre-existing BPE Tokenisers with Backwards-compatible Morphological Semi-supervision 13 Jun 2024 · 1 repository
-
PC-LoRA: Low-Rank Adaptation for Progressive Model Compression with Knowledge Distillation 13 Jun 2024 · 0 repositories · arXiv:2406.09117
-
Ad Auctions for LLMs via Retrieval Augmented Generation 12 Jun 2024 · 0 repositories · arXiv:2406.09459
-
Exploring Fact Memorization and Style Imitation in LLMs Using QLoRA: An Experimental Study and Quality Assessment Methods 12 Jun 2024 · 0 repositories · arXiv:2406.08582
-
Label-aware Hard Negative Sampling Strategies with Momentum Contrastive Learning for Implicit Hate Speech Detection 12 Jun 2024 · 1 repository · arXiv:2406.07886
-
Leveraging Large Language Models for Web Scraping 12 Jun 2024 · 0 repositories · arXiv:2406.08246
-
Multimodal Representation Loss Between Timed Text and Audio for Regularized Speech Separation 12 Jun 2024 · 0 repositories · arXiv:2406.08328
-
Beyond Words: On Large Language Models Actionability in Mission-Critical Risk Analysis 11 Jun 2024 · 0 repositories · arXiv:2406.10273
-
COVID-19 Twitter Sentiment Classification Using Hybrid Deep Learning Model Based on Grid Search Methodology 11 Jun 2024 · 0 repositories · arXiv:2406.10266
-
DR-RAG: Applying Dynamic Document Relevance to Retrieval-Augmented Generation for Question-Answering 11 Jun 2024 · 0 repositories · arXiv:2406.07348
-
Multimodal Belief Prediction 11 Jun 2024 · 1 repository · arXiv:2406.07466
-
Question-Answering (QA) Model for a Personalized Learning Assistant for Arabic Language 11 Jun 2024 · 0 repositories · arXiv:2406.08519
-
Leveraging Large Language Models for Knowledge-free Weak Supervision in Clinical Natural Language Processing 10 Jun 2024 · 0 repositories · arXiv:2406.06723
-
The Impact of Quantization on Retrieval-Augmented Generation: An Analysis of Small LLMs 10 Jun 2024 · 0 repositories · arXiv:2406.10251
-
UMBRELA: UMbrela is the (Open-Source Reproduction of the) Bing RELevance Assessor 10 Jun 2024 · 1 repository · arXiv:2406.06519
-
DomainRAG: A Chinese Benchmark for Evaluating Domain-specific Retrieval-Augmented Generation 9 Jun 2024 · 2 repositories · arXiv:2406.05654Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
Machine Against the RAG: Jamming Retrieval-Augmented Generation with Blocker Documents 9 Jun 2024 · 0 repositories · arXiv:2406.05870
-
RE-RAG: Improving Open-Domain QA Performance and Interpretability with Relevance Estimator in Retrieval-Augmented Generation 9 Jun 2024 · 1 repository · arXiv:2406.05794
-
Advancing Semantic Textual Similarity Modeling: A Regression Framework with Translated ReLU and Smooth K2 Loss 8 Jun 2024 · 2 repositories · arXiv:2406.05326
-
Concept Formation and Alignment in Language Models: Bridging Statistical Patterns in Latent Space to Concept Taxonomy 8 Jun 2024 · 0 repositories · arXiv:2406.05315
-
VP-LLM: Text-Driven 3D Volume Completion with Large Language Models through Patchification 8 Jun 2024 · 0 repositories · arXiv:2406.05543
-
BAMO at SemEval-2024 Task 9: BRAINTEASER: A Novel Task Defying Common Sense 7 Jun 2024 · 1 repository · arXiv:2406.04947
-
Corpus Poisoning via Approximate Greedy Gradient Descent 7 Jun 2024 · 1 repository · arXiv:2406.05087
-
CRAG -- Comprehensive RAG Benchmark 7 Jun 2024 · 2 repositories · arXiv:2406.04744Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Multi-Head RAG: Solving Multi-Aspect Problems with LLMs 7 Jun 2024 · 2 repositories · arXiv:2406.05085
-
VTrans: Accelerating Transformer Compression with Variational Information Bottleneck based Pruning 7 Jun 2024 · 0 repositories · arXiv:2406.05276
-
Empirical Guidelines for Deploying LLMs onto Resource-constrained Edge Devices 6 Jun 2024 · 0 repositories · arXiv:2406.03777
-
PoLYTC: a novel BERT-based classifier to detect political leaning of YouTube videos based on their titles 5 Jun 2024 · 1 repository
-
RICo: Reddit ideological communities 5 Jun 2024 · 1 repository
-
Chain of Agents: Large Language Models Collaborating on Long-Context Tasks 4 Jun 2024 · 0 repositories · arXiv:2406.02818
-
Probing the Category of Verbal Aspect in Transformer Language Models 4 Jun 2024 · 0 repositories · arXiv:2406.02335
-
Randomized Geometric Algebra Methods for Convex Neural Networks 4 Jun 2024 · 1 repository · arXiv:2406.02806Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
SMS Spam Detection and Classification to Combat Abuse in Telephone Networks Using Natural Language Processing 4 Jun 2024 · 0 repositories · arXiv:2406.06578
-
Synergetic Event Understanding: A Collaborative Approach to Cross-Document Event Coreference Resolution with Large Language Models 4 Jun 2024 · 1 repository · arXiv:2406.02148
-
Towards Effective Time-Aware Language Representation: Exploring Enhanced Temporal Understanding in Language Models 4 Jun 2024 · 0 repositories · arXiv:2406.01863
-
Annotation Guidelines-Based Knowledge Augmentation: Towards Enhancing Large Language Models for Educational Text Classification 3 Jun 2024 · 0 repositories · arXiv:2406.00954
-
Ask-EDA: A Design Assistant Empowered by LLM, Hybrid RAG and Abbreviation De-hallucination 3 Jun 2024 · 0 repositories · arXiv:2406.06575
-
BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models 3 Jun 2024 · 0 repositories · arXiv:2406.00083
-
FactGenius: Combining Zero-Shot Prompting and Fuzzy Relation Mining to Improve Fact Verification with Knowledge Graphs 3 Jun 2024 · 1 repository · arXiv:2406.01311
-
Focus on the Core: Efficient Attention via Pruned Token Compression for Document Classification 3 Jun 2024 · 0 repositories · arXiv:2406.01283
-
Luna: An Evaluation Foundation Model to Catch Language Model Hallucinations with High Accuracy and Low Cost 3 Jun 2024 · 0 repositories · arXiv:2406.00975
-
Natural Language Interaction with a Household Electricity Knowledge-based Digital Twin 3 Jun 2024 · 0 repositories · arXiv:2406.06566
-
SoccerRAG: Multimodal Soccer Information Retrieval via Natural Queries 3 Jun 2024 · 1 repository · arXiv:2406.01273
-
A Theory for Token-Level Harmonization in Retrieval-Augmented Generation 3 Jun 2024 · 0 repositories · arXiv:2406.00944
-
Formality Style Transfer in Persian 2 Jun 2024 · 0 repositories · arXiv:2406.00867
-
CASE: Efficient Curricular Data Pre-training for Building Assistive Psychology Expert Models 1 Jun 2024 · 1 repository · arXiv:2406.00314