Methods › Natural Language Processing › Subword Segmentation › WordPiece › Papers, page 16
WordPiece
Papers archive 2025-07-28
archive papers tagged: 7,063 · with a code link: 2,910 · where Syntology ran a sample: 650 (529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (650 of 7,063 tagged: 529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument)
Page 16 of 71: papers 1,501 to 1,600 of 7,063, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Retrieval-Augmented Generation for Natural Language Processing: A Survey 18 Jul 2024 · 0 repositories · arXiv:2407.13193
-
Retrieve, Summarize, Plan: Advancing Multi-hop Question Answering with an Iterative Approach 18 Jul 2024 · 0 repositories · arXiv:2407.13101
-
AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases 17 Jul 2024 · 1 repository · arXiv:2407.12784Syntology official (archive's flag): 16 ran · 16 ran (of which 0 constructed an object rather than computing a result; 15 with no instrument failure: 1 honoured, 0 violated, 14 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 19 harvested samples) · 1 pointer-only (licence)
-
Deep Learning-based Sentiment Analysis of Olympics Tweets 17 Jul 2024 · 0 repositories · arXiv:2407.12376
-
On Initializing Transformers with Pre-trained Embeddings 17 Jul 2024 · 0 repositories · arXiv:2407.12514
-
Optimizing Query Generation for Enhanced Document Retrieval in RAG 17 Jul 2024 · 0 repositories · arXiv:2407.12325
-
Evaluating Search Engines and Large Language Models for Answering Health Questions 17 Jul 2024 · 1 repository · arXiv:2407.12468
-
Sharif-STR at SemEval-2024 Task 1: Transformer as a Regression Model for Fine-Grained Scoring of Textual Semantic Relations 17 Jul 2024 · 1 repository · arXiv:2407.12426
-
Textualized and Feature-based Models for Compound Multimodal Emotion Recognition in the Wild 17 Jul 2024 · 2 repositories · arXiv:2407.12927Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Better RAG using Relevant Information Gain 16 Jul 2024 · 1 repository · arXiv:2407.12101
-
Large Visual-Language Models Are Also Good Classifiers: A Study of In-Context Multimodal Fake News Detection 16 Jul 2024 · 0 repositories · arXiv:2407.12879
-
LLMs-in-the-loop Part-1: Expert Small AI Models for Bio-Medical Text Translation 16 Jul 2024 · 0 repositories · arXiv:2407.12126
-
Mindful-RAG: A Study of Points of Failure in Retrieval Augmented Generation 16 Jul 2024 · 0 repositories · arXiv:2407.12216
-
R-SFLLM: Jamming Resilient Framework for Split Federated Learning with Large Language Models 16 Jul 2024 · 0 repositories · arXiv:2407.11654
-
Scientific QA System with Verifiable Answers 16 Jul 2024 · 1 repository · arXiv:2407.11485
-
Communication- and Computation-Efficient Distributed Submodular Optimization in Robot Mesh Networks 15 Jul 2024 · 1 repository · arXiv:2407.10382
-
Deep Learning-Based Operators for Evolutionary Algorithms 15 Jul 2024 · 0 repositories · arXiv:2407.10477
-
Enhancing Retrieval and Managing Retrieval: A Four-Module Synergy for Improved Quality and Efficiency in RAG Systems 15 Jul 2024 · 1 repository · arXiv:2407.10670Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 1 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
Evaluation of RAG Metrics for Question Answering in the Telecom Domain 15 Jul 2024 · 0 repositories · arXiv:2407.12873
-
MixGR: Enhancing Retriever Generalization for Scientific Domain through Complementary Granularity 15 Jul 2024 · 1 repository · arXiv:2407.10691
-
Think-on-Graph 2.0: Deep and Faithful Large Language Model Reasoning with Knowledge-guided Retrieval Augmented Generation 15 Jul 2024 · 1 repository · arXiv:2407.10805Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 1 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 16 harvested samples)
-
Causality extraction from medical text using Large Language Models (LLMs) 13 Jul 2024 · 0 repositories · arXiv:2407.10020
-
Document-level Clinical Entity and Relation Extraction via Knowledge Base-Guided Generation 13 Jul 2024 · 0 repositories · arXiv:2407.10021
-
Hydra: Bidirectional State Space Models Through Generalized Matrix Mixers 13 Jul 2024 · 1 repository · arXiv:2407.09941Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 10 harvested samples) · 1 pointer-only (licence)
-
Resource Management for Low-latency Cooperative Fine-tuning of Foundation Models at the Network Edge 13 Jul 2024 · 0 repositories · arXiv:2407.09873
-
Deep Bag-of-Words Model: An Efficient and Interpretable Relevance Architecture for Chinese E-Commerce 12 Jul 2024 · 0 repositories · arXiv:2407.09395
-
Enhancing Depressive Post Detection in Bangla: A Comparative Study of TF-IDF, BERT and FastText Embeddings 12 Jul 2024 · 0 repositories · arXiv:2407.09187
-
Human-like Episodic Memory for Infinite Context LLMs 12 Jul 2024 · 1 repository · arXiv:2407.09450Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 13 harvested samples)
-
Movie Recommendation with Poster Attention via Multi-modal Transformer Feature Fusion 12 Jul 2024 · 0 repositories · arXiv:2407.09157
-
Robustness of LLMs to Perturbations in Text 12 Jul 2024 · 0 repositories · arXiv:2407.08989
-
Beyond Benchmarks: Evaluating Embedding Model Similarity for Retrieval Augmented Generation Systems 11 Jul 2024 · 1 repository · arXiv:2407.08275
-
fairBERTs: Erasing Sensitive Information Through Semantic and Fairness-aware Perturbations 11 Jul 2024 · 0 repositories · arXiv:2407.08189
-
Investigating LLMs as Voting Assistants via Contextual Augmentation: A Case Study on the European Parliament Elections 2024 11 Jul 2024 · 0 repositories · arXiv:2407.08495
-
Speculative RAG: Enhancing Retrieval Augmented Generation through Drafting 11 Jul 2024 · 0 repositories · arXiv:2407.08223
-
Attribute or Abstain: Large Language Models as Long Document Assistants 10 Jul 2024 · 1 repository · arXiv:2407.07799Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
DS@GT eRisk 2024: Sentence Transformers for Social Media Risk Assessment 10 Jul 2024 · 1 repository · arXiv:2407.08008
-
FACTS About Building Retrieval Augmented Generation-based Chatbots 10 Jul 2024 · 0 repositories · arXiv:2407.07858
-
FsPONER: Few-shot Prompt Optimization for Named Entity Recognition in Domain-specific Scenarios 10 Jul 2024 · 1 repository · arXiv:2407.08035
-
Examining Long-Context Large Language Models for Environmental Review Document Comprehension 10 Jul 2024 · 0 repositories · arXiv:2407.07321
-
ROSA: Random Subspace Adaptation for Efficient Fine-Tuning 10 Jul 2024 · 1 repository · arXiv:2407.07802
-
A Simple Architecture for Enterprise Large Language Model Applications based on Role based security and Clearance Levels using Retrieval-Augmented Generation or Mixture of Experts 9 Jul 2024 · 0 repositories · arXiv:2407.06718
-
Empirical analysis of Binding Precedent efficiency in the Brazilian Supreme Court via Similar Case Retrieval 9 Jul 2024 · 0 repositories · arXiv:2407.07004
-
An Empirical Comparison of Vocabulary Expansion and Initialization Approaches for Language Models 8 Jul 2024 · 1 repository · arXiv:2407.05841Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Personality Analysis for Social Media Users using Arabic language and its Effect on Sentiment Analysis 8 Jul 2024 · 0 repositories · arXiv:2407.06314
-
How do you know that? Teaching Generative Language Models to Reference Answers to Biomedical Questions 6 Jul 2024 · 1 repository · arXiv:2407.05015
-
RULE: Reliable Multimodal RAG for Factuality in Medical Vision Language Models 6 Jul 2024 · 1 repository · arXiv:2407.05131Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Are LLMs Correctly Integrated into Software Systems? 6 Jul 2024 · 0 repositories · arXiv:2407.05138
-
GPT vs RETRO: Exploring the Intersection of Retrieval and Parameter-Efficient Fine-Tuning 5 Jul 2024 · 0 repositories · arXiv:2407.04528
-
Using LLMs to label medical papers according to the CIViC evidence model 5 Jul 2024 · 1 repository · arXiv:2407.04466
-
Convolutional vs Large Language Models for Software Log Classification in Edge-Deployable Cellular Network Testing 4 Jul 2024 · 0 repositories · arXiv:2407.03759
-
Deep Content Understanding Toward Entity and Aspect Target Sentiment Analysis on Foundation Models 4 Jul 2024 · 1 repository · arXiv:2407.04050
-
DSLR: Document Refinement with Sentence-Level Re-ranking and Reconstruction to Enhance Retrieval-Augmented Generation 4 Jul 2024 · 0 repositories · arXiv:2407.03627
-
QET: Enhancing Quantized LLM Parameters and KV cache Compression through Element Substitution and Residual Clustering 4 Jul 2024 · 0 repositories · arXiv:2407.03637
-
HYBRINFOX at CheckThat! 2024 -- Task 1: Enhancing Language Models with Structured Information for Check-Worthiness Estimation 4 Jul 2024 · 0 repositories · arXiv:2407.03850
-
HYBRINFOX at CheckThat! 2024 -- Task 2: Enriching BERT Models with the Expert System VAGO for Subjectivity Detection 4 Jul 2024 · 0 repositories · arXiv:2407.03770
-
A Comparative Study of DSL Code Generation: Fine-Tuning vs. Optimized Retrieval Augmentation 3 Jul 2024 · 0 repositories · arXiv:2407.02742
-
CATT: Character-based Arabic Tashkeel Transformer 3 Jul 2024 · 1 repository · arXiv:2407.03236
-
Croppable Knowledge Graph Embedding 3 Jul 2024 · 0 repositories · arXiv:2407.02779
-
MLKD-BERT: Multi-level Knowledge Distillation for Pre-trained Language Models 3 Jul 2024 · 0 repositories · arXiv:2407.02775
-
RDBE: Reasoning Distillation-Based Evaluation Enhances Automatic Essay Scoring 3 Jul 2024 · 0 repositories · arXiv:2407.13781
-
Extracting and Encoding: Leveraging Large Language Models and Medical Knowledge to Enhance Radiological Text Representation 2 Jul 2024 · 1 repository · arXiv:2407.01948
-
MeMemo: On-device Retrieval Augmentation for Private and Personalized Text Generation 2 Jul 2024 · 1 repository · arXiv:2407.01972
-
RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs 2 Jul 2024 · 0 repositories · arXiv:2407.02485
-
The Solution for The PST-KDD-2024 OAG-Challenge 2 Jul 2024 · 0 repositories · arXiv:2407.12827
-
BERGEN: A Benchmarking Library for Retrieval-Augmented Generation 1 Jul 2024 · 1 repository · arXiv:2407.01102Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Face4RAG: Factual Consistency Evaluation for Retrieval Augmented Generation in Chinese 1 Jul 2024 · 0 repositories · arXiv:2407.01080
-
Ground Every Sentence: Improving Retrieval-Augmented LLMs with Interleaved Reference-Claim Generation 1 Jul 2024 · 0 repositories · arXiv:2407.01796
-
Hybrid RAG-empowered Multi-modal LLM for Secure Data Management in Internet of Medical Things: A Diffusion-based Contract Approach 1 Jul 2024 · 0 repositories · arXiv:2407.00978
-
Multi-Modal Fusion-Based Multi-Task Semantic Communication System 1 Jul 2024 · 0 repositories · arXiv:2407.00964
-
Retrieval-augmented generation in multilingual settings 1 Jul 2024 · 1 repository · arXiv:2407.01463Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Searching for Best Practices in Retrieval-Augmented Generation 1 Jul 2024 · 1 repository · arXiv:2407.01219Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Summary of a Haystack: A Challenge to Long-Context LLMs and RAG Systems 1 Jul 2024 · 1 repository · arXiv:2407.01370Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Memory³: Language Modeling with Explicit Memory 1 Jul 2024 · 0 repositories · arXiv:2407.01178
-
Characterizing Stereotypical Bias from Privacy-preserving Pre-Training 30 Jun 2024 · 0 repositories · arXiv:2407.00764
-
LegalTurk Optimized BERT for Multi-Label Text Classification and NER 30 Jun 2024 · 0 repositories · arXiv:2407.00648
-
Parm: Efficient Training of Large Sparsely-Activated Models with Dedicated Schedules 30 Jun 2024 · 1 repository · arXiv:2407.00599
-
Answering real-world clinical questions using large language model based systems 29 Jun 2024 · 0 repositories · arXiv:2407.00541
-
From RAG to RICHES: Retrieval Interlaced with Sequence Generation 29 Jun 2024 · 0 repositories · arXiv:2407.00361
-
LLM-Generated Natural Language Meets Scaling Laws: New Explorations and Data Augmentation Methods 29 Jun 2024 · 0 repositories · arXiv:2407.00322
-
BioMNER: A Dataset for Biomedical Method Entity Recognition 28 Jun 2024 · 0 repositories · arXiv:2406.20038
-
Uncertainty Quantification in Large Language Models Through Convex Hull Analysis 28 Jun 2024 · 0 repositories · arXiv:2406.19712
-
AutoPureData: Automated Filtering of Undesirable Web Data to Update LLM Knowledge 27 Jun 2024 · 1 repository · arXiv:2406.19271
-
AutoRAG-HP: Automatic Online Hyper-Parameter Tuning for Retrieval-Augmented Generation 27 Jun 2024 · 0 repositories · arXiv:2406.19251
-
Historia Magistra Vitae: Dynamic Topic Modeling of Roman Literature using Neural Embeddings 27 Jun 2024 · 0 repositories · arXiv:2406.18907
-
IndoToxic2024: A Demographically-Enriched Dataset of Hate Speech and Toxicity Types for Indonesian Language 27 Jun 2024 · 0 repositories · arXiv:2406.19349
-
RAVEN: Multitask Retrieval Augmented Vision-Language Learning 27 Jun 2024 · 0 repositories · arXiv:2406.19150
-
SeaKR: Self-aware Knowledge Retrieval for Adaptive Retrieval Augmented Generation 27 Jun 2024 · 1 repository · arXiv:2406.19215Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
Generating Is Believing: Membership Inference Attacks against Retrieval-Augmented Generation 27 Jun 2024 · 0 repositories · arXiv:2406.19234
-
Evaluating Quality of Answers for Retrieval-Augmented Generation: A Strong LLM Is All You Need 26 Jun 2024 · 0 repositories · arXiv:2406.18064
-
"Glue pizza and eat rocks" -- Exploiting Vulnerabilities in Retrieval-Augmented Generative Models 26 Jun 2024 · 0 repositories · arXiv:2406.19417
-
Knowledge graph enhanced retrieval-augmented generation for failure mode and effects analysis 26 Jun 2024 · 1 repository · arXiv:2406.18114
-
Multi-step Inference over Unstructured Data 26 Jun 2024 · 0 repositories · arXiv:2406.17987
-
Poisoned LangChain: Jailbreak LLMs by LangChain 26 Jun 2024 · 0 repositories · arXiv:2406.18122
-
ResumeAtlas: Revisiting Resume Classification with Large-Scale Datasets and Large Language Models 26 Jun 2024 · 1 repository · arXiv:2406.18125
-
Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented Generation 26 Jun 2024 · 1 repository · arXiv:2406.18676Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Zero-shot prompt-based classification: topic labeling in times of foundation models in German Tweets 26 Jun 2024 · 0 repositories · arXiv:2406.18239
-
SetBERT: Enhancing Retrieval Performance for Boolean Logic and Set Operation Queries 25 Jun 2024 · 0 repositories · arXiv:2406.17282
-
CTBench: A Comprehensive Benchmark for Evaluating Language Model Capabilities in Clinical Trial Design 25 Jun 2024 · 1 repository · arXiv:2406.17888
-
LumberChunker: Long-Form Narrative Document Segmentation 25 Jun 2024 · 1 repository · arXiv:2406.17526Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
RAGBench: Explainable Benchmark for Retrieval-Augmented Generation Systems 25 Jun 2024 · 0 repositories · arXiv:2407.11005