Methods › Natural Language Processing › Subword Segmentation › WordPiece › Papers, page 14
WordPiece
Papers archive 2025-07-28
archive papers tagged: 7,063 · with a code link: 2,910 · where Syntology ran a sample: 650 (529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (650 of 7,063 tagged: 529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument)
Page 14 of 71: papers 1,301 to 1,400 of 7,063, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Comparing Retrieval-Augmentation and Parameter-Efficient Fine-Tuning for Privacy-Preserving Personalization of Large Language Models 14 Sep 2024 · 1 repository · arXiv:2409.09510
-
LLM-Powered Ensemble Learning for Paper Source Tracing: A GPU-Free Approach 14 Sep 2024 · 1 repository · arXiv:2409.09383
-
A RAG Approach for Generating Competency Questions in Ontology Engineering 13 Sep 2024 · 0 repositories · arXiv:2409.08820
-
DomURLs_BERT: Pre-trained BERT-based Model for Malicious Domains and URLs Detection and Classification 13 Sep 2024 · 1 repository · arXiv:2409.09143
-
Exploring Information Retrieval Landscapes: An Investigation of a Novel Evaluation Techniques and Comparative Document Splitting Methods 13 Sep 2024 · 1 repository · arXiv:2409.08479
-
Winning Solution For Meta KDD Cup' 24 13 Sep 2024 · 0 repositories · arXiv:2410.00005
-
AudioBERT: Audio Knowledge Augmented Language Model 12 Sep 2024 · 1 repository · arXiv:2409.08199
-
Enhanced Online Grooming Detection Employing Context Determination and Message-Level Analysis 12 Sep 2024 · 0 repositories · arXiv:2409.07958
-
OmniQuery: Contextually Augmenting Captured Multimodal Memory to Enable Personal Question Answering 12 Sep 2024 · 0 repositories · arXiv:2409.08250
-
On the Vulnerability of Applying Retrieval-Augmented Generation within Knowledge-Intensive Application Domains 12 Sep 2024 · 0 repositories · arXiv:2409.17275
-
Unleashing Worms and Extracting Data: Escalating the Outcome of Attacks against RAG-based Inference in Scale and Severity Using Jailbreaking 12 Sep 2024 · 1 repository · arXiv:2409.08045
-
Bio-Eng-LMM AI Assist chatbot: A Comprehensive Tool for Research and Education 11 Sep 2024 · 1 repository · arXiv:2409.07110
-
Cross-Dialect Text-To-Speech in Pitch-Accent Language Incorporating Multi-Dialect Phoneme-Level BERT 11 Sep 2024 · 0 repositories · arXiv:2409.07265
-
Integrating SPARQL and LLMs for Question Answering over Scholarly Data Sources 11 Sep 2024 · 0 repositories · arXiv:2409.18969
-
Towards Fairer Health Recommendations: finding informative unbiased samples via Word Sense Disambiguation 11 Sep 2024 · 0 repositories · arXiv:2409.07424
-
KAG: Boosting LLMs in Professional Domains via Knowledge Augmented Generation 10 Sep 2024 · 1 repository · arXiv:2409.13731Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
GroUSE: A Benchmark to Evaluate Evaluators in Grounded Question Answering 10 Sep 2024 · 1 repository · arXiv:2409.06595
-
A Small Claims Court for the NLP: Judging Legal Text Classification Strategies With Small Datasets 9 Sep 2024 · 0 repositories · arXiv:2409.05972
-
Application Specific Compression of Deep Learning Models 9 Sep 2024 · 1 repository · arXiv:2409.05368
-
MemoRAG: Moving towards Next-Gen RAG Via Memory-Inspired Knowledge Discovery 9 Sep 2024 · 1 repository · arXiv:2409.05591Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples) · 6 pointer-only (licence)
-
QiBERT -- Classifying Online Conversations Messages with BERT as a Feature 9 Sep 2024 · 0 repositories · arXiv:2409.05530
-
Revisiting the Solution of Meta KDD Cup 2024: CRAG 9 Sep 2024 · 1 repository · arXiv:2409.15337
-
OneGen: Efficient One-Pass Unified Generation and Retrieval for LLMs 8 Sep 2024 · 1 repository · arXiv:2409.05152
-
Constrained Multi-Layer Contrastive Learning for Implicit Discourse Relationship Recognition 7 Sep 2024 · 0 repositories · arXiv:2409.13716
-
Column Vocabulary Association (CVA): semantic interpretation of dataless tables 6 Sep 2024 · 0 repositories · arXiv:2409.13709
-
Hermes: Memory-Efficient Pipeline Inference for Large Models on Edge Devices 6 Sep 2024 · 0 repositories · arXiv:2409.04249
-
Protein sequence classification using natural language processing techniques 6 Sep 2024 · 0 repositories · arXiv:2409.04491
-
A Different Level Text Protection Mechanism With Differential Privacy 5 Sep 2024 · 0 repositories · arXiv:2409.03707
-
CA-BERT: Leveraging Context Awareness for Enhanced Multi-Turn Chat Interaction 5 Sep 2024 · 0 repositories · arXiv:2409.13701
-
CACER: Clinical Concept Annotations for Cancer Events and Relations 5 Sep 2024 · 1 repository · arXiv:2409.03905
-
MARAGS: A Multi-Adapter System for Multi-Task Retrieval Augmented Generation Question Answering 5 Sep 2024 · 0 repositories · arXiv:2409.03171
-
RAG based Question-Answering for Contextual Response Prediction System 5 Sep 2024 · 0 repositories · arXiv:2409.03708
-
Revolutionizing Database Q&A with Large Language Models: Comprehensive Benchmark and Evaluation 5 Sep 2024 · 1 repository · arXiv:2409.04475
-
A Data Selection Approach for Enhancing Low Resource Machine Translation Using Cross-Lingual Sentence Representations 4 Sep 2024 · 0 repositories · arXiv:2409.02712
-
GenDFIR: Advancing Cyber Incident Timeline Analysis Through Retrieval Augmented Generation and Large Language Models 4 Sep 2024 · 0 repositories · arXiv:2409.02572
-
Detecting Calls to Action in Multimodal Content: Analysis of the 2021 German Federal Election Campaign on Instagram 4 Sep 2024 · 0 repositories · arXiv:2409.02690
-
Diversify-verify-adapt: Efficient and Robust Retrieval-Augmented Ambiguous Question Answering 4 Sep 2024 · 0 repositories · arXiv:2409.02361
-
How Privacy-Savvy Are Large Language Models? A Case Study on Compliance and Privacy Technical Review 4 Sep 2024 · 0 repositories · arXiv:2409.02375
-
MoA is All You Need: Building LLM Research Team using Mixture of Agents 4 Sep 2024 · 0 repositories · arXiv:2409.07487
-
OpenFact at CheckThat! 2024: Combining Multiple Attack Methods for Effective Adversarial Text Generation 4 Sep 2024 · 0 repositories · arXiv:2409.02649
-
Pre-training data selection for biomedical domain adaptation using journal impact metrics 4 Sep 2024 · 0 repositories · arXiv:2409.02725
-
Robust Text-to-Cypher Using Combination of BERT, GraphSAGE, and Transformer (CoBGT) Model 4 Sep 2024 · 0 repositories
-
Multi-Source Knowledge Pruning for Retrieval-Augmented Generation: A Benchmark and Empirical Study 3 Sep 2024 · 2 repositories · arXiv:2409.13694
-
You Only Use Reactive Attention Slice For Long Context Retrieval 3 Sep 2024 · 1 repository · arXiv:2409.13695
-
AdaComp: Extractive Context Compression with Adaptive Predictor for Retrieval-Augmented Large Language Models 3 Sep 2024 · 0 repositories · arXiv:2409.01579
-
BEAVER: An Enterprise Benchmark for Text-to-SQL 3 Sep 2024 · 0 repositories · arXiv:2409.02038
-
Benchmarking Cognitive Domains for LLMs: Insights from Taiwanese Hakka Culture 3 Sep 2024 · 0 repositories · arXiv:2409.01556
-
In Defense of RAG in the Era of Long-Context Language Models 3 Sep 2024 · 0 repositories · arXiv:2409.01666
-
Pairing Analogy-Augmented Generation with Procedural Memory for Procedural Q&A 2 Sep 2024 · 1 repository · arXiv:2409.01344
-
Modeling Text-Label Alignment for Hierarchical Text Classification 1 Sep 2024 · 1 repository · arXiv:2409.00788
-
The Design of an LLM-powered Unstructured Analytics System 1 Sep 2024 · 0 repositories · arXiv:2409.00847
-
GenAI-powered Multi-Agent Paradigm for Smart Urban Mobility: Opportunities and Challenges for Integrating Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) with Intelligent Transportation Systems 31 Aug 2024 · 0 repositories · arXiv:2409.00494
-
Improving Extraction of Clinical Event Contextual Properties from Electronic Health Records: A Comparative Study 30 Aug 2024 · 0 repositories · arXiv:2408.17181
-
MaFeRw: Query Rewriting with Multi-Aspect Feedbacks for Retrieval-Augmented Large Language Models 30 Aug 2024 · 1 repository · arXiv:2408.17072
-
RISSOLE: Parameter-efficient Diffusion Models via Block-wise Generation and Retrieval-Guidance 30 Aug 2024 · 0 repositories · arXiv:2408.17095
-
Assessing Large Language Models for Online Extremism Research: Identification, Explanation, and New Knowledge 29 Aug 2024 · 0 repositories · arXiv:2408.16749
-
HyPA-RAG: A Hybrid Parameter Adaptive Retrieval-Augmented Generation System for AI Legal and Policy Applications 29 Aug 2024 · 0 repositories · arXiv:2409.09046
-
A Simple Baseline with Single-encoder for Referring Image Segmentation 28 Aug 2024 · 0 repositories · arXiv:2408.15521
-
An Extremely Data-efficient and Generative LLM-based Reinforcement Learning Agent for Recommenders 28 Aug 2024 · 0 repositories · arXiv:2408.16032
-
CBF-LLM: Safe Control for LLM Alignment 28 Aug 2024 · 1 repository · arXiv:2408.15625
-
Conan-embedding: General Text Embedding with More and Better Negative Samples 28 Aug 2024 · 0 repositories · arXiv:2408.15710
-
Is Personality Prediction Possible Based on Reddit Comments? 28 Aug 2024 · 0 repositories · arXiv:2408.16089
-
LRP4RAG: Detecting Hallucinations in Retrieval-Augmented Generation via Layer-wise Relevance Propagation 28 Aug 2024 · 3 repositories · arXiv:2408.15533
-
Into the Unknown Unknowns: Engaged Human Learning through Participation in Language Model Agent Conversations 27 Aug 2024 · 1 repository · arXiv:2408.15232
-
Question answering system of bridge design specification based on large language model 26 Aug 2024 · 1 repository · arXiv:2408.13282
-
LowCLIP: Adapting the CLIP Model Architecture for Low-Resource Languages in Multimodal Image Retrieval Task 25 Aug 2024 · 0 repositories · arXiv:2408.13909
-
Pandora's Box or Aladdin's Lamp: A Comprehensive Analysis Revealing the Role of RAG Noise in Large Language Models 24 Aug 2024 · 1 repository · arXiv:2408.13533Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
GRATR: Zero-Shot Evidence Graph Retrieval-Augmented Trustworthiness Reasoning 22 Aug 2024 · 1 repository · arXiv:2408.12333
-
Large Language Models as Foundations for Next-Gen Dense Retrieval: A Comprehensive Empirical Assessment 22 Aug 2024 · 0 repositories · arXiv:2408.12194
-
LLMs are not Zero-Shot Reasoners for Biomedical Information Extraction 22 Aug 2024 · 0 repositories · arXiv:2408.12249
-
Unlearning Trojans in Large Language Models: A Comparison Between Natural Language and Source Code 22 Aug 2024 · 0 repositories · arXiv:2408.12416
-
A Quick, trustworthy spectral knowledge Q&A system leveraging retrieval-augmented generation on LLM 21 Aug 2024 · 1 repository · arXiv:2408.11557
-
Ancient Wisdom, Modern Tools: Exploring Retrieval-Augmented LLMs for Ancient Indian Philosophy 21 Aug 2024 · 1 repository · arXiv:2408.11903
-
WeQA: A Benchmark for Retrieval Augmented Generation in Wind Energy Domain 21 Aug 2024 · 0 repositories · arXiv:2408.11800
-
RAG-Optimized Tibetan Tourism LLMs: Enhancing Accuracy and Personalization 21 Aug 2024 · 0 repositories · arXiv:2408.12003
-
RAGLAB: A Modular and Research-Oriented Unified Framework for Retrieval-Augmented Generation 21 Aug 2024 · 1 repository · arXiv:2408.11381
-
The Self-Contained Negation Test Set 21 Aug 2024 · 0 repositories · arXiv:2408.11469
-
Language Modeling on Tabular Data: A Survey of Foundations, Techniques and Evolution 20 Aug 2024 · 1 repository · arXiv:2408.10548
-
Reading with Intent 20 Aug 2024 · 0 repositories · arXiv:2408.11189
-
Reconciling Methodological Paradigms: Employing Large Language Models as Novice Qualitative Research Assistants in Talent Management Research 20 Aug 2024 · 0 repositories · arXiv:2408.11043
-
A Strategy to Combine 1stGen Transformers and Open LLMs for Automatic Text Classification 19 Aug 2024 · 0 repositories · arXiv:2408.09629
-
Acquiring Bidirectionality via Large and Small Language Models 19 Aug 2024 · 1 repository · arXiv:2408.09640
-
Active Learning for Identifying Disaster-Related Tweets: A Comparison with Keyword Filtering and Generic Fine-Tuning 19 Aug 2024 · 0 repositories · arXiv:2408.09914
-
Carbon Footprint Accounting Driven by Large Language Models and Retrieval-augmented Generation 19 Aug 2024 · 0 repositories · arXiv:2408.09713
-
Enhanced document retrieval with topic embeddings 19 Aug 2024 · 0 repositories · arXiv:2408.10435
-
LegalBench-RAG: A Benchmark for Retrieval-Augmented Generation in the Legal Domain 19 Aug 2024 · 1 repository · arXiv:2408.10343Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples)
-
SSDTrain: An Activation Offloading Framework to SSDs for Faster Large Language Model Training 19 Aug 2024 · 0 repositories · arXiv:2408.10013
-
Agentic Retrieval-Augmented Generation for Time Series Analysis 18 Aug 2024 · 0 repositories · arXiv:2408.14484
-
Sentiment analysis of preservice teachers' reflections using a large language model 17 Aug 2024 · 0 repositories · arXiv:2408.11862
-
TC-RAG:Turing-Complete RAG's Case study on Medical LLM Systems 17 Aug 2024 · 2 repositories · arXiv:2408.09199
-
CIKMar: A Dual-Encoder Approach to Prompt-Based Reranking in Educational Dialogue Systems 16 Aug 2024 · 0 repositories · arXiv:2408.08805
-
CommunityKG-RAG: Leveraging Community Structures in Knowledge Graphs for Advanced Retrieval-Augmented Generation in Fact-Checking 16 Aug 2024 · 1 repository · arXiv:2408.08535
-
Improving VTE Identification through Language Models from Radiology Reports: A Comparative Study of Mamba, Phi-3 Mini, and BERT 16 Aug 2024 · 0 repositories · arXiv:2408.09043
-
Meta Knowledge for Retrieval Augmented Large Language Models 16 Aug 2024 · 0 repositories · arXiv:2408.09017
-
Quantifying the Effectiveness of Student Organization Activities using Natural Language Processing 16 Aug 2024 · 0 repositories · arXiv:2408.08694
-
VERA: Validation and Evaluation of Retrieval-Augmented Systems 16 Aug 2024 · 0 repositories · arXiv:2409.03759
-
Graph Retrieval-Augmented Generation: A Survey 15 Aug 2024 · 1 repository · arXiv:2408.08921
-
Plan with Code: Comparing approaches for robust NL to DSL generation 15 Aug 2024 · 0 repositories · arXiv:2408.08335
-
RAGChecker: A Fine-grained Framework for Diagnosing Retrieval-Augmented Generation 15 Aug 2024 · 1 repository · arXiv:2408.08067Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples)
-
DataVisT5: A Pre-trained Language Model for Jointly Understanding Text and Data Visualization 14 Aug 2024 · 1 repository · arXiv:2408.07401