Methods › Natural Language Processing › Subword Segmentation › WordPiece › Papers, page 15
WordPiece
Papers archive 2025-07-28
archive papers tagged: 7,063 · with a code link: 2,910 · where Syntology ran a sample: 650 (529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (650 of 7,063 tagged: 529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument)
Page 15 of 71: papers 1,401 to 1,500 of 7,063, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Enhancing Visual Question Answering through Ranking-Based Hybrid Training and Multimodal Fusion 14 Aug 2024 · 0 repositories · arXiv:2408.07303
-
Exploring Retrieval Augmented Generation in Arabic 14 Aug 2024 · 1 repository · arXiv:2408.07425
-
LiPCoT: Linear Predictive Coding based Tokenizer for Self-supervised Learning of Time Series Data via Language Models 14 Aug 2024 · 1 repository · arXiv:2408.07292
-
Transformers and Large Language Models for Efficient Intrusion Detection Systems: A Comprehensive Survey 14 Aug 2024 · 0 repositories · arXiv:2408.07583
-
BERT's Conceptual Cartography: Mapping the Landscapes of Meaning 13 Aug 2024 · 0 repositories · arXiv:2408.07190
-
Pragmatic inference of scalar implicature by LLMs 13 Aug 2024 · 0 repositories · arXiv:2408.06673
-
TableGuard -- Securing Structured & Unstructured Data 13 Aug 2024 · 0 repositories · arXiv:2408.07045
-
Bayesian inference to improve quality of Retrieval Augmented Generation 12 Aug 2024 · 0 repositories · arXiv:2408.08901
-
Improving Structural Diversity of Blackbox LLMs via Chain-of-Specification Prompting 12 Aug 2024 · 0 repositories · arXiv:2408.06186
-
LOLgorithm: Integrating Semantic,Syntactic and Contextual Elements for Humor Classification 12 Aug 2024 · 0 repositories · arXiv:2408.06335
-
Optimizing RAG Techniques for Automotive Industry PDF Chatbots: A Case Study with Locally Deployed Ollama Models 12 Aug 2024 · 0 repositories · arXiv:2408.05933
-
The Language of Trauma: Modeling Traumatic Event Descriptions Across Domains with Explainable AI 12 Aug 2024 · 0 repositories · arXiv:2408.05977
-
PhishLang: A Real-Time, Fully Client-Side Phishing Detection Framework Using MobileBERT 11 Aug 2024 · 2 repositories · arXiv:2408.05667
-
A Hybrid RAG System with Comprehensive Enhancement on Complex Reasoning 9 Aug 2024 · 0 repositories · arXiv:2408.05141
-
ConfusedPilot: Confused Deputy Risks in RAG-based LLMs 9 Aug 2024 · 0 repositories · arXiv:2408.04870
-
Ensemble BERT: A student social network text sentiment classification model based on ensemble learning and BERT architecture 9 Aug 2024 · 0 repositories · arXiv:2408.04849
-
HybridRAG: Integrating Knowledge Graphs and Vector Retrieval Augmented Generation for Efficient Information Extraction 9 Aug 2024 · 0 repositories · arXiv:2408.04948
-
Rag and Roll: An End-to-End Evaluation of Indirect Prompt Manipulations in LLM-based Application Frameworks 9 Aug 2024 · 0 repositories · arXiv:2408.05025
-
Analysis of Argument Structure Constructions in the Large Language Model BERT 8 Aug 2024 · 0 repositories · arXiv:2408.04270
-
EfficientRAG: Efficient Retriever for Multi-Hop Question Answering 8 Aug 2024 · 1 repository · arXiv:2408.04259Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Hybrid Student-Teacher Large Language Model Refinement for Cancer Toxicity Symptom Extraction 8 Aug 2024 · 0 repositories · arXiv:2408.04775
-
Medical Graph RAG: Towards Safe Medical Large Language Model via Graph Retrieval-Augmented Generation 8 Aug 2024 · 1 repository · arXiv:2408.04187Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
SCENE: Evaluating Explainable AI Techniques Using Soft Counterfactuals 8 Aug 2024 · 0 repositories · arXiv:2408.04575
-
A Comparison of LLM Finetuning Methods & Evaluation Metrics with Travel Chatbot Use Case 7 Aug 2024 · 0 repositories · arXiv:2408.03562
-
VulScribeR: Exploring RAG-based Vulnerability Augmentation with LLMs 7 Aug 2024 · 1 repository · arXiv:2408.04125
-
Is Child-Directed Speech Effective Training Data for Language Models? 7 Aug 2024 · 1 repository · arXiv:2408.03617Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 1 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples) · 2 pointer-only (licence)
-
MaxMind: A Memory Loop Network to Enhance Software Productivity based on Large Language Models 7 Aug 2024 · 0 repositories · arXiv:2408.03841
-
Topic Modeling with Fine-tuning LLMs and Bag of Sentences 6 Aug 2024 · 1 repository · arXiv:2408.03099Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Evaluating the Translation Performance of Large Language Models Based on Euas-20 6 Aug 2024 · 0 repositories · arXiv:2408.03119
-
Leveraging Parameter Efficient Training Methods for Low Resource Text Classification: A Case Study in Marathi 6 Aug 2024 · 0 repositories · arXiv:2408.03172
-
Training LLMs to Recognize Hedges in Spontaneous Narratives 6 Aug 2024 · 1 repository · arXiv:2408.03319
-
RAG Foundry: A Framework for Enhancing LLMs for Retrieval Augmented Generation 5 Aug 2024 · 2 repositories · arXiv:2408.02545
-
Wiping out the limitations of Large Language Models -- A Taxonomy for Retrieval Augmented Generation 5 Aug 2024 · 0 repositories · arXiv:2408.02854
-
AppAgent v2: Advanced Agent for Flexible Mobile Interactions 5 Aug 2024 · 0 repositories · arXiv:2408.11824
-
LLM Agents Improve Semantic Code Search 5 Aug 2024 · 0 repositories · arXiv:2408.11058
-
Defining and Evaluating Decision and Composite Risk in Language Models Applied to Natural Language Inference 4 Aug 2024 · 0 repositories · arXiv:2408.01935
-
Indexing and Visualization of Climate Change Narratives Using BERT and Causal Extraction 3 Aug 2024 · 0 repositories · arXiv:2408.01745
-
Tracking Emotional Dynamics in Chat Conversations: A Hybrid Approach using DistilBERT and Emoji Sentiment Analysis 3 Aug 2024 · 0 repositories · arXiv:2408.01838
-
Cross-domain Named Entity Recognition via Graph Matching 2 Aug 2024 · 0 repositories · arXiv:2408.00981
-
BioRAG: A RAG-LLM Framework for Biological Question Reasoning 2 Aug 2024 · 0 repositories · arXiv:2408.01107
-
RAGEval: Scenario Specific RAG Evaluation Dataset Generation Framework 2 Aug 2024 · 1 repository · arXiv:2408.01262
-
Evaluating the Impact of Advanced LLM Techniques on AI-Lecture Tutors for a Robotics Course 2 Aug 2024 · 0 repositories · arXiv:2408.04645
-
Improving Retrieval-Augmented Generation in Medicine with Iterative Follow-up Questions 1 Aug 2024 · 1 repository · arXiv:2408.00727
-
Guiding Sentiment Analysis with Hierarchical Text Clustering: Analyzing the German X/Twitter Discourse on Face Masks in the 2020 COVID-19 Pandemic 1 Aug 2024 · 1 repository
-
Ontological Relations from Word Embeddings 1 Aug 2024 · 0 repositories · arXiv:2408.00444
-
Multi-Level Querying using A Knowledge Pyramid 31 Jul 2024 · 0 repositories · arXiv:2407.21276
-
SAKR: Enhancing Retrieval-Augmented Generation via Streaming Algorithm and K-Means Clustering 31 Jul 2024 · 0 repositories · arXiv:2407.21300
-
MetaOpenFOAM: an LLM-based multi-agent framework for CFD 31 Jul 2024 · 1 repository · arXiv:2407.21320Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Adaptive Retrieval-Augmented Generation for Conversational Systems 31 Jul 2024 · 0 repositories · arXiv:2407.21712
-
Recording First-person Experiences to Build a New Type of Foundation Model 31 Jul 2024 · 0 repositories · arXiv:2408.02680
-
A New Type of Foundation Model Based on Recordings of People's Emotions and Physiology 31 Jul 2024 · 0 repositories · arXiv:2408.00030
-
Event-Arguments Extraction Corpus and Modeling using BERT for Arabic 30 Jul 2024 · 0 repositories · arXiv:2407.21153
-
BERT and LLMs-Based avGFP Brightness Prediction and Mutation Design 30 Jul 2024 · 0 repositories · arXiv:2407.20534
-
CLEFT: Language-Image Contrastive Learning with Efficient Large Language Model and Prompt Fine-Tuning 30 Jul 2024 · 1 repository · arXiv:2407.21011
-
A Study on the Implementation Method of an Agent-Based Advanced RAG System Using Graph 29 Jul 2024 · 0 repositories · arXiv:2407.19994
-
AutoScale: Scale-Aware Data Mixing for Pre-Training LLMs 29 Jul 2024 · 1 repository · arXiv:2407.20177Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Enhancing Code Translation in Language Models with Few-Shot Learning via Retrieval-Augmented Generation 29 Jul 2024 · 0 repositories · arXiv:2407.19619
-
Introducing a new hyper-parameter for RAG: Context Window Utilization 29 Jul 2024 · 0 repositories · arXiv:2407.19794
-
Sentiment Analysis of Lithuanian Online Reviews Using Large Language Models 29 Jul 2024 · 0 repositories · arXiv:2407.19914
-
Faculty Perspectives on the Potential of RAG in Computer Science Higher Education 28 Jul 2024 · 0 repositories · arXiv:2408.01462
-
Exploring Genre and Success Classification through Song Lyrics using DistilBERT: A Fun NLP Venture 28 Jul 2024 · 0 repositories · arXiv:2407.21068
-
Motamot: A Dataset for Revealing the Supremacy of Large Language Models over Transformer Models in Bengali Political Sentiment Analysis 28 Jul 2024 · 1 repository · arXiv:2407.19528
-
FarSSiBERT: A Novel Transformer-based Model for Semantic Similarity Measurement of Persian Social Networks Informal Texts 27 Jul 2024 · 0 repositories · arXiv:2407.19173
-
Modular RAG: Transforming RAG Systems into LEGO-like Reconfigurable Frameworks 26 Jul 2024 · 0 repositories · arXiv:2407.21059
-
Human-artificial intelligence teaming for scientific information extraction from data-driven additive manufacturing research using large language models 26 Jul 2024 · 0 repositories · arXiv:2407.18827
-
ClinicRealm: Re-evaluating Large Language Models with Conventional Machine Learning for Non-Generative Clinical Prediction Tasks 26 Jul 2024 · 1 repository · arXiv:2407.18525Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 1 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
MistralBSM: Leveraging Mistral-7B for Vehicular Networks Misbehavior Detection 26 Jul 2024 · 0 repositories · arXiv:2407.18462
-
REAPER: Reasoning based Retrieval Planning for Complex RAG Systems 26 Jul 2024 · 0 repositories · arXiv:2407.18553
-
Banyan: Improved Representation Learning with Explicit Structure 25 Jul 2024 · 0 repositories · arXiv:2407.17771
-
RoBERTa, ResNeXt and BiLSTM with self-attention: The ultimate trio for customer sentiment analysis 25 Jul 2024 · 0 repositories
-
The Geometry of Queries: Query-Based Innovations in Retrieval-Augmented Generation 25 Jul 2024 · 0 repositories · arXiv:2407.18044
-
Understanding the Interplay of Scale, Data, and Bias in Language Models: A Case Study with BERT 25 Jul 2024 · 0 repositories · arXiv:2407.21058
-
What Matters in Explanations: Towards Explainable Fake Review Detection Focusing on Transformers 24 Jul 2024 · 0 repositories · arXiv:2407.21056
-
A Comprehensive Approach to Misspelling Correction with BERT and Levenshtein Distance 24 Jul 2024 · 0 repositories · arXiv:2407.17383
-
A Novel Two-Step Fine-Tuning Pipeline for Cold-Start Active Learning in Text Classification Tasks 24 Jul 2024 · 0 repositories · arXiv:2407.17284
-
Bailicai: A Domain-Optimized Retrieval-Augmented Generation Framework for Medical Applications 24 Jul 2024 · 0 repositories · arXiv:2407.21055
-
Artificial Intelligence in Extracting Diagnostic Data from Dental Records 23 Jul 2024 · 0 repositories · arXiv:2407.21050
-
LawLuo: A Multi-Agent Collaborative Framework for Multi-Round Chinese Legal Consultation 23 Jul 2024 · 0 repositories · arXiv:2407.16252
-
Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid Approach 23 Jul 2024 · 0 repositories · arXiv:2407.16833
-
TookaBERT: A Step Forward for Persian NLU 23 Jul 2024 · 0 repositories · arXiv:2407.16382
-
An Empirical Comparison of Video Frame Sampling Methods for Multi-Modal RAG Retrieval 22 Jul 2024 · 0 repositories · arXiv:2408.03340
-
Customized Retrieval Augmented Generation and Benchmarking for EDA Tool Documentation QA 22 Jul 2024 · 0 repositories · arXiv:2407.15353
-
Inverted Activations: Reducing Memory Footprint in Neural Network Training 22 Jul 2024 · 1 repository · arXiv:2407.15545
-
LLMmap: Fingerprinting For Large Language Models 22 Jul 2024 · 1 repository · arXiv:2407.15847Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 7 unverified (of 16 harvested samples)
-
MoRSE: Bridging the Gap in Cybersecurity Expertise with Retrieval Augmented Generation 22 Jul 2024 · 0 repositories · arXiv:2407.15748
-
RadioRAG: Factual large language models for enhanced diagnostics in radiology using online retrieval augmented generation 22 Jul 2024 · 1 repository · arXiv:2407.15621
-
ZZU-NLP at SIGHAN-2024 dimABSA Task: Aspect-Based Sentiment Analysis with Coarse-to-Fine In-context Learning 22 Jul 2024 · 0 repositories · arXiv:2407.15341
-
A multi-level multi-label text classification dataset of 19th century Ottoman and Russian literary and critical texts 21 Jul 2024 · 0 repositories · arXiv:2407.15136
-
Golden-Retriever: High-Fidelity Agentic Retrieval Augmented Generation for Industrial Knowledge Base 20 Jul 2024 · 0 repositories · arXiv:2408.00798
-
Automatic Generation of Fashion Images using Prompting in Generative Machine Learning Models 20 Jul 2024 · 1 repository · arXiv:2407.14944
-
Differential Privacy of Cross-Attention with Provable Guarantee 20 Jul 2024 · 0 repositories · arXiv:2407.14717
-
Adversarial Databases Improve Success in Retrieval-based Large Language Models 19 Jul 2024 · 0 repositories · arXiv:2407.14609
-
ChatQA 2: Bridging the Gap to Proprietary LLMs in Long Context and RAG Capabilities 19 Jul 2024 · 0 repositories · arXiv:2407.14482
-
Unipa-GPT: Large Language Models for university-oriented QA in Italian 19 Jul 2024 · 1 repository · arXiv:2407.14246
-
Black-Box Opinion Manipulation Attacks to Retrieval-Augmented Generation of Large Language Models 18 Jul 2024 · 0 repositories · arXiv:2407.13757
-
Can Open-Source LLMs Compete with Commercial Models? Exploring the Few-Shot Performance of Current GPT Models in Biomedical Tasks 18 Jul 2024 · 1 repository · arXiv:2407.13511
-
Evaluating Large Language Models for Anxiety and Depression Classification using Counseling and Psychotherapy Transcripts 18 Jul 2024 · 1 repository · arXiv:2407.13228
-
PRAGyan -- Connecting the Dots in Tweets 18 Jul 2024 · 0 repositories · arXiv:2407.13909
-
Qalam : A Multimodal LLM for Arabic Optical Character and Handwriting Recognition 18 Jul 2024 · 0 repositories · arXiv:2407.13559
-
Reconstruct the Pruned Model without Any Retraining 18 Jul 2024 · 0 repositories · arXiv:2407.13331