Methods › Natural Language Processing › Subword Segmentation › WordPiece › Papers, page 8
WordPiece
Papers archive 2025-07-28
archive papers tagged: 7,063 · with a code link: 2,910 · where Syntology ran a sample: 650 (529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (650 of 7,063 tagged: 529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument)
Page 8 of 71: papers 701 to 800 of 7,063, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
AirRAG: Activating Intrinsic Reasoning for Retrieval Augmented Generation via Tree-based Search 17 Jan 2025 · 0 repositories · arXiv:2501.10053
-
BBPOS: BERT-based Part-of-Speech Tagging for Uzbek 17 Jan 2025 · 0 repositories · arXiv:2501.10107
-
Passage Segmentation of Documents for Extractive Question Answering 17 Jan 2025 · 0 repositories · arXiv:2501.09940
-
On Learning Informative Trajectory Embeddings for Imitation, Classification and Regression 16 Jan 2025 · 1 repository · arXiv:2501.09327
-
Sentiment Analysis in Twitter Social Network Centered on Cryptocurrencies Using Machine Learning 16 Jan 2025 · 0 repositories · arXiv:2501.09777
-
Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG 15 Jan 2025 · 1 repository · arXiv:2501.09136
-
Augmenting Human-Annotated Training Data with Large Language Model Generation and Distillation in Open-Response Assessment 15 Jan 2025 · 0 repositories · arXiv:2501.09126
-
Expanding Vietnamese SentiWordNet to Improve Performance of Vietnamese Sentiment Analysis Models 15 Jan 2025 · 0 repositories · arXiv:2501.08758
-
A Driver Advisory System Based on Large Language Model for High-speed Train 14 Jan 2025 · 0 repositories · arXiv:2501.07837
-
ASTRID -- An Automated and Scalable TRIaD for the Evaluation of RAG-based Clinical Question Answering Systems 14 Jan 2025 · 0 repositories · arXiv:2501.08208
-
Comparative Analysis of Efficient Adapter-Based Fine-Tuning of State-of-the-Art Transformer Models 14 Jan 2025 · 0 repositories · arXiv:2501.08271
-
Eliciting In-context Retrieval and Reasoning for Long-context Large Language Models 14 Jan 2025 · 0 repositories · arXiv:2501.08248
-
Exploring Narrative Clustering in Large Language Models: A Layerwise Analysis of BERT 14 Jan 2025 · 0 repositories · arXiv:2501.08053
-
READ: Reinforcement-based Adversarial Learning for Text Classification with Limited Labeled Data 14 Jan 2025 · 0 repositories · arXiv:2501.08035
-
ReARTeR: Retrieval-Augmented Reasoning with Trustworthy Process Rewarding 14 Jan 2025 · 1 repository · arXiv:2501.07861
-
Enhancing Retrieval-Augmented Generation: A Study of Best Practices 13 Jan 2025 · 1 repository · arXiv:2501.07391Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Enhancing Talent Employment Insights Through Feature Extraction with LLM Finetuning 13 Jan 2025 · 0 repositories · arXiv:2501.07663
-
Parallel Key-Value Cache Fusion for Position Invariant RAG 13 Jan 2025 · 0 repositories · arXiv:2501.07523
-
WebWalker: Benchmarking LLMs in Web Traversal 13 Jan 2025 · 2 repositories · arXiv:2501.07572
-
Eliza: A Web3 friendly AI Agent Operating System 12 Jan 2025 · 2 repositories · arXiv:2501.06781
-
MiniRAG: Towards Extremely Simple Retrieval-Augmented Generation 12 Jan 2025 · 1 repository · arXiv:2501.06713Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
First Token Probability Guided RAG for Telecom Question Answering 11 Jan 2025 · 0 repositories · arXiv:2501.06468
-
CognoSpeak: an automatic, remote assessment of early cognitive decline in real-world conversational speech 10 Jan 2025 · 0 repositories · arXiv:2501.05755
-
Aligning Brain Activity with Advanced Transformer Models: Exploring the Role of Punctuation in Semantic Processing 10 Jan 2025 · 1 repository · arXiv:2501.06278
-
VideoRAG: Retrieval-Augmented Generation over Video Corpus 10 Jan 2025 · 1 repository · arXiv:2501.05874Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
A General Retrieval-Augmented Generation Framework for Multimodal Case-Based Reasoning Applications 9 Jan 2025 · 0 repositories · arXiv:2501.05030
-
Biomedical Relation Extraction via Adaptive Document-Relation Cross-Mapping and Concept Unique Identifier 9 Jan 2025 · 0 repositories · arXiv:2501.05155
-
DisSim-FinBERT: Text Simplification for Core Message Extraction in Complex Financial Texts 9 Jan 2025 · 0 repositories · arXiv:2501.04959
-
Enhancing Plagiarism Detection in Marathi with a Weighted Ensemble of TF-IDF and BERT Embeddings for Low-Resource Language Processing 9 Jan 2025 · 1 repository · arXiv:2501.05260
-
LLMQuoter: Enhancing RAG Capabilities Through Efficient Quote Extraction From Large Contexts 9 Jan 2025 · 1 repository · arXiv:2501.05554
-
RAG-WM: An Efficient Black-Box Watermarking Approach for Retrieval-Augmented Generation of Large Language Models 9 Jan 2025 · 0 repositories · arXiv:2501.05249
-
The more polypersonal the better -- a short look on space geometry of fine-tuned layers 9 Jan 2025 · 0 repositories · arXiv:2501.05503
-
Advancing Retrieval-Augmented Generation for Persian: Development of Language Models, Comprehensive Benchmarks, and Best Practices for Optimization 8 Jan 2025 · 0 repositories · arXiv:2501.04858
-
Integrating LLMs with ITS: Recent Advances, Potentials, Challenges, and Future Directions 8 Jan 2025 · 0 repositories · arXiv:2501.04437
-
Knowledge Retrieval Based on Generative AI 8 Jan 2025 · 0 repositories · arXiv:2501.04635
-
Multi-task retriever fine-tuning for domain-specific and efficient RAG 8 Jan 2025 · 0 repositories · arXiv:2501.04652
-
Quantum-inspired Embeddings Projection and Similarity Metrics for Representation Learning 8 Jan 2025 · 1 repository · arXiv:2501.04591
-
Re-ranking the Context for Multimodal Retrieval Augmented Generation 8 Jan 2025 · 0 repositories · arXiv:2501.04695
-
HP-BERT: A framework for longitudinal study of Hinduphobia on social media via LLMs 7 Jan 2025 · 1 repository · arXiv:2501.05482
-
IntegrityAI at GenAI Detection Task 2: Detecting Machine-Generated Academic Essays in English and Arabic Using ELECTRA and Stylometry 7 Jan 2025 · 0 repositories · arXiv:2501.05476
-
MTRAG: A Multi-Turn Conversational Benchmark for Evaluating Retrieval-Augmented Generation Systems 7 Jan 2025 · 1 repository · arXiv:2501.03468
-
Practical Design and Benchmarking of Generative AI Applications for Surgical Billing and Coding 7 Jan 2025 · 0 repositories · arXiv:2501.05479
-
RAG-Check: Evaluating Multimodal Retrieval Augmented Generation Performance 7 Jan 2025 · 0 repositories · arXiv:2501.03995
-
Reading with Intent -- Neutralizing Intent 7 Jan 2025 · 0 repositories · arXiv:2501.03475
-
Text to Band Gap: Pre-trained Language Models as Encoders for Semiconductor Band Gap Prediction 7 Jan 2025 · 1 repository · arXiv:2501.03456
-
Developing an Artificial Intelligence Tool for Personalized Breast Cancer Treatment Plans based on the NCCN Guidelines 6 Jan 2025 · 0 repositories · arXiv:2502.15698
-
FlippedRAG: Black-Box Opinion Manipulation Adversarial Attacks to Retrieval-Augmented Generation Models 6 Jan 2025 · 0 repositories · arXiv:2501.02968
-
Graph-based Retrieval Augmented Generation for Dynamic Few-shot Text Classification 6 Jan 2025 · 0 repositories · arXiv:2501.02844
-
Multi-Modal One-Shot Federated Ensemble Learning for Medical Data with Vision Large Language Model 6 Jan 2025 · 0 repositories · arXiv:2501.03292
-
Political Events using RAG with LLMs 6 Jan 2025 · 0 repositories · arXiv:2502.15701
-
QuIM-RAG: Advancing Retrieval-Augmented Generation with Inverted Question Matching for Enhanced QA Performance 6 Jan 2025 · 0 repositories · arXiv:2501.02702
-
Tree-based RAG-Agent Recommendation System: A Case Study in Medical Test Data 6 Jan 2025 · 0 repositories · arXiv:2501.02727
-
Context Aware Lemmatization and Morphological Tagging Method in Turkish 4 Jan 2025 · 0 repositories · arXiv:2501.02361
-
Knowledge Graph Retrieval-Augmented Generation for LLM-based Recommendation 4 Jan 2025 · 0 repositories · arXiv:2501.02226
-
LLM Content Moderation and User Satisfaction: Evidence from Response Refusals in Chatbot Arena 4 Jan 2025 · 0 repositories · arXiv:2501.03266
-
BERT4MIMO: A Foundation Model using BERT Architecture for Massive MIMO Channel State Information Prediction 3 Jan 2025 · 1 repository · arXiv:2501.01802
-
GoBERT: Gene Ontology Graph Informed BERT for Universal Gene Function Prediction 3 Jan 2025 · 0 repositories · arXiv:2501.01930
-
Improving Transducer-Based Spoken Language Understanding with Self-Conditioned CTC and Knowledge Transfer 3 Jan 2025 · 0 repositories · arXiv:2501.01936
-
PersonaAI: Leveraging Retrieval-Augmented Generation and Personalized Context for AI-Driven Digital Avatars 3 Jan 2025 · 0 repositories · arXiv:2503.15489
-
An Efficient Attention Mechanism for Sequential Recommendation Tasks: HydraRec 2 Jan 2025 · 0 repositories · arXiv:2501.01242
-
Disambiguation of Chinese Polyphones in an End-to-End Framework with Semantic Features Extracted by Pre-trained BERT 2 Jan 2025 · 0 repositories · arXiv:2501.01102
-
Multi-Head Explainer: A General Framework to Improve Explainability in CNNs and Transformers 2 Jan 2025 · 0 repositories · arXiv:2501.01311
-
Multi-Modal Video Feature Extraction for Popularity Prediction 2 Jan 2025 · 0 repositories · arXiv:2501.01422
-
Beyond Words: AuralLLM and SignMST-C for Precise Sign Language Production and Bidirectional Accessibility 1 Jan 2025 · 0 repositories · arXiv:2501.00765
-
Column Property Annotation using Large Language Models 1 Jan 2025 · 1 repository
-
Decoding the Flow: CauseMotion for Emotional Causality Analysis in Long-form Conversations 1 Jan 2025 · 0 repositories · arXiv:2501.00778
-
Docopilot: Improving Multimodal Models for Document-Level Understanding 1 Jan 2025 · 1 repository
-
DocSAM: Unified Document Image Segmentation via Query Decomposition and Heterogeneous Mixed Learning 1 Jan 2025 · 0 repositories
-
Labels Generated by Large Language Model Helps Measuring People's Empathy in Vitro 1 Jan 2025 · 1 repository · arXiv:2501.00691
-
On Importance of Layer Pruning for Smaller BERT Models and Low Resource Languages 1 Jan 2025 · 0 repositories · arXiv:2501.00733
-
Dementia Detection using Multi-modal Methods on Audio Data 31 Dec 2024 · 0 repositories · arXiv:2501.00465
-
Exploring Variability in Fine-Tuned Models for Text Classification with DistilBERT 31 Dec 2024 · 0 repositories · arXiv:2501.00241
-
MAIN-RAG: Multi-Agent Filtering Retrieval-Augmented Generation 31 Dec 2024 · 0 repositories · arXiv:2501.00332
-
RAG-Instruct: Boosting LLMs with Diverse Retrieval-Augmented Instructions 31 Dec 2024 · 1 repository · arXiv:2501.00353Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 6 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
Retrieval-Augmented Generation with Graphs (GraphRAG) 31 Dec 2024 · 0 repositories · arXiv:2501.00309
-
Plancraft: an evaluation dataset for planning with LLM agents 30 Dec 2024 · 1 repository · arXiv:2412.21033
-
Text Classification: Neural Networks VS Machine Learning Models VS Pre-trained Models 30 Dec 2024 · 0 repositories · arXiv:2412.21022
-
ELECTRA and GPT-4o: Cost-Effective Partners for Sentiment Analysis 29 Dec 2024 · 1 repository · arXiv:2501.00062
-
On Adversarial Robustness of Language Models in Transfer Learning 29 Dec 2024 · 0 repositories · arXiv:2501.00066
-
Understanding the Impact of Confidence in Retrieval Augmented Generation: A Case Study in the Medical Domain 29 Dec 2024 · 1 repository · arXiv:2412.20309
-
Assessing Text Classification Methods for Cyberbullying Detection on Social Media Platforms 27 Dec 2024 · 0 repositories · arXiv:2412.19928
-
Gradient Weight-normalized Low-rank Projection for Efficient LLM Training 27 Dec 2024 · 1 repository · arXiv:2412.19616
-
Long Context vs. RAG for LLMs: An Evaluation and Revisits 27 Dec 2024 · 1 repository · arXiv:2501.01880
-
Text2Insight: Transform natural language text into insights seamlessly using multi-model architecture 27 Dec 2024 · 0 repositories · arXiv:2412.19718
-
Indonesian-English Code-Switching Speech Synthesizer Utilizing Multilingual STEN-TTS and Bert LID 26 Dec 2024 · 0 repositories · arXiv:2412.19043
-
RAG with Differential Privacy 26 Dec 2024 · 1 repository · arXiv:2412.19291
-
Sentiment trading with large language models 26 Dec 2024 · 0 repositories · arXiv:2412.19245
-
Injecting Bias into Text Classification Models using Backdoor Attacks 25 Dec 2024 · 0 repositories · arXiv:2412.18975
-
Optimizing Large Language Models with an Enhanced LoRA Fine-Tuning Algorithm for Efficiency and Robustness in NLP Tasks 25 Dec 2024 · 0 repositories · arXiv:2412.18729
-
Resource-Efficient Transformer Architecture: Optimizing Memory and Execution Time for Real-Time Applications 25 Dec 2024 · 0 repositories · arXiv:2501.00042
-
Comprehensive Assessment of BERT-Based Methods for Predicting Antimicrobial Peptides 24 Dec 2024 · 1 repository
-
GeAR: Graph-enhanced Agent for Retrieval-augmented Generation 24 Dec 2024 · 0 repositories · arXiv:2412.18431
-
Improving Factuality with Explicit Working Memory 24 Dec 2024 · 0 repositories · arXiv:2412.18069
-
Molly: Making Large Language Model Agents Solve Python Problem More Logically 24 Dec 2024 · 0 repositories · arXiv:2412.18093
-
Pirates of the RAG: Adaptively Attacking LLMs to Leak Knowledge Bases 24 Dec 2024 · 0 repositories · arXiv:2412.18295
-
Research on the Proximity Relationships of Psychosomatic Disease Knowledge Graph Modules Extracted by Large Language Models 24 Dec 2024 · 0 repositories · arXiv:2412.18419
-
Unlocking the Potential of Multiple BERT Models for Bangla Question Answering in NCTB Textbooks 24 Dec 2024 · 0 repositories · arXiv:2412.18440
-
A Survey of Query Optimization in Large Language Models 23 Dec 2024 · 0 repositories · arXiv:2412.17558
-
Comparative Analysis of Document-Level Embedding Methods for Similarity Scoring on Shakespeare Sonnets and Taylor Swift Lyrics 23 Dec 2024 · 0 repositories · arXiv:2412.17552
-
Efficient fine-tuning methodology of text embedding models for information retrieval: contrastive learning penalty (clp) 23 Dec 2024 · 1 repository · arXiv:2412.17364