Methods › Natural Language Processing › Subword Segmentation › WordPiece › Papers, page 18
WordPiece
Papers archive 2025-07-28
archive papers tagged: 7,063 · with a code link: 2,910 · where Syntology ran a sample: 650 (529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (650 of 7,063 tagged: 529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument)
Page 18 of 71: papers 1,701 to 1,800 of 7,063, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Mix-of-Granularity: Optimize the Chunking Granularity for Retrieval-Augmented Generation 1 Jun 2024 · 1 repository · arXiv:2406.00456
-
Pseudo-label Based Domain Adaptation for Zero-Shot Text Steganalysis 1 Jun 2024 · 0 repositories · arXiv:2406.18565
-
RoBERTa-BiLSTM: A Context-Aware Hybrid Model for Sentiment Analysis 1 Jun 2024 · 1 repository · arXiv:2406.00367
-
A comparison of correspondence analysis with PMI-based word embedding methods 31 May 2024 · 1 repository · arXiv:2405.20895
-
Bi-Directional Transformers vs. word2vec: Discovering Vulnerabilities in Lifted Compiled Code 31 May 2024 · 0 repositories · arXiv:2405.20611
-
Effect of antibody levels on the spread of disease in multiple infections 31 May 2024 · 0 repositories · arXiv:2405.20702
-
Enhancing Noise Robustness of Retrieval-Augmented Language Models with Adaptive Adversarial Training 31 May 2024 · 1 repository · arXiv:2405.20978Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
RAG Does Not Work for Enterprises 31 May 2024 · 0 repositories · arXiv:2406.04369
-
Retrieval Meets Reasoning: Even High-school Textbook Knowledge Benefits Multimodal Reasoning 31 May 2024 · 0 repositories · arXiv:2405.20834
-
Ensemble Model With Bert,Roberta and Xlnet For Molecular property prediction 30 May 2024 · 0 repositories · arXiv:2406.06553
-
GNN-RAG: Graph Neural Retrieval for Large Language Model Reasoning 30 May 2024 · 1 repository · arXiv:2405.20139Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Is My Data in Your Retrieval Database? Membership Inference Attacks Against Retrieval Augmented Generation 30 May 2024 · 0 repositories · arXiv:2405.20446
-
KerasCV and KerasNLP: Vision and Language Power-Ups 30 May 2024 · 0 repositories · arXiv:2405.20247
-
One Token Can Help! Learning Scalable and Pluggable Virtual Tokens for Retrieval-Augmented Large Language Models 30 May 2024 · 2 repositories · arXiv:2405.19670
-
Phantom: General Trigger Attacks on Retrieval Augmented Language Generation 30 May 2024 · 0 repositories · arXiv:2405.20485
-
Student Answer Forecasting: Transformer-Driven Answer Choice Prediction for Language Learning 30 May 2024 · 1 repository · arXiv:2405.20079
-
A Multi-Source Retrieval Question Answering Framework Based on RAG 29 May 2024 · 0 repositories · arXiv:2405.19207
-
Can GPT Redefine Medical Understanding? Evaluating GPT on Biomedical Machine Reading Comprehension 29 May 2024 · 0 repositories · arXiv:2405.18682
-
CtrlA: Adaptive Retrieval-Augmented Generation via Inherent Control 29 May 2024 · 1 repository · arXiv:2405.18727Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 12 harvested samples)
-
STAT: Shrinking Transformers After Training 29 May 2024 · 0 repositories · arXiv:2406.00061
-
Toward Conversational Agents with Context and Time Sensitive Long-term Memory 29 May 2024 · 1 repository · arXiv:2406.00057
-
Two-Layer Retrieval-Augmented Generation Framework for Low-Resource Medical Question Answering Using Reddit Data: Proof-of-Concept Study 29 May 2024 · 0 repositories · arXiv:2405.19519
-
ATM: Adversarial Tuning Multi-agent System Makes a Robust Retrieval-Augmented Generator 28 May 2024 · 1 repository · arXiv:2405.18111Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Attention-based sequential recommendation system using multimodal data 28 May 2024 · 0 repositories · arXiv:2405.17959
-
Don't Forget to Connect! Improving RAG with Graph-based Reranking 28 May 2024 · 0 repositories · arXiv:2405.18414
-
WIDIn: Wording Image for Domain-Invariant Representation in Single-Source Domain Generalization 28 May 2024 · 0 repositories · arXiv:2405.18405
-
Augmenting Textual Generation via Topology Aware Retrieval 27 May 2024 · 0 repositories · arXiv:2405.17602
-
DeeperImpact: Optimizing Sparse Learned Index Structures 27 May 2024 · 1 repository · arXiv:2405.17093
-
Detecting Deceptive Dark Patterns in E-commerce Platforms 27 May 2024 · 0 repositories · arXiv:2406.01608
-
Exploiting the Layered Intrinsic Dimensionality of Deep Models for Practical Adversarial Training 27 May 2024 · 0 repositories · arXiv:2405.17130
-
NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models 27 May 2024 · 0 repositories · arXiv:2405.17428
-
PAE: LLM-based Product Attribute Extraction for E-Commerce Fashion Trends 27 May 2024 · 0 repositories · arXiv:2405.17533
-
Performance evaluation of Reddit Comments using Machine Learning and Natural Language Processing methods in Sentiment Analysis 27 May 2024 · 0 repositories · arXiv:2405.16810
-
QUB-Cirdan at "Discharge Me!": Zero shot discharge letter generation by open-source LLM 27 May 2024 · 0 repositories · arXiv:2406.00041
-
Video Enriched Retrieval Augmented Generation Using Aligned Video Captions 27 May 2024 · 1 repository · arXiv:2405.17706
-
AI-Generated Text Detection and Classification Based on BERT Deep Learning Algorithm 26 May 2024 · 0 repositories · arXiv:2405.16422
-
GRAG: Graph Retrieval-Augmented Generation 26 May 2024 · 1 repository · arXiv:2405.16506
-
M-RAG: Reinforcing Large Language Model Performance through Retrieval-Augmented Generation with Multiple Partitions 26 May 2024 · 0 repositories · arXiv:2405.16420
-
Accelerating Inference of Retrieval-Augmented Generation via Sparse Context Selection 25 May 2024 · 0 repositories · arXiv:2405.16178
-
Incremental Comprehension of Garden-Path Sentences by Large Language Models: Semantic Interpretation, Syntactic Re-Analysis, and Attention 25 May 2024 · 0 repositories · arXiv:2405.16042
-
Towards Unlocking Insights from Logbooks Using AI 25 May 2024 · 0 repositories · arXiv:2406.12881
-
Enhancing Augmentative and Alternative Communication with Card Prediction and Colourful Semantics 24 May 2024 · 0 repositories · arXiv:2405.15896
-
Comet: A Communication-efficient and Performant Approximation for Private Transformer Inference 24 May 2024 · 0 repositories · arXiv:2405.17485
-
A Structure-Aware Framework for Learning Device Placements on Computation Graphs 23 May 2024 · 1 repository · arXiv:2405.14185
-
CEEBERT: Cross-Domain Inference in Early Exit BERT 23 May 2024 · 1 repository · arXiv:2405.15039Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language Models 23 May 2024 · 2 repositories · arXiv:2405.14831Syntology community repositories only · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
RaFe: Ranking Feedback Improves Query Rewriting for RAG 23 May 2024 · 0 repositories · arXiv:2405.14431
-
ViHateT5: Enhancing Hate Speech Detection in Vietnamese With A Unified Text-to-Text Transformer Model 23 May 2024 · 1 repository · arXiv:2405.14141
-
Automated Evaluation of Retrieval-Augmented Language Models with Task-Specific Exam Generation 22 May 2024 · 1 repository · arXiv:2405.13622
-
FlashRAG: A Modular Toolkit for Efficient Retrieval-Augmented Generation Research 22 May 2024 · 1 repository · arXiv:2405.13576
-
TrojanRAG: Retrieval-Augmented Generation Can Be Backdoor Driver in Large Language Models 22 May 2024 · 1 repository · arXiv:2405.13401
-
Unleashing the Power of Unlabeled Data: A Self-supervised Learning Framework for Cyber Attack Detection in Smart Grids 22 May 2024 · 0 repositories · arXiv:2405.13965
-
Generative AI in Cybersecurity: A Comprehensive Review of LLM Applications and Vulnerabilities 21 May 2024 · 0 repositories · arXiv:2405.12750
-
How Reliable AI Chatbots are for Disease Prediction from Patient Complaints? 21 May 2024 · 0 repositories · arXiv:2405.13219
-
The 2nd FutureDial Challenge: Dialog Systems with Retrieval Augmented Generation (FutureDial-RAG) 21 May 2024 · 1 repository · arXiv:2405.13084
-
A review on the use of large language models as virtual tutors 20 May 2024 · 0 repositories · arXiv:2405.11983
-
CReMa: Crisis Response through Computational Identification and Matching of Cross-Lingual Requests and Offers Shared on Social Media 20 May 2024 · 0 repositories · arXiv:2405.11897
-
Degree of Irrationality: Sentiment and Implied Volatility Surface 20 May 2024 · 0 repositories · arXiv:2405.11730
-
Question-Based Retrieval using Atomic Units for Enterprise RAG 20 May 2024 · 0 repositories · arXiv:2405.12363
-
Exploring speech style spaces with language models: Emotional TTS without emotion labels 18 May 2024 · 0 repositories · arXiv:2405.11413
-
ActiveLLM: Large Language Model-based Active Learning for Textual Few-Shot Scenarios 17 May 2024 · 0 repositories · arXiv:2405.10808
-
Empowering Prior to Court Legal Analysis: A Transparent and Accessible Dataset for Defensive Statement Classification and Interpretation 17 May 2024 · 0 repositories · arXiv:2405.10702
-
NeuroAssist: Enhancing Cognitive-Computer Synergy with Adaptive AI and Advanced Neural Decoding for Efficient EEG Signal Classification 17 May 2024 · 0 repositories · arXiv:2406.01600
-
Tailoring Vaccine Messaging with Common-Ground Opinions 17 May 2024 · 1 repository · arXiv:2405.10861
-
Bridging the gap in online hate speech detection: a comparative analysis of BERT and traditional models for homophobic content identification on X/Twitter 15 May 2024 · 0 repositories · arXiv:2405.09221
-
IM-RAG: Multi-Round Retrieval-Augmented Generation Through Learning Inner Monologues 15 May 2024 · 0 repositories · arXiv:2405.13021
-
Transfer Learning in Pre-Trained Large Language Models for Malware Detection Based on System Calls 15 May 2024 · 0 repositories · arXiv:2405.09318
-
Control Token with Dense Passage Retrieval 13 May 2024 · 0 repositories · arXiv:2405.13008
-
Evaluation of Retrieval-Augmented Generation: A Survey 13 May 2024 · 1 repository · arXiv:2405.07437
-
From Questions to Insightful Answers: Building an Informed Chatbot for University Resources 13 May 2024 · 0 repositories · arXiv:2405.08120
-
DuetRAG: Collaborative Retrieval-Augmented Generation 12 May 2024 · 0 repositories · arXiv:2405.13002
-
ExplainableDetector: Exploring Transformer-based Language Modeling Approach for SMS Spam Detection with Explainability Analysis 12 May 2024 · 0 repositories · arXiv:2405.08026
-
L(u)PIN: LLM-based Political Ideology Nowcasting 12 May 2024 · 0 repositories · arXiv:2405.07320
-
TacoERE: Cluster-aware Compression for Event Relation Extraction 11 May 2024 · 0 repositories · arXiv:2405.06890
-
A Survey on RAG Meeting LLMs: Towards Retrieval-Augmented Large Language Models 10 May 2024 · 0 repositories · arXiv:2405.06211
-
CANAL -- Cyber Activity News Alerting Language Model: Empirical Approach vs. Expensive LLM 10 May 2024 · 0 repositories · arXiv:2405.06772
-
Characterizing the Accuracy -- Efficiency Trade-off of Low-rank Decomposition in Language Models 10 May 2024 · 0 repositories · arXiv:2405.06626
-
Ditto: Quantization-aware Secure Inference of Transformers upon MPC 9 May 2024 · 1 repository · arXiv:2405.05525
-
Reddit-Impacts: A Named Entity Recognition Dataset for Analyzing Clinical and Social Effects of Substance Use Derived from Social Media 9 May 2024 · 0 repositories · arXiv:2405.06145
-
Evaluating Students' Open-ended Written Responses with LLMs: Using the RAG Framework for GPT-3.5, GPT-4, Claude-3, and Mistral-Large 8 May 2024 · 0 repositories · arXiv:2405.05444
-
Utilizing Large Language Models to Generate Synthetic Data to Increase the Performance of BERT-Based Neural Networks 8 May 2024 · 0 repositories · arXiv:2405.06695
-
Enriched BERT Embeddings for Scholarly Publication Classification 7 May 2024 · 1 repository · arXiv:2405.04136
-
ERATTA: Extreme RAG for Table To Answers with Large Language Models 7 May 2024 · 0 repositories · arXiv:2405.03963
-
Evaluating Text Summaries Generated by Large Language Models Using OpenAI's GPT 7 May 2024 · 0 repositories · arXiv:2405.04053
-
Remote Diffusion 7 May 2024 · 0 repositories · arXiv:2405.04717
-
Revisiting Character-level Adversarial Attacks for Language Models 7 May 2024 · 1 repository · arXiv:2405.04346Syntology official (archive's flag): 23 ran · 23 ran (of which 3 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 10 where Syntology's instrument failed) · 8 unverified (of 31 harvested samples)
-
Robust Implementation of Retrieval-Augmented Generation on Edge-based Computing-in-Memory Architectures 7 May 2024 · 0 repositories · arXiv:2405.04700
-
Utilizing GPT to Enhance Text Summarization: A Strategy to Minimize Hallucinations 7 May 2024 · 0 repositories · arXiv:2405.04039
-
Characterizing the Dilemma of Performance and Index Size in Billion-Scale Vector Search and Breaking It with Second-Tier Memory 6 May 2024 · 0 repositories · arXiv:2405.03267
-
Compressing Long Context for Enhancing RAG with AMR-based Concept Distillation 6 May 2024 · 0 repositories · arXiv:2405.03085
-
Detecting Android Malware: From Neural Embeddings to Hands-On Validation with BERTroid 6 May 2024 · 0 repositories · arXiv:2405.03620
-
Detecting Anti-Semitic Hate Speech using Transformer-based Large Language Models 6 May 2024 · 0 repositories · arXiv:2405.03794
-
ERAGent: Enhancing Retrieval-Augmented Language Models with Improved Accuracy, Efficiency, and Personalization 6 May 2024 · 1 repository · arXiv:2405.06683
-
Leveraging Lecture Content for Improved Feedback: Explorations with GPT-4 and Retrieval Augmented Generation 5 May 2024 · 0 repositories · arXiv:2405.06681
-
Stochastic RAG: End-to-End Retrieval-Augmented Generation through Expected Utility Maximization 5 May 2024 · 0 repositories · arXiv:2405.02816
-
Unraveling the Dominance of Large Language Models Over Transformer Models for Bangla Natural Language Inference: A Comprehensive Study 5 May 2024 · 1 repository · arXiv:2405.02937
-
A Combination of BERT and Transformer for Vietnamese Spelling Correction 4 May 2024 · 0 repositories · arXiv:2405.02573
-
Analyzing Narrative Processing in Large Language Models (LLMs): Using GPT4 to test BERT 3 May 2024 · 0 repositories · arXiv:2405.02024
-
Comparative Analysis of Retrieval Systems in the Real World 3 May 2024 · 0 repositories · arXiv:2405.02048
-
DALLMi: Domain Adaption for LLM-based Multi-label Classifier 3 May 2024 · 1 repository · arXiv:2405.01883