Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 5
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 5 of 190: papers 401 to 500 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Multi-modal brain encoding models for multi-modal stimuli 26 May 2025 · 1 repository · arXiv:2505.20027Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 3 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning 26 May 2025 · 0 repositories · arXiv:2505.19938
-
NeuSym-RAG: Hybrid Neural Symbolic Retrieval with Multiview Structuring for PDF Question Answering 26 May 2025 · 1 repository · arXiv:2505.19754
-
One Surrogate to Fool Them All: Universal, Transferable, and Targeted Adversarial Attacks with CLIP 26 May 2025 · 1 repository · arXiv:2505.19840
-
R3-RAG: Learning Step-by-Step Reasoning and Retrieval for LLMs via Reinforcement Learning 26 May 2025 · 1 repository · arXiv:2505.23794
-
REARANK: Reasoning Re-ranking Agent via Reinforcement Learning 26 May 2025 · 1 repository · arXiv:2505.20046Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 2 where Syntology's instrument failed) · 6 unverified (of 18 harvested samples) · 4 pointer-only (licence)
-
Structured Initialization for Vision Transformers 26 May 2025 · 0 repositories · arXiv:2505.19985
-
syftr: Pareto-Optimal Generative AI 26 May 2025 · 1 repository · arXiv:2505.20266
-
Synthetic Time Series Forecasting with Transformer Architectures: Extensive Simulation Benchmarks 26 May 2025 · 1 repository · arXiv:2505.20048
-
The Avengers: A Simple Recipe for Uniting Smaller Language Models to Challenge Proprietary Giants 26 May 2025 · 1 repository · arXiv:2505.19797
-
The Missing Point in Vision Transformers for Universal Image Segmentation 26 May 2025 · 1 repository · arXiv:2505.19795
-
Training LLM-Based Agents with Synthetic Self-Reflected Trajectories and Partial Masking 26 May 2025 · 0 repositories · arXiv:2505.20023
-
Transformers in Protein: A Survey 26 May 2025 · 0 repositories · arXiv:2505.20098
-
Understanding Transformer from the Perspective of Associative Memory 26 May 2025 · 0 repositories · arXiv:2505.19488
-
VADER: A Human-Evaluated Benchmark for Vulnerability Assessment, Detection, Explanation, and Remediation 26 May 2025 · 1 repository · arXiv:2505.19395
-
A Smart Healthcare System for Monkeypox Skin Lesion Detection and Tracking 25 May 2025 · 0 repositories · arXiv:2505.19023
-
AI4Math: A Native Spanish Benchmark for University-Level Mathematical Reasoning in Large Language Models 25 May 2025 · 0 repositories · arXiv:2505.18978
-
Assistant-Guided Mitigation of Teacher Preference Bias in LLM-as-a-Judge 25 May 2025 · 1 repository · arXiv:2505.19176
-
Benchmarking Large Language Models for Cyberbullying Detection in Real-World YouTube Comments 25 May 2025 · 0 repositories · arXiv:2505.18927
-
Communication-Efficient Multi-Device Inference Acceleration for Transformer Models 25 May 2025 · 1 repository · arXiv:2505.19342
-
Exploring Magnitude Preservation and Rotation Modulation in Diffusion Transformers 25 May 2025 · 0 repositories · arXiv:2505.19122
-
GhostPrompt: Jailbreaking Text-to-image Generative Models based on Dynamic Optimization 25 May 2025 · 0 repositories · arXiv:2505.18979
-
Hypercube-RAG: Hypercube-Based Retrieval-Augmented Generation for In-domain Scientific Question-Answering 25 May 2025 · 1 repository · arXiv:2505.19288
-
Investigating Pedagogical Teacher and Student LLM Agents: Genetic Adaptation Meets Retrieval Augmented Generation Across Learning Style 25 May 2025 · 0 repositories · arXiv:2505.19173
-
POQD: Performance-Oriented Query Decomposer for Multi-vector retrieval 25 May 2025 · 1 repository · arXiv:2505.19189Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 3 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Retrieval-Augmented Generation for Service Discovery: Chunking Strategies and Benchmarking 25 May 2025 · 0 repositories · arXiv:2505.19310
-
System-1.5 Reasoning: Traversal in Language and Latent Spaces with Dynamic Shortcuts 25 May 2025 · 0 repositories · arXiv:2505.18962
-
A Survey of LLM × DATA 24 May 2025 · 2 repositories · arXiv:2505.18458
-
Benchmarking Poisoning Attacks against Retrieval-Augmented Generation 24 May 2025 · 0 repositories · arXiv:2505.18543
-
BRIT: Bidirectional Retrieval over Unified Image-Text Graph 24 May 2025 · 0 repositories · arXiv:2505.18450
-
Federated Retrieval-Augmented Generation: A Systematic Mapping Study 24 May 2025 · 0 repositories · arXiv:2505.18906
-
From Reddit to Generative AI: Evaluating Large Language Models for Anxiety Support Fine-tuned on Social Media Data 24 May 2025 · 0 repositories · arXiv:2505.18464
-
GainRAG: Preference Alignment in Retrieval-Augmented Generation through Gain Signal Synthesis 24 May 2025 · 1 repository · arXiv:2505.18710
-
How Does Sequence Modeling Architecture Influence Base Capabilities of Pre-trained Language Models? Exploring Key Architecture Design Principles to Avoid Base Capabilities Degradation 24 May 2025 · 0 repositories · arXiv:2505.18522
-
LLMs for Supply Chain Management 24 May 2025 · 0 repositories · arXiv:2505.18597
-
Localizing Knowledge in Diffusion Transformers 24 May 2025 · 0 repositories · arXiv:2505.18832
-
Removal of Hallucination on Hallucination: Debate-Augmented RAG 24 May 2025 · 1 repository · arXiv:2505.18581Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 4 harvested samples)
-
Security Concerns for Large Language Models: A Survey 24 May 2025 · 0 repositories · arXiv:2505.18889
-
Smart Energy Guardian: A Hybrid Deep Learning Model for Detecting Fraudulent PV Generation 24 May 2025 · 0 repositories · arXiv:2505.18755
-
Strong Membership Inference Attacks on Massive Datasets and (Moderately) Large Language Models 24 May 2025 · 0 repositories · arXiv:2505.18773
-
SW-ViT: A Spatio-Temporal Vision Transformer Network with Post Denoiser for Sequential Multi-Push Ultrasound Shear Wave Elastography 24 May 2025 · 0 repositories · arXiv:2505.18865
-
The Silent Saboteur: Imperceptible Adversarial Attacks against Black-Box Retrieval-Augmented Generation Systems 24 May 2025 · 0 repositories · arXiv:2505.18583
-
TrajMoE: Spatially-Aware Mixture of Experts for Unified Human Mobility Modeling 24 May 2025 · 0 repositories · arXiv:2505.18670
-
Unleashing Diffusion Transformers for Visual Correspondence by Modulating Massive Activations 24 May 2025 · 0 repositories · arXiv:2505.18584
-
COLORA: Efficient Fine-Tuning for Convolutional Models with a Study Case on Optical Coherence Tomography Image Classification 23 May 2025 · 0 repositories · arXiv:2505.18315
-
Contrastive Distillation of Emotion Knowledge from LLMs for Zero-Shot Emotion Recognition 23 May 2025 · 1 repository · arXiv:2505.18040
-
Direct3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse Attention 23 May 2025 · 1 repository · arXiv:2505.17412Syntology 6 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples)
-
Explainable Anatomy-Guided AI for Prostate MRI: Foundation Models and In Silico Clinical Trials for Virtual Biopsy-based Risk Assessment 23 May 2025 · 0 repositories · arXiv:2505.17971
-
FinRAGBench-V: A Benchmark for Multimodal RAG with Visual Citation in the Financial Domain 23 May 2025 · 0 repositories · arXiv:2505.17471
-
FreqU-FNet: Frequency-Aware U-Net for Imbalanced Medical Image Segmentation 23 May 2025 · 0 repositories · arXiv:2505.17544
-
Gaming Tool Preferences in Agentic LLMs 23 May 2025 · 1 repository · arXiv:2505.18135
-
Hybrid Mamba-Transformer Decoder for Error-Correcting Codes 23 May 2025 · 0 repositories · arXiv:2505.17834
-
Is It Bad to Work All the Time? Cross-Cultural Evaluation of Social Norm Biases in GPT-4 23 May 2025 · 0 repositories · arXiv:2505.18322
-
LLM assisted web application functional requirements generation: A case study of four popular LLMs over a Mess Management System 23 May 2025 · 0 repositories · arXiv:2505.18019
-
MMMG: a Comprehensive and Reliable Evaluation Suite for Multitask Multimodal Generation 23 May 2025 · 0 repositories · arXiv:2505.17613
-
Multi-Scale Probabilistic Generation Theory: A Hierarchical Framework for Interpreting Large Language Models 23 May 2025 · 0 repositories · arXiv:2505.18244
-
One Model Transfer to All: On Robust Jailbreak Prompts Generation against LLMs 23 May 2025 · 1 repository · arXiv:2505.17598
-
QwenLong-CPRS: Towards ∞-LLMs with Dynamic Context Optimization 23 May 2025 · 0 repositories · arXiv:2505.18092
-
Resolving Conflicting Evidence in Automated Fact-Checking: A Study on Retrieval-Augmented LLMs 23 May 2025 · 1 repository · arXiv:2505.17762
-
ShIOEnv: A CLI Behavior-Capturing Environment Enabling Grammar-Guided Command Synthesis for Dataset Curation 23 May 2025 · 1 repository · arXiv:2505.18374
-
Token Reduction Should Go Beyond Efficiency in Generative Models -- From Vision, Language to Multimodality 23 May 2025 · 1 repository · arXiv:2505.18227
-
AdamS: Momentum Itself Can Be A Normalizer for LLM Pretraining and Post-training 22 May 2025 · 1 repository · arXiv:2505.16363Syntology official: no sample here; runs from other or unrecorded repositories · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Align-GRAG: Reasoning-Guided Dual Alignment for Graph Retrieval-Augmented Generation 22 May 2025 · 0 repositories · arXiv:2505.16237
-
Attributing Response to Context: A Jensen-Shannon Divergence Driven Mechanistic Study of Context Attribution in Retrieval-Augmented Generation 22 May 2025 · 0 repositories · arXiv:2505.16415Syntology 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Augmenting LLM Reasoning with Dynamic Notes Writing for Complex QA 22 May 2025 · 0 repositories · arXiv:2505.16293
-
Beamforming-Codebook-Aware Channel Knowledge Map Construction for Multi-Antenna Systems 22 May 2025 · 1 repository · arXiv:2505.16132
-
Bottlenecked Transformers: Periodic KV Cache Abstraction for Generalised Reasoning 22 May 2025 · 0 repositories · arXiv:2505.16950
-
Breaking Complexity Barriers: High-Resolution Image Restoration with Rank Enhanced Linear Attention 22 May 2025 · 0 repositories · arXiv:2505.16157
-
CAIFormer: A Causal Informed Transformer for Multivariate Time Series Forecasting 22 May 2025 · 0 repositories · arXiv:2505.16308
-
Chain-of-Thought Poisoning Attacks against R1-based Retrieval-Augmented Generation Systems 22 May 2025 · 0 repositories · arXiv:2505.16367
-
Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering Tasks 22 May 2025 · 0 repositories · arXiv:2505.16901
-
Comparative analysis of subword tokenization approaches for Indian languages 22 May 2025 · 0 repositories · arXiv:2505.16868
-
DailyQA: A Benchmark to Evaluate Web Retrieval Augmented LLMs Based on Capturing Real-World Changes 22 May 2025 · 0 repositories · arXiv:2505.17162
-
Data-Driven Breakthroughs and Future Directions in AI Infrastructure: A Comprehensive Review 22 May 2025 · 0 repositories · arXiv:2505.16771
-
Date Fragments: A Hidden Bottleneck of Tokenization for Temporal Reasoning 22 May 2025 · 0 repositories · arXiv:2505.16088
-
Explain Less, Understand More: Jargon Detection via Personalized Parameter-Efficient Fine-tuning 22 May 2025 · 0 repositories · arXiv:2505.16227
-
Fusion of Foundation and Vision Transformer Model Features for Dermatoscopic Image Classification 22 May 2025 · 0 repositories · arXiv:2505.16338
-
Generative AI and Creativity: A Systematic Literature Review and Meta-Analysis 22 May 2025 · 1 repository · arXiv:2505.17241
-
Learning Normal Patterns in Musical Loops 22 May 2025 · 0 repositories · arXiv:2505.23784
-
Native Segmentation Vision Transformers 22 May 2025 · 0 repositories · arXiv:2505.16993
-
Personalizing Student-Agent Interactions Using Log-Contextualized Retrieval Augmented Generation (RAG) 22 May 2025 · 0 repositories · arXiv:2505.17238
-
R1-Searcher++: Incentivizing the Dynamic Knowledge Acquisition of LLMs via Reinforcement Learning 22 May 2025 · 3 repositories · arXiv:2505.17005Syntology community repositories only · 9 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 10 harvested samples)
-
Scalable Graph Generative Modeling via Substructure Sequences 22 May 2025 · 1 repository · arXiv:2505.16130Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Search Wisely: Mitigating Sub-optimal Agentic Searches By Reducing Uncertainty 22 May 2025 · 0 repositories · arXiv:2505.17281
-
Swin Transformer for Robust CGI Images Detection: Intra- and Inter-Dataset Analysis across Multiple Color Spaces 22 May 2025 · 0 repositories · arXiv:2505.16253
-
The Polar Express: Optimal Matrix Sign Methods and Their Application to the Muon Algorithm 22 May 2025 · 0 repositories · arXiv:2505.16932
-
Three Minds, One Legend: Jailbreak Large Reasoning Model with Adaptive Stacked Ciphers 22 May 2025 · 0 repositories · arXiv:2505.16241
-
Training-Free Efficient Video Generation via Dynamic Token Carving 22 May 2025 · 1 repository · arXiv:2505.16864
-
Transformer Copilot: Learning from The Mistake Log in LLM Fine-tuning 22 May 2025 · 1 repository · arXiv:2505.16270Syntology official (archive's flag): 4 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 3 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
Understanding Differential Transformer Unchains Pretrained Self-Attentions 22 May 2025 · 0 repositories · arXiv:2505.16333
-
VoxRAG: A Step Toward Transcription-Free RAG Systems in Spoken Question Answering 22 May 2025 · 0 repositories · arXiv:2505.17326
-
Walk&Retrieve: Simple Yet Effective Zero-shot Retrieval-Augmented Generation via Knowledge Graph Walks 22 May 2025 · 1 repository · arXiv:2505.16849
-
An Efficient Private GPT Never Autoregressively Decodes 21 May 2025 · 0 repositories · arXiv:2505.15252
-
An Exploratory Approach Towards Investigating and Explaining Vision Transformer and Transfer Learning for Brain Disease Detection 21 May 2025 · 0 repositories · arXiv:2505.16039
-
BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems 21 May 2025 · 0 repositories · arXiv:2505.15216
-
BR-TaxQA-R: A Dataset for Question Answering with References for Brazilian Personal Income Tax Law, including case law 21 May 2025 · 0 repositories · arXiv:2505.15916
-
HDLxGraph: Bridging Large Language Models and HDL Repositories via HDL Graph Databases 21 May 2025 · 1 repository · arXiv:2505.15701
-
InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation 21 May 2025 · 0 repositories · arXiv:2505.15872
-
Leveraging Large Language Models for Command Injection Vulnerability Analysis in Python: An Empirical Study on Popular Open-Source Projects 21 May 2025 · 0 repositories · arXiv:2505.15088
-
LogiCase: Effective Test Case Generation from Logical Description in Competitive Programming 21 May 2025 · 0 repositories · arXiv:2505.15039Syntology 11 ran (of which 3 constructed an object rather than computing a result; 5 with no instrument failure: 2 honoured, 0 violated, 3 with no contract checked; 6 where Syntology's instrument failed) · 8 unverified (of 19 harvested samples) · 19 pointer-only (licence)