Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 14
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 14 of 190: papers 1,301 to 1,400 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Vision Transformer Based Semantic Communications for Next Generation Wireless Networks 21 Mar 2025 · 0 repositories · arXiv:2503.17275
-
When Words Outperform Vision: VLMs Can Self-Improve Via Text-Only Training For Human-Centered Decision Making 21 Mar 2025 · 0 repositories · arXiv:2503.16965Syntology 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Zero-Shot Styled Text Image Generation, but Make It Autoregressive 21 Mar 2025 · 0 repositories · arXiv:2503.17074
-
Binarized Mamba-Transformer for Lightweight Quad Bayer HybridEVS Demosaicing 20 Mar 2025 · 1 repository · arXiv:2503.16134
-
Design and Implementation of an FPGA-Based Hardware Accelerator for Transformer 20 Mar 2025 · 1 repository · arXiv:2503.16731
-
DocVideoQA: Towards Comprehensive Understanding of Document-Centric Videos through Question Answering 20 Mar 2025 · 0 repositories · arXiv:2503.15887
-
EDiT: Efficient Diffusion Transformers with Linear Compressed Attention 20 Mar 2025 · 0 repositories · arXiv:2503.16726
-
Financial Analysis: Intelligent Financial Data Analysis System Based on LLM-RAG 20 Mar 2025 · 0 repositories · arXiv:2504.06279
-
FreeFlux: Understanding and Exploiting Layer-Specific Roles in RoPE-Based MMDiT for Versatile Image Editing 20 Mar 2025 · 0 repositories · arXiv:2503.16153
-
GraPLUS: Graph-based Placement Using Semantics for Image Composition 20 Mar 2025 · 0 repositories · arXiv:2503.15761
-
Hyperspectral Imaging for Identifying Foreign Objects on Pork Belly 20 Mar 2025 · 0 repositories · arXiv:2503.16086
-
iFlame: Interleaving Full and Linear Attention for Efficient Mesh Generation 20 Mar 2025 · 0 repositories · arXiv:2503.16653
-
Iterative Optimal Attention and Local Model for Single Image Rain Streak Removal 20 Mar 2025 · 1 repository · arXiv:2503.16165
-
Deep learning framework for action prediction reveals multi-timescale locomotor control 20 Mar 2025 · 0 repositories · arXiv:2503.16340
-
Parameters vs. Context: Fine-Grained Control of Knowledge Reliance in Language Models 20 Mar 2025 · 1 repository · arXiv:2503.15888
-
PromptHash: Affinity-Prompted Collaborative Cross-Modal Learning for Adaptive Hashing Retrieval 20 Mar 2025 · 0 repositories · arXiv:2503.16064
-
SenseExpo: Efficient Autonomous Exploration with Prediction Information from Lightweight Neural Networks 20 Mar 2025 · 0 repositories · arXiv:2503.16000
-
SpiLiFormer: Enhancing Spiking Transformers with Lateral Inhibition 20 Mar 2025 · 0 repositories · arXiv:2503.15986
-
The Lighthouse of Language: Enhancing LLM Agents via Critique-Guided Improvement 20 Mar 2025 · 0 repositories · arXiv:2503.16024
-
Towards Lighter and Robust Evaluation for Retrieval Augmented Generation 20 Mar 2025 · 1 repository · arXiv:2503.16161Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Transformer-based Wireless Symbol Detection Over Fading Channels 20 Mar 2025 · 0 repositories · arXiv:2503.16594
-
Tuning LLMs by RAG Principles: Towards LLM-native Memory 20 Mar 2025 · 1 repository · arXiv:2503.16071
-
Typed-RAG: Type-aware Multi-Aspect Decomposition for Non-Factoid Question Answering 20 Mar 2025 · 1 repository · arXiv:2503.15879
-
Unify and Triumph: Polyglot, Diverse, and Self-Consistent Generation of Unit Tests with LLMs 20 Mar 2025 · 0 repositories · arXiv:2503.16144
-
UniHDSA: A Unified Relation Prediction Approach for Hierarchical Document Structure Analysis 20 Mar 2025 · 1 repository · arXiv:2503.15893
-
XAttention: Block Sparse Attention with Antidiagonal Scoring 20 Mar 2025 · 1 repository · arXiv:2503.16428
-
A Novel Channel Boosted Residual CNN-Transformer with Regional-Boundary Learning for Breast Cancer Detection 19 Mar 2025 · 0 repositories · arXiv:2503.15008
-
ChatGPT or A Silent Everywhere Helper: A Survey of Large Language Models 19 Mar 2025 · 0 repositories · arXiv:2503.17403
-
ELTEX: A Framework for Domain-Driven Synthetic Data Generation 19 Mar 2025 · 1 repository · arXiv:2503.15055
-
Enhancing Code LLM Training with Programmer Attention 19 Mar 2025 · 0 repositories · arXiv:2503.14936
-
Enhancing Pancreatic Cancer Staging with Large Language Models: The Role of Retrieval-Augmented Generation 19 Mar 2025 · 0 repositories · arXiv:2503.15664
-
Bias Evaluation and Mitigation in Retrieval-Augmented Medical Question-Answering Systems 19 Mar 2025 · 0 repositories · arXiv:2503.15454
-
FP4DiT: Towards Effective Floating Point Quantization for Diffusion Transformers 19 Mar 2025 · 1 repository · arXiv:2503.15465Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
GenM³: Generative Pretrained Multi-path Motion Model for Text Conditional Human Motion Generation 19 Mar 2025 · 0 repositories · arXiv:2503.14919
-
Optimizing Retrieval Strategies for Financial Question Answering Documents in Retrieval-Augmented Generation Systems 19 Mar 2025 · 1 repository · arXiv:2503.15191Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
RAG-based User Profiling for Precision Planning in Mixed-precision Over-the-Air Federated Learning 19 Mar 2025 · 0 repositories · arXiv:2503.15569
-
TROVE: A Challenge for Fine-Grained Text Provenance via Source Sentence Tracing and Relationship Classification 19 Mar 2025 · 0 repositories · arXiv:2503.15289
-
TruthLens:A Training-Free Paradigm for DeepFake Detection 19 Mar 2025 · 0 repositories · arXiv:2503.15342
-
Understanding the Generalization of In-Context Learning in Transformers: An Empirical Study 19 Mar 2025 · 1 repository · arXiv:2503.15579
-
Binary AddiVortes: (Bayesian) Additive Voronoi Tessellations for Binary Classification with an application to Predicting Home Mortgage Application Outcomes 18 Mar 2025 · 0 repositories · arXiv:2503.21792
-
BurTorch: Revisiting Training from First Principles by Coupling Autodiff, Math Optimization, and Systems 18 Mar 2025 · 1 repository · arXiv:2503.13795
-
CTSAC: Curriculum-Based Transformer Soft Actor-Critic for Goal-Oriented Robot Exploration 18 Mar 2025 · 0 repositories · arXiv:2503.14254
-
Dynamic Accumulated Attention Map for Interpreting Evolution of Decision-Making in Vision Transformer 18 Mar 2025 · 1 repository · arXiv:2503.14640
-
Enhancing LLM Generation with Knowledge Hypergraph for Evidence-Based Medicine 18 Mar 2025 · 0 repositories · arXiv:2503.16530
-
Fast Autoregressive Video Generation with Diagonal Decoding 18 Mar 2025 · 0 repositories · arXiv:2503.14070
-
Good/Evil Reputation Judgment of Celebrities by LLMs via Retrieval Augmented Generation 18 Mar 2025 · 0 repositories · arXiv:2503.14382
-
Gricean Norms as a Basis for Effective Collaboration 18 Mar 2025 · 1 repository · arXiv:2503.14484
-
JuDGE: Benchmarking Judgment Document Generation for Chinese Legal System 18 Mar 2025 · 1 repository · arXiv:2503.14258
-
Beyond Single Pass, Looping Through Time: KG-IRAG with Iterative Knowledge Retrieval 18 Mar 2025 · 0 repositories · arXiv:2503.14234
-
Large Language Models for Virtual Human Gesture Selection 18 Mar 2025 · 0 repositories · arXiv:2503.14408
-
MDocAgent: A Multi-Modal Multi-Agent Framework for Document Understanding 18 Mar 2025 · 1 repository · arXiv:2503.13964Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
MoK-RAG: Mixture of Knowledge Paths Enhanced Retrieval-Augmented Generation for Embodied AI Environments 18 Mar 2025 · 0 repositories · arXiv:2503.13882
-
Multimodal Feature-Driven Deep Learning for the Prediction of Duck Body Dimensions and Weight 18 Mar 2025 · 0 repositories · arXiv:2503.14001
-
PENCIL: Long Thoughts with Short Memory 18 Mar 2025 · 1 repository · arXiv:2503.14337Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
RAGO: Systematic Performance Optimization for Retrieval-Augmented Generation Serving 18 Mar 2025 · 1 repository · arXiv:2503.14649
-
Self-Vocabularizing Training for Neural Machine Translation 18 Mar 2025 · 0 repositories · arXiv:2503.13837
-
Splintering Nonconcatenative Languages for Better Tokenization 18 Mar 2025 · 1 repository · arXiv:2503.14433Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Theoretical Foundation of Flow-Based Time Series Generation: Provable Approximation, Generalization, and Efficiency 18 Mar 2025 · 0 repositories · arXiv:2503.14076
-
XOXO: Stealthy Cross-Origin Context Poisoning Attacks against AI Coding Assistants 18 Mar 2025 · 0 repositories · arXiv:2503.14281
-
A Reinforcement Learning-Driven Transformer GAN for Molecular Generation 17 Mar 2025 · 0 repositories · arXiv:2503.12796
-
A Survey on Transformer Context Extension: Approaches and Evaluation 17 Mar 2025 · 0 repositories · arXiv:2503.13299
-
Advancing Chronic Tuberculosis Diagnostics Using Vision-Language Models: A Multi modal Framework for Precision Analysis 17 Mar 2025 · 0 repositories · arXiv:2503.14536
-
An interpretable approach to automating the assessment of biofouling in video footage 17 Mar 2025 · 1 repository · arXiv:2503.12875
-
Are LLMs (Really) Ideological? An IRT-based Analysis and Alignment Tool for Perceived Socio-Economic Bias in LLMs 17 Mar 2025 · 0 repositories · arXiv:2503.13149
-
Can Language Models Follow Multiple Turns of Entangled Instructions? 17 Mar 2025 · 1 repository · arXiv:2503.13222
-
Feature Extraction and Analysis for GPT-Generated Text 17 Mar 2025 · 0 repositories · arXiv:2503.13687
-
Generative AI for Software Architecture. Applications, Trends, Challenges, and Future Directions 17 Mar 2025 · 0 repositories · arXiv:2503.13310
-
Humanoid Policy ~ Human Policy 17 Mar 2025 · 0 repositories · arXiv:2503.13441
-
MES-RAG: Bringing Multi-modal, Entity-Storage, and Secure Enhancements to RAG 17 Mar 2025 · 1 repository · arXiv:2503.13563
-
OSCAR: Online Soft Compression And Reranking 17 Mar 2025 · 0 repositories · arXiv:2504.07109
-
Privacy-Aware RAG: Secure and Isolated Knowledge Retrieval 17 Mar 2025 · 0 repositories · arXiv:2503.15548
-
SeisRDT: Latent Diffusion Model Based On Representation Learning For Seismic Data Interpolation And Reconstruction 17 Mar 2025 · 0 repositories · arXiv:2503.21791
-
SuperBPE: Space Travel for Language Models 17 Mar 2025 · 0 repositories · arXiv:2503.13423
-
Towards Scalable Foundation Model for Multi-modal and Hyperspectral Geospatial Data 17 Mar 2025 · 0 repositories · arXiv:2503.12843
-
Fourier-Based 3D Multistage Transformer for Aberration Correction in Multicellular Specimens 16 Mar 2025 · 2 repositories · arXiv:2503.12593
-
Fragile Mastery: Are Domain-Specific Trade-Offs Undermining On-Device Language Models? 16 Mar 2025 · 0 repositories · arXiv:2503.22698
-
PA-CFL: Privacy-Adaptive Clustered Federated Learning for Transformer-Based Sales Forecasting on Heterogeneous Retail Data 15 Mar 2025 · 0 repositories · arXiv:2503.12220
-
Fast Critical Clearing Time Calculation for Power Systems with Synchronous and Asynchronous Generation 15 Mar 2025 · 0 repositories · arXiv:2503.12132
-
Integrating Chain-of-Thought and Retrieval Augmented Generation Enhances Rare Disease Diagnosis from Clinical Notes 15 Mar 2025 · 0 repositories · arXiv:2503.12286
-
LLM & HPC:Benchmarking DeepSeek's Performance in High-Performance Computing Tasks 15 Mar 2025 · 1 repository · arXiv:2504.03665
-
Maritime Mission Planning for Unmanned Surface Vessel using Large Language Model 15 Mar 2025 · 0 repositories · arXiv:2503.12065
-
Addressing Information Loss and Interaction Collapse: A Dual Enhanced Attention Framework for Feature Interaction 14 Mar 2025 · 0 repositories · arXiv:2503.11233
-
Alzheimer's Disease Classification Using Retinal OCT: TransnetOCT and Swin Transformer Models 14 Mar 2025 · 0 repositories · arXiv:2503.11511
-
Asynchronous Sharpness-Aware Minimization For Fast and Accurate Deep Learning 14 Mar 2025 · 0 repositories · arXiv:2503.11147
-
Augmenting Image Annotation: A Human-LMM Collaborative Framework for Efficient Object Selection and Label Generation 14 Mar 2025 · 0 repositories · arXiv:2503.11096
-
BEVDiffLoc: End-to-End LiDAR Global Localization in BEV View based on Diffusion Model 14 Mar 2025 · 1 repository · arXiv:2503.11372
-
Combining Causal Models for More Accurate Abstractions of Neural Networks 14 Mar 2025 · 1 repository · arXiv:2503.11429Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Context-Aware Rule Mining Using a Dynamic Transformer-Based Framework 14 Mar 2025 · 0 repositories · arXiv:2503.11125
-
DynRsl-VLM: Enhancing Autonomous Driving Perception with Dynamic Resolution Vision-Language Models 14 Mar 2025 · 0 repositories · arXiv:2503.11265
-
Exploring the Potential of Large Multimodal Models as Effective Alternatives for Pronunciation Assessment 14 Mar 2025 · 0 repositories · arXiv:2503.11229
-
Time and Memory Trade-off of KV-Cache Compression in Tensor Transformer Decoding 14 Mar 2025 · 0 repositories · arXiv:2503.11108
-
MEET: A Million-Scale Dataset for Fine-Grained Geospatial Scene Classification with Zoom-Free Remote Sensing Imagery 14 Mar 2025 · 0 repositories · arXiv:2503.11219
-
Prompt Sentiment: The Catalyst for LLM Change 14 Mar 2025 · 0 repositories · arXiv:2503.13510
-
RAG-KG-IL: A Multi-Agent Hybrid Framework for Reducing Hallucinations and Enhancing LLM Reasoning through RAG and Incremental Knowledge Graph Learning Integration 14 Mar 2025 · 0 repositories · arXiv:2503.13514
-
Relevance Isn't All You Need: Scaling RAG Systems With Inference-Time Compute Via Multi-Criteria Reranking 14 Mar 2025 · 2 repositories · arXiv:2504.07104
-
RESPONSE: Benchmarking the Ability of Language Models to Undertake Commonsense Reasoning in Crisis Situation 14 Mar 2025 · 0 repositories · arXiv:2503.11348
-
Solution for 8th Competition on Affective & Behavior Analysis in-the-wild 14 Mar 2025 · 0 repositories · arXiv:2503.11115
-
Text Compression for Efficient Language Generation 14 Mar 2025 · 0 repositories · arXiv:2503.11426
-
TransiT: Transient Transformer for Non-line-of-sight Videography 14 Mar 2025 · 0 repositories · arXiv:2503.11328
-
TreeMeshGPT: Artistic Mesh Generation with Autoregressive Tree Sequencing 14 Mar 2025 · 1 repository · arXiv:2503.11629Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 6 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples) · 2 pointer-only (licence)