Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 11
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 11 of 190: papers 1,001 to 1,100 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
DiTSE: High-Fidelity Generative Speech Enhancement via Latent Diffusion Transformers 13 Apr 2025 · 0 repositories · arXiv:2504.09381
-
Enhanced Filterless Multi-Color VLC via QCT 13 Apr 2025 · 0 repositories · arXiv:2504.09743
-
Ensemble-Enhanced Graph Autoencoder with GAT and Transformer-Based Encoders for Robust Fault Diagnosis 13 Apr 2025 · 0 repositories · arXiv:2504.09427
-
HD-RAG: Retrieval-Augmented Generation for Hybrid Documents Containing Text and Hierarchical Tables 13 Apr 2025 · 0 repositories · arXiv:2504.09554
-
HM-RAG: Hierarchical Multi-Agent Multimodal Retrieval Augmented Generation 13 Apr 2025 · 1 repository · arXiv:2504.12330
-
Integrating Large Language Models for Automated Structural Analysis 13 Apr 2025 · 0 repositories · arXiv:2504.09754
-
Iterative Self-Training for Code Generation via Reinforced Re-Ranking 13 Apr 2025 · 0 repositories · arXiv:2504.09643
-
Trajectory-guided Motion Perception for Facial Expression Quality Assessment in Neurological Disorders 13 Apr 2025 · 1 repository · arXiv:2504.09530
-
Accurate Diagnosis of Respiratory Viruses Using an Explainable Machine Learning with Mid-Infrared Biomolecular Fingerprinting of Nasopharyngeal Secretions 12 Apr 2025 · 0 repositories · arXiv:2504.09211
-
HeteRAG: A Heterogeneous Retrieval-augmented Generation Framework with Decoupled Knowledge Representations 12 Apr 2025 · 0 repositories · arXiv:2504.10529
-
Learning Occlusion-Robust Vision Transformers for Real-Time UAV Tracking 12 Apr 2025 · 1 repository · arXiv:2504.09228Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
Lumos: Efficient Performance Modeling and Estimation for Large-scale LLM Training 12 Apr 2025 · 0 repositories · arXiv:2504.09307
-
Multi-Modal Brain Tumor Segmentation via 3D Multi-Scale Self-attention and Cross-attention 12 Apr 2025 · 0 repositories · arXiv:2504.09088
-
Multi-scale Activation, Refinement, and Aggregation: Exploring Diverse Cues for Fine-Grained Bird Recognition 12 Apr 2025 · 0 repositories · arXiv:2504.09215
-
NetTAG: A Multimodal RTL-and-Layout-Aligned Netlist Foundation Model via Text-Attributed Graph 12 Apr 2025 · 1 repository · arXiv:2504.09260
-
Pneuma: Leveraging LLMs for Tabular Data Representation and Retrieval in an End-to-End System 12 Apr 2025 · 1 repository · arXiv:2504.09207
-
Semantic Commit: Helping Users Update Intent Specifications for AI Memory at Scale 12 Apr 2025 · 0 repositories · arXiv:2504.09283
-
Adaptive Additive Parameter Updates of Vision Transformers for Few-Shot Continual Learning 11 Apr 2025 · 0 repositories · arXiv:2504.08982
-
Adopting Large Language Models to Automated System Integration 11 Apr 2025 · 0 repositories · arXiv:2504.08490
-
DreamFuse: Adaptive Image Fusion with Diffusion Transformer 11 Apr 2025 · 0 repositories · arXiv:2504.08291
-
DrivAer Transformer: A high-precision and fast prediction method for vehicle aerodynamic drag coefficient based on the DrivAerNet++ dataset 11 Apr 2025 · 0 repositories · arXiv:2504.08217
-
Examining GPT's Capability to Generate and Map Course Concepts and Their Relationship 11 Apr 2025 · 0 repositories · arXiv:2504.08856
-
HyperCore: The Core Framework for Building Hyperbolic Foundation Models with Comprehensive Modules 11 Apr 2025 · 1 repository · arXiv:2504.08912Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Hypergraph Vision Transformers: Images are More than Nodes, More than Edges 11 Apr 2025 · 0 repositories · arXiv:2504.08710
-
Learning from Elders: Making an LLM-powered Chatbot for Retirement Communities more Accessible through User-centered Design 11 Apr 2025 · 0 repositories · arXiv:2504.08985
-
LLM for Comparative Narrative Analysis 11 Apr 2025 · 0 repositories · arXiv:2504.08211
-
LLMTaxo: Leveraging Large Language Models for Constructing Taxonomy of Factual Claims from Social Media 11 Apr 2025 · 0 repositories · arXiv:2504.12325
-
Millions of States: Designing a Scalable MoE Architecture with RWKV-7 Meta-learner 11 Apr 2025 · 0 repositories · arXiv:2504.08247
-
MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft 11 Apr 2025 · 0 repositories · arXiv:2504.08388
-
MixDiT: Accelerating Image Diffusion Transformer Inference with Mixed-Precision MX Quantization 11 Apr 2025 · 0 repositories · arXiv:2504.08398
-
Out of Style: RAG's Fragility to Linguistic Variation 11 Apr 2025 · 1 repository · arXiv:2504.08231Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 14 harvested samples)
-
PCA-RAG: Principal Component Analysis for Efficient Retrieval-Augmented Generation 11 Apr 2025 · 0 repositories · arXiv:2504.08386
-
RTLRepoCoder: Repository-Level RTL Code Completion through the Combination of Fine-Tuning and Retrieval Augmentation 11 Apr 2025 · 0 repositories · arXiv:2504.08862
-
SARFormer -- An Acquisition Parameter Aware Vision Transformer for Synthetic Aperture Radar Data 11 Apr 2025 · 0 repositories · arXiv:2504.08441
-
SWAN-GPT: An Efficient and Scalable Approach for Long-Context Language Modeling 11 Apr 2025 · 0 repositories · arXiv:2504.08719
-
The Other Side of the Coin: Exploring Fairness in Retrieval-Augmented Generation 11 Apr 2025 · 1 repository · arXiv:2504.12323
-
VLMT: Vision-Language Multimodal Transformer for Multimodal Multi-hop Question Answering 11 Apr 2025 · 0 repositories · arXiv:2504.08269
-
ZipIR: Latent Pyramid Diffusion Transformer for High-Resolution Image Restoration 11 Apr 2025 · 0 repositories · arXiv:2504.08591
-
A System for Comprehensive Assessment of RAG Frameworks 10 Apr 2025 · 1 repository · arXiv:2504.07803
-
AI Coding with Few-Shot Prompting for Thematic Analysis 10 Apr 2025 · 0 repositories · arXiv:2504.07408
-
AttentionDefense: Leveraging System Prompt Attention for Explainable Defense Against Novel Jailbreaks 10 Apr 2025 · 0 repositories · arXiv:2504.12321
-
Beating Transformers using Synthetic Cognition 10 Apr 2025 · 0 repositories · arXiv:2504.07619
-
Beyond Feature Importance: Feature Interactions in Predicting Post-Stroke Rigidity with Graph Explainable AI 10 Apr 2025 · 0 repositories · arXiv:2504.08150
-
Breaking the Barriers: Video Vision Transformers for Word-Level Sign Language Recognition 10 Apr 2025 · 0 repositories · arXiv:2504.07792
-
Can Reasoning LLMs Enhance Clinical Document Classification? 10 Apr 2025 · 1 repository · arXiv:2504.08040
-
ConceptFormer: Towards Efficient Use of Knowledge-Graph Embeddings in Large Language Models 10 Apr 2025 · 0 repositories · arXiv:2504.07624
-
Deep Learning Meets Teleconnections: Improving S2S Predictions for European Winter Weather 10 Apr 2025 · 1 repository · arXiv:2504.07625
-
Distilling Knowledge from Heterogeneous Architectures for Semantic Segmentation 10 Apr 2025 · 0 repositories · arXiv:2504.07691
-
Genetic Programming with Reinforcement Learning Trained Transformer for Real-World Dynamic Scheduling Problems 10 Apr 2025 · 0 repositories · arXiv:2504.07779
-
Has the Creativity of Large-Language Models peaked? An analysis of inter- and intra-LLM variability 10 Apr 2025 · 0 repositories · arXiv:2504.12320
-
Heart Failure Prediction using Modal Decomposition and Masked Autoencoders for Scarce Echocardiography Databases 10 Apr 2025 · 1 repository · arXiv:2504.07606
-
JEPA4Rec: Learning Effective Language Representations for Sequential Recommendation via Joint Embedding Predictive Architecture 10 Apr 2025 · 0 repositories · arXiv:2504.10512
-
MRD-RAG: Enhancing Medical Diagnosis with Multi-Round Retrieval-Augmented Generation 10 Apr 2025 · 1 repository · arXiv:2504.07724
-
Novel Pooling-based VGG-Lite for Pneumonia and Covid-19 Detection from Imbalanced Chest X-Ray Datasets 10 Apr 2025 · 0 repositories · arXiv:2504.07468
-
On the Practice of Deep Hierarchical Ensemble Network for Ad Conversion Rate Prediction 10 Apr 2025 · 0 repositories · arXiv:2504.08169
-
Pangu Ultra: Pushing the Limits of Dense Large Language Models on Ascend NPUs 10 Apr 2025 · 0 repositories · arXiv:2504.07866
-
PatchTrAD: A Patch-Based Transformer focusing on Patch-Wise Reconstruction Error for Time Series Anomaly Detection 10 Apr 2025 · 0 repositories · arXiv:2504.08827
-
PoGO: A Scalable Proof of Useful Work via Quantized Gradient Descent and Merkle Proofs 10 Apr 2025 · 0 repositories · arXiv:2504.07540
-
RadZero: Similarity-Based Cross-Attention for Explainable Vision-Language Alignment in Radiology with Zero-Shot Multi-Task Capability 10 Apr 2025 · 0 repositories · arXiv:2504.07416Syntology 5 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Revisiting Prompt Optimization with Large Reasoning Models-A Case Study on Event Extraction 10 Apr 2025 · 0 repositories · arXiv:2504.07357
-
Synthetic Fluency: Hallucinations, Confabulations, and the Creation of Irish Words in LLM-Generated Translations 10 Apr 2025 · 0 repositories · arXiv:2504.07680
-
AMAD: AutoMasked Attention for Unsupervised Multivariate Time Series Anomaly Detection 9 Apr 2025 · 0 repositories · arXiv:2504.06643
-
DyDiT++: Dynamic Diffusion Transformers for Efficient Visual Generation 9 Apr 2025 · 1 repository · arXiv:2504.06803
-
Endowing Embodied Agents with Spatial Reasoning Capabilities for Vision-and-Language Navigation 9 Apr 2025 · 0 repositories · arXiv:2504.08806
-
Evaluating Retrieval Augmented Generative Models for Document Queries in Transportation Safety 9 Apr 2025 · 0 repositories · arXiv:2504.07022
-
Face-LLaVA: Facial Expression and Attribute Understanding through Instruction Tuning 9 Apr 2025 · 0 repositories · arXiv:2504.07198
-
GenDoP: Auto-regressive Camera Trajectory Generation as a Director of Photography 9 Apr 2025 · 0 repositories · arXiv:2504.07083
-
Kaleidoscope: In-language Exams for Massively Multilingual Vision Evaluation 9 Apr 2025 · 1 repository · arXiv:2504.07072
-
Linguistic Interpretability of Transformer-based Language Models: a systematic review 9 Apr 2025 · 1 repository · arXiv:2504.08001
-
Poly-Vector Retrieval: Reference and Content Embeddings for Legal Documents 9 Apr 2025 · 0 repositories · arXiv:2504.10508
-
Text-to-Image Models and Their Representation of People from Different Nationalities Engaging in Activities 8 Apr 2025 · 0 repositories · arXiv:2504.06313
-
Assessing how hyperparameters impact Large Language Models' sarcasm detection performance 8 Apr 2025 · 0 repositories · arXiv:2504.06166
-
Fusing Global and Local: Transformer-CNN Synergy for Next-Gen Current Estimation 8 Apr 2025 · 0 repositories · arXiv:2504.07996
-
Graph-based Approaches and Functionalities in Retrieval-Augmented Generation: A Comprehensive Survey 8 Apr 2025 · 0 repositories · arXiv:2504.10499
-
Leveraging Auto-Distillation and Generative Self-Supervised Learning in Residual Graph Transformers for Enhanced Recommender Systems 8 Apr 2025 · 0 repositories · arXiv:2504.10500
-
PathGPT: Leveraging Large Language Models for Personalized Route Generation 8 Apr 2025 · 0 repositories · arXiv:2504.05846
-
Rethinking the Nested U-Net Approach: Enhancing Biomarker Segmentation with Attention Mechanisms and Multiscale Feature Fusion 8 Apr 2025 · 1 repository · arXiv:2504.06158
-
Retrieval Augmented Generation with Collaborative Filtering for Personalized Text Generation 8 Apr 2025 · 1 repository · arXiv:2504.05731
-
AI for Climate Finance: Agentic Retrieval and Multi-Step Reasoning for Early Warning System Investments 7 Apr 2025 · 0 repositories · arXiv:2504.05104
-
Boundary representation learning via Transformer 7 Apr 2025 · 0 repositories · arXiv:2504.07134
-
Collab-RAG: Boosting Retrieval-Augmented Generation for Complex Question Answering via White-Box and Black-Box LLM Collaboration 7 Apr 2025 · 1 repository · arXiv:2504.04915Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Content-Aware Transformer for All-in-one Image Restoration 7 Apr 2025 · 1 repository · arXiv:2504.04869
-
InstructionBench: An Instructional Video Understanding Benchmark 7 Apr 2025 · 0 repositories · arXiv:2504.05040
-
Leveraging LLMs for Utility-Focused Annotation: Reducing Manual Effort for Retrieval and RAG 7 Apr 2025 · 0 repositories · arXiv:2504.05220
-
Lumina-OmniLV: A Unified Multimodal Framework for General Low-Level Vision 7 Apr 2025 · 0 repositories · arXiv:2504.04903
-
OmniEcon Nexus: Global Microeconomic Simulation Engine 7 Apr 2025 · 1 repository
-
One-Minute Video Generation with Test-Time Training 7 Apr 2025 · 0 repositories · arXiv:2504.05298
-
One Quantizer is Enough: Toward a Lightweight Audio Codec 7 Apr 2025 · 1 repository · arXiv:2504.04949
-
Provable Failure of Language Models in Learning Majority Boolean Logic via Gradient Descent 7 Apr 2025 · 0 repositories · arXiv:2504.04702
-
REWIND: Real-Time Egocentric Whole-Body Motion Diffusion with Exemplar-Based Identity Conditioning 7 Apr 2025 · 0 repositories · arXiv:2504.04956
-
Weak-for-Strong: Training Weak Meta-Agent to Harness Strong Executors 7 Apr 2025 · 1 repository · arXiv:2504.04785
-
Multimodal Cinematic Video Synthesis Using Text-to-Image and Audio Generation Models 6 Apr 2025 · 0 repositories · arXiv:2506.10005
-
DanceMosaic: High-Fidelity Dance Generation with Multimodal Editability 6 Apr 2025 · 0 repositories · arXiv:2504.04634
-
Driving-RAG: Driving Scenarios Embedding, Search, and RAG Applications 6 Apr 2025 · 0 repositories · arXiv:2504.04419
-
Hierarchical Planning for Complex Tasks with Knowledge Graph-RAG and Symbolic Verification 6 Apr 2025 · 0 repositories · arXiv:2504.04578
-
CATS: Mitigating Correlation Shift for Multivariate Time Series Classification 5 Apr 2025 · 0 repositories · arXiv:2504.04283
-
QE-RAG: A Robust Retrieval-Augmented Generation Benchmark for Query Entry Errors 5 Apr 2025 · 0 repositories · arXiv:2504.04062
-
Quantum Adaptive Self-Attention for Quantum Transformer Models 5 Apr 2025 · 0 repositories · arXiv:2504.05336
-
Sigma: A dataset for text-to-code semantic parsing with statistical analysis 5 Apr 2025 · 1 repository · arXiv:2504.04301
-
Transformer representation learning is necessary for dynamic multi-modal physiological data on small-cohort patients 5 Apr 2025 · 0 repositories · arXiv:2504.04120