Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 16
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 16 of 190: papers 1,501 to 1,600 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
End-to-End HOI Reconstruction Transformer with Graph-based Encoding 8 Mar 2025 · 0 repositories · arXiv:2503.06012
-
Fish2Mesh Transformer: 3D Human Mesh Recovery from Egocentric Vision 8 Mar 2025 · 0 repositories · arXiv:2503.06089
-
Lightweight Software Kernels and Hardware Extensions for Efficient Sparse Deep Neural Networks on Microcontrollers 8 Mar 2025 · 0 repositories · arXiv:2503.06183
-
LimTopic: LLM-based Topic Modeling and Text Summarization for Analyzing Scientific Articles limitations 8 Mar 2025 · 1 repository · arXiv:2503.10658
-
MoEMoE: Question Guided Dense and Scalable Sparse Mixture-of-Expert for Multi-source Multi-modal Answering 8 Mar 2025 · 0 repositories · arXiv:2503.06296
-
Optimizing Generative AI's Accuracy and Transparency in Inductive Thematic Analysis: A Human-AI Comparison 8 Mar 2025 · 0 repositories · arXiv:2503.16485
-
Poisoned-MRAG: Knowledge Poisoning Attacks to Multimodal Retrieval Augmented Generation 8 Mar 2025 · 0 repositories · arXiv:2503.06254
-
X2I: Seamless Integration of Multimodal Understanding into Diffusion Transformer via Attention Distillation 8 Mar 2025 · 1 repository · arXiv:2503.06134Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
A Hybrid Model/Data-Driven Solution to Channel, Position and Orientation Tracking in mmWave Vehicular Systems 7 Mar 2025 · 0 repositories · arXiv:2503.05091
-
A Real-time Multimodal Transformer Neural Network-powered Wildfire Forecasting System 7 Mar 2025 · 0 repositories · arXiv:2503.05971
-
BARK: A Fully Bayesian Tree Kernel for Black-box Optimization 7 Mar 2025 · 0 repositories · arXiv:2503.05574
-
Evaluating Large Language Models in Code Generation: INFINITE Methodology for Defining the Inference Index 7 Mar 2025 · 0 repositories · arXiv:2503.05852
-
Simulating and Analysing Human Survey Responses with Large Language Models: A Case Study in Energy Stated Preference 7 Mar 2025 · 0 repositories · arXiv:2503.10652
-
FastMap: Fast Queries Initialization Based Vectorized HD Map Reconstruction Framework 7 Mar 2025 · 1 repository · arXiv:2503.05492
-
FMCHS: Advancing Traditional Chinese Medicine Herb Recommendation with Fusion of Multiscale Correlations of Herbs and Symptoms 7 Mar 2025 · 0 repositories · arXiv:2503.05167
-
FMT:A Multimodal Pneumonia Detection Model Based on Stacking MOE Framework 7 Mar 2025 · 0 repositories · arXiv:2503.05626
-
Language modelling techniques for analysing the impact of human genetic variation 7 Mar 2025 · 0 repositories · arXiv:2503.10655
-
Leveraging Approximate Caching for Faster Retrieval-Augmented Generation 7 Mar 2025 · 0 repositories · arXiv:2503.05530
-
MagicInfinite: Generating Infinite Talking Videos with Your Words and Voice 7 Mar 2025 · 0 repositories · arXiv:2503.05978
-
Quantifying the Robustness of Retrieval-Augmented Language Models Against Spurious Features in Grounding Data 7 Mar 2025 · 0 repositories · arXiv:2503.05587
-
R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning 7 Mar 2025 · 5 repositories · arXiv:2503.05592Syntology 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Tractable Representations for Convergent Approximation of Distributional HJB Equations 7 Mar 2025 · 0 repositories · arXiv:2503.05563
-
Zero-shot Medical Event Prediction Using a Generative Pre-trained Transformer on Electronic Health Records 7 Mar 2025 · 0 repositories · arXiv:2503.05893
-
A Generalist Cross-Domain Molecular Learning Framework for Structure-Based Drug Discovery 6 Mar 2025 · 0 repositories · arXiv:2503.04362
-
Beyond RAG: Task-Aware KV Cache Compression for Comprehensive Knowledge Reasoning 6 Mar 2025 · 0 repositories · arXiv:2503.04973
-
BicliqueEncoder: An Efficient Method for Link Prediction in Bipartite Networks using Formal Concept Analysis and Transformer Encoder 6 Mar 2025 · 0 repositories · arXiv:2503.07645
-
Can We Optimize Deep RL Policy Weights as Trajectory Modeling? 6 Mar 2025 · 0 repositories · arXiv:2503.04074
-
Collapse of Dense Retrievers: Short, Early, and Literal Biases Outranking Factual Evidence 6 Mar 2025 · 0 repositories · arXiv:2503.05037
-
Compositional Causal Reasoning Evaluation in Language Models 6 Mar 2025 · 0 repositories · arXiv:2503.04556
-
DB-Explore: Automated Database Exploration and Instruction Synthesis for Text-to-SQL 6 Mar 2025 · 0 repositories · arXiv:2503.04959
-
GBT-SAM: Adapting a Foundational Deep Learning Model for Generalizable Brain Tumor Segmentation via Efficient Integration of Multi-Parametric MRI Data 6 Mar 2025 · 1 repository · arXiv:2503.04325
-
Hedging with Sparse Reward Reinforcement Learning 6 Mar 2025 · 0 repositories · arXiv:2503.04218
-
High-Precision Transformer-Based Visual Servoing for Humanoid Robots in Aligning Tiny Objects 6 Mar 2025 · 0 repositories · arXiv:2503.04862
-
HILGEN: Hierarchically-Informed Data Generation for Biomedical NER Using Knowledgebases and Large Language Models 6 Mar 2025 · 0 repositories · arXiv:2503.04930
-
In-depth Analysis of Graph-based RAG in a Unified Framework 6 Mar 2025 · 0 repositories · arXiv:2503.04338
-
Incentivizing Multi-Tenant Split Federated Learning for Foundation Models at the Network Edge 6 Mar 2025 · 0 repositories · arXiv:2503.04971
-
Learning Transformer-based World Models with Contrastive Predictive Coding 6 Mar 2025 · 0 repositories · arXiv:2503.04416
-
Leveraging Large Language Models to Address Data Scarcity in Machine Learning: Applications in Graphene Synthesis 6 Mar 2025 · 1 repository · arXiv:2503.04870
-
Toward Lightweight and Fast Decoders for Diffusion Models in Image and Video Generation 6 Mar 2025 · 1 repository · arXiv:2503.04871
-
Towards Autonomous Reinforcement Learning for Real-World Robotic Manipulation with Large Language Models 6 Mar 2025 · 0 repositories · arXiv:2503.04280
-
A Multimodal Framework for Topic Propagation Classification in Social Networks 5 Mar 2025 · 0 repositories · arXiv:2503.03112
-
All-atom Diffusion Transformers: Unified generative modelling of molecules and materials 5 Mar 2025 · 1 repository · arXiv:2503.03965Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
DTU-Net: A Multi-Scale Dilated Transformer Network for Nonlinear Hyperspectral Unmixing 5 Mar 2025 · 0 repositories · arXiv:2503.03465
-
Large language models in finance : what is financial sentiment? 5 Mar 2025 · 0 repositories · arXiv:2503.03612
-
MA-LoT: Multi-Agent Lean-based Long Chain-of-Thought Reasoning enhances Formal Theorem Proving 5 Mar 2025 · 1 repository · arXiv:2503.03205
-
PathRWKV: Enabling Whole Slide Prediction with Recurrent-Transformer 5 Mar 2025 · 0 repositories · arXiv:2503.03199
-
Pretrained LLMs as Real-Time Controllers for Robot Operated Serial Production Line 5 Mar 2025 · 0 repositories · arXiv:2503.03889
-
RiskAgent: Autonomous Medical AI Copilot for Generalist Risk Prediction 5 Mar 2025 · 0 repositories · arXiv:2503.03802
-
ScaleFusionNet: Transformer-Guided Multi-Scale Feature Fusion for Skin Lesion Segmentation 5 Mar 2025 · 1 repository · arXiv:2503.03327
-
The Box is in the Pen: Evaluating Commonsense Reasoning in Neural Machine Translation 5 Mar 2025 · 1 repository · arXiv:2503.03308
-
A Transformer Model for Predicting Chemical Reaction Products from Generic Templates 4 Mar 2025 · 0 repositories · arXiv:2503.05810
-
BHViT: Binarized Hybrid Vision Transformer 4 Mar 2025 · 1 repository · arXiv:2503.02394Syntology official (archive's flag): 16 ran · 16 ran (of which 13 constructed an object rather than computing a result; 15 with no instrument failure: 0 honoured, 0 violated, 15 with no contract checked; 1 where Syntology's instrument failed) · 13 unverified (of 29 harvested samples)
-
CoServe: Efficient Collaboration-of-Experts (CoE) Model Inference with Limited Memory 4 Mar 2025 · 0 repositories · arXiv:2503.02354
-
Developing a PET/CT Foundation Model for Cross-Modal Anatomical and Functional Imaging 4 Mar 2025 · 0 repositories · arXiv:2503.02824
-
Effectively Steer LLM To Follow Preference via Building Confident Directions 4 Mar 2025 · 0 repositories · arXiv:2503.02989
-
FourierNAT: A Fourier-Mixing-Based Non-Autoregressive Transformer for Parallel Sequence Generation 4 Mar 2025 · 0 repositories · arXiv:2503.07630
-
Graph Transformer with Disease Subgraph Positional Encoding for Improved Comorbidity Prediction 4 Mar 2025 · 1 repository · arXiv:2503.03046
-
Interpretable Few-Shot Retinal Disease Diagnosis with Concept-Guided Prompting of Vision-Language Models 4 Mar 2025 · 0 repositories · arXiv:2503.02917
-
Learning Precoding in Multi-user Multi-antenna Systems: Transformer or Graph Transformer? 4 Mar 2025 · 0 repositories · arXiv:2503.02998
-
LLaVE: Large Language and Vision Embedding Models with Hardness-Weighted Contrastive Learning 4 Mar 2025 · 0 repositories · arXiv:2503.04812
-
Network Traffic Classification Using Machine Learning, Transformer, and Large Language Models 4 Mar 2025 · 0 repositories · arXiv:2503.02141
-
Optimizing open-domain question answering with graph-based retrieval augmented generation 4 Mar 2025 · 0 repositories · arXiv:2503.02922
-
PennyLang: Pioneering LLM-Based Quantum Code Generation with a Novel PennyLane-Centric Dataset 4 Mar 2025 · 0 repositories · arXiv:2503.02497
-
Tabby: Tabular Data Synthesis with Language Models 4 Mar 2025 · 0 repositories · arXiv:2503.02152
-
Target Return Optimizer for Multi-Game Decision Transformer 4 Mar 2025 · 0 repositories · arXiv:2503.02311
-
TeTRA-VPR: A Ternary Transformer Approach for Compact Visual Place Recognition 4 Mar 2025 · 0 repositories · arXiv:2503.02511
-
Use Me Wisely: AI-Driven Assessment for LLM Prompting Skills Development 4 Mar 2025 · 0 repositories · arXiv:2503.02532
-
Weak-to-Strong Generalization Even in Random Feature Networks, Provably 4 Mar 2025 · 0 repositories · arXiv:2503.02877
-
Wikipedia in the Era of LLMs: Evolution and Risks 4 Mar 2025 · 1 repository · arXiv:2503.02879
-
Wyckoff Transformer: Generation of Symmetric Crystals 4 Mar 2025 · 1 repository · arXiv:2503.02407
-
Zero-Shot Multi-Label Classification of Bangla Documents: Large Decoders Vs. Classic Encoders 4 Mar 2025 · 0 repositories · arXiv:2503.02993
-
From Claims to Evidence: A Unified Framework and Critical Analysis of CNN vs. Transformer vs. Mamba in Medical Image Segmentation 3 Mar 2025 · 1 repository · arXiv:2503.01306
-
SrSv: Integrating Sequential Rollouts with Sequential Value Estimation for Multi-agent Reinforcement Learning 3 Mar 2025 · 0 repositories · arXiv:2503.01458
-
An Efficient Approach to Detecting Lung Nodules Using Swin Transformer 3 Mar 2025 · 0 repositories · arXiv:2503.01592
-
Machine Learners Should Acknowledge the Legal Implications of Large Language Models as Personal Data 3 Mar 2025 · 0 repositories · arXiv:2503.01630
-
SAGE: A Framework of Precise Retrieval for RAG 3 Mar 2025 · 0 repositories · arXiv:2503.01713
-
LLMInit: A Free Lunch from Large Language Models for Selective Initialization of Recommendation 3 Mar 2025 · 0 repositories · arXiv:2503.01814
-
A Hybrid CNN-Transformer Model for Heart Disease Prediction Using Life History Data 3 Mar 2025 · 0 repositories · arXiv:2503.02124
-
Architectural and Inferential Inductive Biases For Exchangeable Sequence Modeling 3 Mar 2025 · 1 repository · arXiv:2503.01215
-
AskToAct: Enhancing LLMs Tool Use via Self-Correcting Clarification 3 Mar 2025 · 0 repositories · arXiv:2503.01940
-
Attention Condensation via Sparsity Induced Regularized Training 3 Mar 2025 · 0 repositories · arXiv:2503.01564
-
Cancer Type, Stage and Prognosis Assessment from Pathology Reports using LLMs 3 Mar 2025 · 1 repository · arXiv:2503.01194
-
Dementia Insights: A Context-Based MultiModal Approach 3 Mar 2025 · 0 repositories · arXiv:2503.01226
-
Forgetting Transformer: Softmax Attention with a Forget Gate 3 Mar 2025 · 1 repository · arXiv:2503.02130Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples)
-
HeterRec: Heterogeneous Information Transformer for Scalable Sequential Recommendation 3 Mar 2025 · 0 repositories · arXiv:2503.01469
-
HoH: A Dynamic Benchmark for Evaluating the Impact of Outdated Information on Retrieval-Augmented Generation 3 Mar 2025 · 0 repositories · arXiv:2503.04800Syntology 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
How simple can you go? An off-the-shelf transformer approach to molecular dynamics 3 Mar 2025 · 1 repository · arXiv:2503.01431
-
Interactive Gadolinium-Free MRI Synthesis: A Transformer with Localization Prompt Learning 3 Mar 2025 · 1 repository · arXiv:2503.01265
-
MeshPad: Interactive Sketch-Conditioned Artist-Designed Mesh Generation and Editing 3 Mar 2025 · 0 repositories · arXiv:2503.01425
-
Primus: Enforcing Attention Usage for 3D Medical Image Segmentation 3 Mar 2025 · 0 repositories · arXiv:2503.01835
-
Retrieval-Augmented Perception: High-Resolution Image Perception Meets Visual RAG 3 Mar 2025 · 1 repository · arXiv:2503.01222
-
SePer: Measure Retrieval Utility Through The Lens Of Semantic Perplexity Reduction 3 Mar 2025 · 1 repository · arXiv:2503.01478Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples)
-
SRAG: Structured Retrieval-Augmented Generation for Multi-Entity Question Answering over Wikipedia Graph 3 Mar 2025 · 0 repositories · arXiv:2503.01346
-
Streaming Piano Transcription Based on Consistent Onset and Offset Decoding with Sustain Pedal Detection 3 Mar 2025 · 0 repositories · arXiv:2503.01362
-
Syntactic Learnability of Echo State Neural Language Models at Scale 3 Mar 2025 · 0 repositories · arXiv:2503.01724
-
Unify and Anchor: A Context-Aware Transformer for Cross-Domain Time Series Forecasting 3 Mar 2025 · 0 repositories · arXiv:2503.01157
-
Using (Not so) Large Language Models for Generating Simulation Models in a Formal DSL -- A Study on Reaction Networks 3 Mar 2025 · 0 repositories · arXiv:2503.01675
-
ViKANformer: Embedding Kolmogorov Arnold Networks in Vision Transformers for Pattern-Based Learning 3 Mar 2025 · 0 repositories · arXiv:2503.01124
-
SemViQA: A Semantic Question Answering System for Vietnamese Information Fact-Checking 2 Mar 2025 · 1 repository · arXiv:2503.00955
-
ER-RAG: Enhance RAG with ER-Based Unified Modeling of Heterogeneous Data Sources 2 Mar 2025 · 0 repositories · arXiv:2504.06271