Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 13
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 13 of 190: papers 1,201 to 1,300 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
MHTS: Multi-Hop Tree Structure Framework for Generating Difficulty-Controllable QA Datasets for RAG Evaluation 29 Mar 2025 · 0 repositories · arXiv:2504.08756
-
Multimodal machine learning with large language embedding model for polymer property prediction 29 Mar 2025 · 1 repository · arXiv:2503.22962
-
The geomagnetic storm and Kp prediction using Wasserstein transformer 29 Mar 2025 · 0 repositories · arXiv:2503.23102
-
The realization of tones in spontaneous spoken Taiwan Mandarin: a corpus-based survey and theory-driven computational modeling 29 Mar 2025 · 0 repositories · arXiv:2503.23163
-
An Advanced Ensemble Deep Learning Framework for Stock Price Prediction Using VAE, Transformer, and LSTM Model 28 Mar 2025 · 0 repositories · arXiv:2503.22192
-
AnnoPage Dataset: Dataset of Non-Textual Elements in Documents with Fine-Grained Categorization 28 Mar 2025 · 0 repositories · arXiv:2503.22526
-
Bridging the Dimensional Chasm: Uncover Layer-wise Dimensional Reduction in Transformers through Token Correlation 28 Mar 2025 · 0 repositories · arXiv:2503.22547
-
Camera Model Identification with SPAIR-Swin and Entropy based Non-Homogeneous Patches 28 Mar 2025 · 0 repositories · arXiv:2503.22120
-
Correlation-Attention Masked Temporal Transformer for User Identity Linkage Using Heterogeneous Mobility Data 28 Mar 2025 · 1 repository · arXiv:2504.01979
-
DREMnet: An Interpretable Denoising Framework for Semi-Airborne Transient Electromagnetic Signal 28 Mar 2025 · 0 repositories · arXiv:2503.22223
-
EdgeInfinite: A Memory-Efficient Infinite-Context Transformer for Edge Devices 28 Mar 2025 · 0 repositories · arXiv:2503.22196
-
How Well Can Vison-Language Models Understand Humans' Intention? An Open-ended Theory of Mind Question Evaluation Benchmark 28 Mar 2025 · 0 repositories · arXiv:2503.22093
-
Integrating Artificial Intelligence with Human Expertise: An In-depth Analysis of ChatGPT's Capabilities in Generating Metamorphic Relations 28 Mar 2025 · 0 repositories · arXiv:2503.22141
-
Leveraging LLMs for Predicting Unknown Diagnoses from Clinical Notes 28 Mar 2025 · 0 repositories · arXiv:2503.22092
-
Scenario Dreamer: Vectorized Latent Diffusion for Generating Driving Simulation Environments 28 Mar 2025 · 0 repositories · arXiv:2503.22496
-
Understanding Inequality of LLM Fact-Checking over Geographic Regions with Agent and Retrieval models 28 Mar 2025 · 0 repositories · arXiv:2503.22877
-
An evaluation of LLMs and Google Translate for translation of selected Indian languages via sentiment and semantic analyses 27 Mar 2025 · 0 repositories · arXiv:2503.21393
-
Collab: Controlled Decoding using Mixture of Agents for LLM Alignment 27 Mar 2025 · 0 repositories · arXiv:2503.21720
-
HyperGraphRAG: Retrieval-Augmented Generation with Hypergraph-Structured Knowledge Representation 27 Mar 2025 · 1 repository · arXiv:2503.21322Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Integrating Travel Behavior Forecasting and Generative Modeling for Predicting Future Urban Mobility and Spatial Transformations 27 Mar 2025 · 0 repositories · arXiv:2503.21158
-
MemInsight: Autonomous Memory Augmentation for LLM Agents 27 Mar 2025 · 0 repositories · arXiv:2503.21760
-
Molecular Quantum Transformer 27 Mar 2025 · 0 repositories · arXiv:2503.21686
-
Real-Time Evaluation Models for RAG: Who Detects Hallucinations Best? 27 Mar 2025 · 0 repositories · arXiv:2503.21157
-
ReaRAG: Knowledge-guided Reasoning Enhances Factuality of Large Reasoning Models with Iterative Retrieval Augmented Generation 27 Mar 2025 · 1 repository · arXiv:2503.21729
-
ReCoM: Realistic Co-Speech Motion Generation with Recurrent Embedded Transformer 27 Mar 2025 · 0 repositories · arXiv:2503.21847
-
Retinal Fundus Multi-Disease Image Classification using Hybrid CNN-Transformer-Ensemble Architectures 27 Mar 2025 · 1 repository · arXiv:2503.21465
-
Using large language models to produce literature reviews: Usages and systematic biases of microphysics parametrizations in 2699 publications 27 Mar 2025 · 0 repositories · arXiv:2503.21352
-
VALLR: Visual ASR Language Model for Lip Reading 27 Mar 2025 · 0 repositories · arXiv:2503.21408
-
Vision Language Models versus Machine Learning Models Performance on Polyp Detection and Classification in Colonoscopy Images 27 Mar 2025 · 1 repository · arXiv:2503.21840
-
A Survey of Multimodal Retrieval-Augmented Generation 26 Mar 2025 · 0 repositories · arXiv:2504.08748
-
Advancements in Natural Language Processing: Exploring Transformer-Based Architectures for Text Understanding 26 Mar 2025 · 0 repositories · arXiv:2503.20227
-
Can We Make Code Green? Understanding Trade-Offs in LLMs vs. Human Code Optimizations 26 Mar 2025 · 0 repositories · arXiv:2503.20126
-
CNN+Transformer Based Anomaly Traffic Detection in UAV Networks for Emergency Rescue 26 Mar 2025 · 0 repositories · arXiv:2503.20355
-
Novel Deep Neural OFDM Receiver Architectures for LLR Estimation 26 Mar 2025 · 1 repository · arXiv:2503.20500
-
Devil is in the Uniformity: Exploring Diverse Learners within Transformer for Image Restoration 26 Mar 2025 · 1 repository · arXiv:2503.20174Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
In vitro 2 In vivo : Bidirectional and High-Precision Generation of In Vitro and In Vivo Neuronal Spike Data 26 Mar 2025 · 0 repositories · arXiv:2503.20841
-
ITA-MDT: Image-Timestep-Adaptive Masked Diffusion Transformer Framework for Image-Based Virtual Try-On 26 Mar 2025 · 0 repositories · arXiv:2503.20418
-
Iterative Prompting with Persuasion Skills in Jailbreaking Large Language Models 26 Mar 2025 · 0 repositories · arXiv:2503.20320
-
MCTS-RAG: Enhancing Retrieval-Augmented Generation with Monte Carlo Tree Search 26 Mar 2025 · 1 repository · arXiv:2503.20757
-
MVFNet: Multipurpose Video Forensics Network using Multiple Forms of Forensic Evidence 26 Mar 2025 · 0 repositories · arXiv:2503.20991
-
Patients Speak, AI Listens: LLM-based Analysis of Online Reviews Uncovers Key Drivers for Urgent Care Satisfaction 26 Mar 2025 · 0 repositories · arXiv:2503.20981
-
Progressive Focused Transformer for Single Image Super-Resolution 26 Mar 2025 · 1 repository · arXiv:2503.20337Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 6 where Syntology's instrument failed) · 1 unverified (of 10 harvested samples) · 3 pointer-only (licence)
-
RALLRec+: Retrieval Augmented Large Language Model Recommendation with Reasoning 26 Mar 2025 · 1 repository · arXiv:2503.20430
-
RSRWKV: A Linear-Complexity 2D Attention Mechanism for Efficient Remote Sensing Vision Task 26 Mar 2025 · 0 repositories · arXiv:2503.20382
-
RGL: A Graph-Centric, Modular Framework for Efficient Retrieval-Augmented Generation on Graphs 25 Mar 2025 · 1 repository · arXiv:2503.19314
-
A novel forecasting framework combining virtual samples and enhanced Transformer models for tourism demand forecasting 25 Mar 2025 · 0 repositories · arXiv:2503.19423
-
BiblioPage: A Dataset of Scanned Title Pages for Bibliographic Metadata Extraction 25 Mar 2025 · 1 repository · arXiv:2503.19658
-
CausalRAG: Integrating Causal Graphs into Retrieval-Augmented Generation 25 Mar 2025 · 0 repositories · arXiv:2503.19878
-
Context-Aware Semantic Segmentation: Enhancing Pixel-Level Understanding with Large Language Models for Advanced Vision Applications 25 Mar 2025 · 0 repositories · arXiv:2503.19276
-
Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy 25 Mar 2025 · 1 repository · arXiv:2503.19757Syntology 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Enabling Rapid Shared Human-AI Mental Model Alignment via the After-Action Review 25 Mar 2025 · 1 repository · arXiv:2503.19607
-
Face Spoofing Detection using Deep Learning 25 Mar 2025 · 1 repository · arXiv:2503.19223
-
Fundamental Limits of Perfect Concept Erasure 25 Mar 2025 · 1 repository · arXiv:2503.20098
-
GPT Meets Graphs and KAN Splines: Testing Novel Frameworks on Multitask Fine-Tuned GPT-2 with LoRA 25 Mar 2025 · 0 repositories · arXiv:2504.10490
-
iNatAg: Multi-Class Classification Models Enabled by a Large-Scale Benchmark Dataset with 4.7M Images of 2,959 Crop and Weed Species 25 Mar 2025 · 1 repository · arXiv:2503.20068
-
M²CD: A Unified MultiModal Framework for Optical-SAR Change Detection with Mixture of Experts and Self-Distillation 25 Mar 2025 · 0 repositories · arXiv:2503.19406
-
Mask²DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video Generation 25 Mar 2025 · 0 repositories · arXiv:2503.19881
-
RoboFlamingo-Plus: Fusion of Depth and RGB Perception with Vision-Language Models for Enhanced Robotic Manipulation 25 Mar 2025 · 0 repositories · arXiv:2503.19510
-
Scaling Down Text Encoders of Text-to-Image Diffusion Models 25 Mar 2025 · 1 repository · arXiv:2503.19897
-
SCI-IDEA: Context-Aware Scientific Ideation Using Token and Sentence Embeddings 25 Mar 2025 · 0 repositories · arXiv:2503.19257
-
Social Network User Profiling for Anomaly Detection Based on Graph Neural Networks 25 Mar 2025 · 0 repositories · arXiv:2503.19380
-
Taxonomy Inference for Tabular Data Using Large Language Models 25 Mar 2025 · 0 repositories · arXiv:2503.21810
-
VGAT: A Cancer Survival Analysis Framework Transitioning from Generative Visual Question Answering to Genomic Reconstruction 25 Mar 2025 · 1 repository · arXiv:2503.19367
-
Chirp Localization via Fine-Tuned Transformer Model: A Proof-of-Concept Study 24 Mar 2025 · 0 repositories · arXiv:2503.22713
-
Context-Enhanced Memory-Refined Transformer for Online Action Detection 24 Mar 2025 · 1 repository · arXiv:2503.18359
-
Distil-xLSTM: Learning Attention Mechanisms through Recurrent Structures 24 Mar 2025 · 0 repositories · arXiv:2503.18565
-
Exploring the Integration of Key-Value Attention Into Pure and Hybrid Transformers for Semantic Segmentation 24 Mar 2025 · 0 repositories · arXiv:2503.18862
-
Exploring Training and Inference Scaling Laws in Generative Retrieval 24 Mar 2025 · 1 repository · arXiv:2503.18941
-
Global-Local Tree Search in VLMs for 3D Indoor Scene Generation 24 Mar 2025 · 1 repository · arXiv:2503.18476Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
How to Capture and Study Conversations Between Research Participants and ChatGPT: GPT for Researchers (g4r.org) 24 Mar 2025 · 0 repositories · arXiv:2503.18303
-
Improving RAG for Personalization with Author Features and Contrastive Examples 24 Mar 2025 · 1 repository · arXiv:2504.08745
-
LGI-DETR: Local-Global Interaction for UAV Object Detection 24 Mar 2025 · 0 repositories · arXiv:2503.18785
-
LoTUS: Large-Scale Machine Unlearning with a Taste of Uncertainty 24 Mar 2025 · 1 repository · arXiv:2503.18314
-
Predicting the Road Ahead: A Knowledge Graph based Foundation Model for Scene Understanding in Autonomous Driving 24 Mar 2025 · 0 repositories · arXiv:2503.18730
-
REALM: A Dataset of Real-World LLM Use Cases 24 Mar 2025 · 0 repositories · arXiv:2503.18792
-
SPMTrack: Spatio-Temporal Parameter-Efficient Fine-Tuning with Mixture of Experts for Scalable Visual Tracking 24 Mar 2025 · 1 repository · arXiv:2503.18338Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Synthetic Function Demonstrations Improve Generation in Low-Resource Programming Languages 24 Mar 2025 · 0 repositories · arXiv:2503.18760
-
U-REPA: Aligning Diffusion U-Nets to ViTs 24 Mar 2025 · 1 repository · arXiv:2503.18414
-
Your ViT is Secretly an Image Segmentation Model 24 Mar 2025 · 1 repository · arXiv:2503.19108Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
ZeroLM: Data-Free Transformer Architecture Search for Language Models 24 Mar 2025 · 0 repositories · arXiv:2503.18646
-
Adaptive Rank Allocation: Speeding Up Modern Transformers with RaNA Adapters 23 Mar 2025 · 1 repository · arXiv:2503.18216Syntology official (archive's flag): 3 ran · 3 ran (of which 1 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
End-to-End Implicit Neural Representations for Classification 23 Mar 2025 · 1 repository · arXiv:2503.18123
-
ExpertRAG: Efficient RAG with Mixture of Experts -- Optimizing Context Retrieval for Adaptive LLM Responses 23 Mar 2025 · 0 repositories · arXiv:2504.08744
-
Investigating Recent Large Language Models for Vietnamese Machine Reading Comprehension 23 Mar 2025 · 0 repositories · arXiv:2503.18062
-
PathoHR: Breast Cancer Survival Prediction on High-Resolution Pathological Images 23 Mar 2025 · 1 repository · arXiv:2503.17970
-
Retrieval Augmented Generation and Understanding in Vision: A Survey and New Outlook 23 Mar 2025 · 1 repository · arXiv:2503.18016
-
SymmCompletion: High-Fidelity and High-Consistency Point Cloud Completion with Symmetry Guidance 23 Mar 2025 · 1 repository · arXiv:2503.18007
-
A Modular Dataset to Demonstrate LLM Abstraction Capability 22 Mar 2025 · 0 repositories · arXiv:2503.17645
-
Automated diagnosis of lung diseases using vision transformer: a comparative study on chest x-ray classification 22 Mar 2025 · 0 repositories · arXiv:2503.18973
-
Bandwidth Reservation for Time-Critical Vehicular Applications: A Multi-Operator Environment 22 Mar 2025 · 0 repositories · arXiv:2503.17756
-
EMPLACE: Self-Supervised Urban Scene Change Detection 22 Mar 2025 · 1 repository · arXiv:2503.17716Syntology official (archive's flag): 4 ran · 4 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified; every one of the 4 samples that ran constructed an object rather than computing a result (of 7 harvested samples) · 7 pointer-only (licence)
-
Hierarchy-Aware and Channel-Adaptive Semantic Communication for Bandwidth-Limited Data Fusion 22 Mar 2025 · 0 repositories · arXiv:2503.17777
-
TDRI: Two-Phase Dialogue Refinement and Co-Adaptation for Interactive Image Generation 22 Mar 2025 · 0 repositories · arXiv:2503.17669
-
Assessing the Reliability and Validity of GPT-4 in Annotating Emotion Appraisal Ratings 21 Mar 2025 · 0 repositories · arXiv:2503.16883
-
Autonomous Radiotherapy Treatment Planning Using DOLA: A Privacy-Preserving, LLM-Based Optimization Agent 21 Mar 2025 · 0 repositories · arXiv:2503.17553
-
CoKe: Customizable Fine-Grained Story Evaluation via Chain-of-Keyword Rationalization 21 Mar 2025 · 0 repositories · arXiv:2503.17136
-
Feature-Based Dual Visual Feature Extraction Model for Compound Multimodal Emotion Recognition 21 Mar 2025 · 1 repository · arXiv:2503.17453
-
Federated Cross-Domain Click-Through Rate Prediction With Large Language Model Augmentation 21 Mar 2025 · 0 repositories · arXiv:2503.16875
-
KL3M Tokenizers: A Family of Domain-Specific and Character-Level Tokenizers for Legal, Financial, and Preprocessing Applications 21 Mar 2025 · 1 repository · arXiv:2503.17247
-
SaudiCulture: A Benchmark for Evaluating Large Language Models Cultural Competence within Saudi Arabia 21 Mar 2025 · 0 repositories · arXiv:2503.17485