Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 62
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 62 of 190: papers 6,101 to 6,200 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Are PPO-ed Language Models Hackable? 28 May 2024 · 0 repositories · arXiv:2406.02577
-
ATM: Adversarial Tuning Multi-agent System Makes a Robust Retrieval-Augmented Generator 28 May 2024 · 1 repository · arXiv:2405.18111Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Benchmarks Underestimate the Readiness of Multi-lingual Dialogue Agents 28 May 2024 · 0 repositories · arXiv:2405.17840
-
Delving into Differentially Private Transformer 28 May 2024 · 0 repositories · arXiv:2405.18194
-
Don't Forget to Connect! Improving RAG with Graph-based Reranking 28 May 2024 · 0 repositories · arXiv:2405.18414
-
Dual-Path Multi-Scale Transformer for High-Quality Image Deraining 28 May 2024 · 0 repositories · arXiv:2405.18124
-
Edinburgh Clinical NLP at MEDIQA-CORR 2024: Guiding Large Language Models with Hints 28 May 2024 · 0 repositories · arXiv:2405.18028
-
FASTopic: Pretrained Transformer is a Fast, Adaptive, Stable, and Transferable Topic Model 28 May 2024 · 2 repositories · arXiv:2405.17978Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
ForecastGrapher: Redefining Multivariate Time Series Forecasting with Graph Neural Networks 28 May 2024 · 0 repositories · arXiv:2405.18036
-
HarmoDT: Harmony Multi-Task Decision Transformer for Offline Reinforcement Learning 28 May 2024 · 1 repository · arXiv:2405.18080
-
IAPT: Instruction-Aware Prompt Tuning for Large Language Models 28 May 2024 · 0 repositories · arXiv:2405.18203
-
LDMol: Text-to-Molecule Diffusion Model with Structurally Informative Latent Space 28 May 2024 · 1 repository · arXiv:2405.17829Syntology official (archive's flag): 5 ran · 6 ran (of which 2 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples) · 1 pointer-only (licence)
-
LLMs and Memorization: On Quality and Specificity of Copyright Compliance 28 May 2024 · 1 repository · arXiv:2405.18492
-
Modeling Long Sequences in Bladder Cancer Recurrence: A Comparative Evaluation of LSTM,Transformer,and Mamba 28 May 2024 · 0 repositories · arXiv:2405.18518
-
MindFormer: Semantic Alignment of Multi-Subject fMRI for Brain Decoding 28 May 2024 · 0 repositories · arXiv:2405.17720
-
Multi-objective Representation for Numbers in Clinical Narratives: A CamemBERT-Bio-Based Alternative to Large-Scale LLMs 28 May 2024 · 0 repositories · arXiv:2405.18448
-
Notes on Applicability of GPT-4 to Document Understanding 28 May 2024 · 0 repositories · arXiv:2405.18433
-
ORLM: A Customizable Framework in Training Large Models for Automated Optimization Modeling 28 May 2024 · 1 repository · arXiv:2405.17743Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples)
-
OV-DQUO: Open-Vocabulary DETR with Denoising Text Query Training and Open-World Unknown Objects Supervision 28 May 2024 · 1 repository · arXiv:2405.17913Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 1 honoured, 0 violated, 6 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples) · 2 pointer-only (licence)
-
Peering into the Mind of Language Models: An Approach for Attribution in Contextual Question Answering 28 May 2024 · 1 repository · arXiv:2405.17980
-
Proof of Quality: A Costless Paradigm for Trustless Generative AI Model Inference on Blockchains 28 May 2024 · 0 repositories · arXiv:2405.17934
-
RealitySummary: Exploring On-Demand Mixed Reality Text Summarization and Question Answering using Large Language Models 28 May 2024 · 0 repositories · arXiv:2405.18620
-
Thai Winograd Schemas: A Benchmark for Thai Commonsense Reasoning 28 May 2024 · 1 repository · arXiv:2405.18375
-
The Battle of LLMs: A Comparative Study in Conversational QA Tasks 28 May 2024 · 0 repositories · arXiv:2405.18344
-
Understanding Intrinsic Socioeconomic Biases in Large Language Models 28 May 2024 · 0 repositories · arXiv:2405.18662
-
ViG: Linear-complexity Visual Sequence Learning with Gated Linear Attention 28 May 2024 · 1 repository · arXiv:2405.18425
-
Visual Anchors Are Strong Information Aggregators For Multimodal Large Language Model 28 May 2024 · 1 repository · arXiv:2405.17815
-
VITON-DiT: Learning In-the-Wild Video Try-On from Human Dance Videos via Diffusion Transformers 28 May 2024 · 0 repositories · arXiv:2405.18326
-
Wavelet-Based Image Tokenizer for Vision Transformers 28 May 2024 · 0 repositories · arXiv:2405.18616
-
A One-Layer Decoder-Only Transformer is a Two-Layer RNN: With an Application to Certified Robustness 27 May 2024 · 0 repositories · arXiv:2405.17361
-
Advanced Language Model-based Translator for English-Vietnamese Translation 27 May 2024 · 1 repository
-
Are Self-Attentions Effective for Time Series Forecasting? 27 May 2024 · 1 repository · arXiv:2405.16877Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 8 unverified (of 13 harvested samples) · 3 pointer-only (licence)
-
Assessing LLMs Suitability for Knowledge Graph Completion 27 May 2024 · 1 repository · arXiv:2405.17249
-
Augmenting Textual Generation via Topology Aware Retrieval 27 May 2024 · 0 repositories · arXiv:2405.17602
-
Autoformalizing Euclidean Geometry 27 May 2024 · 1 repository · arXiv:2405.17216Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 2 violated, 1 with no contract checked; 5 where Syntology's instrument failed) · 4 unverified (of 13 harvested samples)
-
Automatic Domain Adaptation by Transformers in In-Context Learning 27 May 2024 · 0 repositories · arXiv:2405.16819
-
BehaviorGPT: Smart Agent Simulation for Autonomous Driving with Next-Patch Prediction 27 May 2024 · 0 repositories · arXiv:2405.17372
-
CHESS: Contextual Harnessing for Efficient SQL Synthesis 27 May 2024 · 2 repositories · arXiv:2405.16755Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
Cost-efficient Knowledge-based Question Answering with Large Language Models 27 May 2024 · 0 repositories · arXiv:2405.17337
-
Deciphering Movement: Unified Trajectory Generation Model for Multi-Agent 27 May 2024 · 1 repository · arXiv:2405.17680Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples)
-
DeeperImpact: Optimizing Sparse Learned Index Structures 27 May 2024 · 1 repository · arXiv:2405.17093
-
Exploiting the Layered Intrinsic Dimensionality of Deep Models for Practical Adversarial Training 27 May 2024 · 0 repositories · arXiv:2405.17130
-
Galaxy: A Resource-Efficient Collaborative Edge AI System for In-situ Transformer Inference 27 May 2024 · 0 repositories · arXiv:2405.17245
-
Interesting Scientific Idea Generation using Knowledge Graphs and LLMs: Evaluations with 100 Research Group Leaders 27 May 2024 · 1 repository · arXiv:2405.17044
-
How Do the Architecture and Optimizer Affect Representation Learning? On the Training Dynamics of Representations in Deep Neural Networks 27 May 2024 · 0 repositories · arXiv:2405.17377
-
InversionView: A General-Purpose Method for Reading Information from Neural Activations 27 May 2024 · 1 repository · arXiv:2405.17653Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified; the one sample that ran constructed an object rather than computing a result (of 6 harvested samples) · 6 pointer-only (licence)
-
LCM: Locally Constrained Compact Point Cloud Model for Masked Point Modeling 27 May 2024 · 1 repository · arXiv:2405.17149Syntology official (archive's flag): 9 ran · 12 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 3 honoured, 0 violated, 2 with no contract checked; 7 where Syntology's instrument failed) · 0 unverified (of 12 harvested samples) · 10 pointer-only (licence)
-
REVECA: Adaptive Planning and Trajectory-based Validation in Cooperative Language Agents using Information Relevance and Relative Proximity 27 May 2024 · 0 repositories · arXiv:2405.16751
-
LoReTrack: Efficient and Accurate Low-Resolution Transformer Tracking 27 May 2024 · 1 repository · arXiv:2405.17660
-
Masked Face Recognition with Generative-to-Discriminative Representations 27 May 2024 · 0 repositories · arXiv:2405.16761
-
Motion-Agent: A Conversational Framework for Human Motion Generation with LLMs 27 May 2024 · 1 repository · arXiv:2405.17013Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
Novel Approaches for ML-Assisted Particle Track Reconstruction and Hit Clustering 27 May 2024 · 0 repositories · arXiv:2405.17325
-
Performance evaluation of Reddit Comments using Machine Learning and Natural Language Processing methods in Sentiment Analysis 27 May 2024 · 0 repositories · arXiv:2405.16810
-
PivotMesh: Generic 3D Mesh Generation via Pivot Vertices Guidance 27 May 2024 · 0 repositories · arXiv:2405.16890
-
Q-value Regularized Transformer for Offline Reinforcement Learning 27 May 2024 · 2 repositories · arXiv:2405.17098Syntology 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples) · 3 pointer-only (licence)
-
QUB-Cirdan at "Discharge Me!": Zero shot discharge letter generation by open-source LLM 27 May 2024 · 0 repositories · arXiv:2406.00041
-
ReflectionCoder: Learning from Reflection Sequence for Enhanced One-off Code Generation 27 May 2024 · 1 repository · arXiv:2405.17057Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Rethinking Transformers in Solving POMDPs 27 May 2024 · 1 repository · arXiv:2405.17358Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
RTL-Repo: A Benchmark for Evaluating LLMs on Large-Scale RTL Design Projects 27 May 2024 · 1 repository · arXiv:2405.17378
-
Safe LoRA: the Silver Lining of Reducing Safety Risks when Fine-tuning Large Language Models 27 May 2024 · 1 repository · arXiv:2405.16833Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Supervised Batch Normalization 27 May 2024 · 0 repositories · arXiv:2405.17027
-
SWAT: Scalable and Efficient Window Attention-based Transformers Acceleration on FPGAs 27 May 2024 · 0 repositories · arXiv:2405.17025
-
The Scaling Law in Stellar Light Curves 27 May 2024 · 0 repositories · arXiv:2405.17156
-
THREAD: Thinking Deeper with Recursive Spawning 27 May 2024 · 1 repository · arXiv:2405.17402Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
On Understanding Attention-Based In-Context Learning for Categorical Data 27 May 2024 · 0 repositories · arXiv:2405.17248
-
UIT-DarkCow team at ImageCLEFmedical Caption 2024: Diagnostic Captioning for Radiology Images Efficiency with Transformer Models 27 May 2024 · 0 repositories · arXiv:2405.17002
-
Unisolver: PDE-Conditional Transformers Are Universal PDE Solvers 27 May 2024 · 1 repository · arXiv:2405.17527Syntology 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 1 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 1 pointer-only (licence)
-
Video Enriched Retrieval Augmented Generation Using Aligned Video Captions 27 May 2024 · 1 repository · arXiv:2405.17706
-
Vision-and-Language Navigation Generative Pretrained Transformer 27 May 2024 · 0 repositories · arXiv:2405.16994
-
Explaining Modern Gated-Linear RNNs via a Unified Implicit Attention Formulation 26 May 2024 · 1 repository · arXiv:2405.16504
-
DarijaBanking: A New Resource for Overcoming Language Barriers in Banking Intent Detection for Moroccan Arabic Speakers 26 May 2024 · 1 repository · arXiv:2405.16482
-
Demystify Mamba in Vision: A Linear Attention Perspective 26 May 2024 · 1 repository · arXiv:2405.16605Syntology official (archive's flag): 6 ran · 6 ran (of which 6 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified; every one of the 6 samples that ran constructed an object rather than computing a result (of 8 harvested samples)
-
Disentangling and Integrating Relational and Sensory Information in Transformer Architectures 26 May 2024 · 2 repositories · arXiv:2405.16727Syntology official (archive's flag): 2 ran · 14 ran (of which 9 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 1 violated, 9 with no contract checked; 4 where Syntology's instrument failed) · 7 unverified (of 21 harvested samples)
-
GRAG: Graph Retrieval-Augmented Generation 26 May 2024 · 1 repository · arXiv:2405.16506
-
M-RAG: Reinforcing Large Language Model Performance through Retrieval-Augmented Generation with Multiple Partitions 26 May 2024 · 0 repositories · arXiv:2405.16420
-
MambaTS: Improved Selective State Space Models for Long-term Time Series Forecasting 26 May 2024 · 1 repository · arXiv:2405.16440Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples) · 2 pointer-only (licence)
-
Planning with Multi-Constraints via Collaborative Language Agents 26 May 2024 · 1 repository · arXiv:2405.16510
-
Scalable Numerical Embeddings for Multivariate Time Series: Enhancing Healthcare Data Representation Learning 26 May 2024 · 0 repositories · arXiv:2405.16557
-
SpinQuant: LLM quantization with learned rotations 26 May 2024 · 3 repositories · arXiv:2405.16406Syntology 10 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Accelerating Inference of Retrieval-Augmented Generation via Sparse Context Selection 25 May 2024 · 0 repositories · arXiv:2405.16178
-
Accelerating Transformers with Spectrum-Preserving Token Merging 25 May 2024 · 1 repository · arXiv:2405.16148Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
AutoManual: Constructing Instruction Manuals by LLM Agents via Interactive Environmental Learning 25 May 2024 · 1 repository · arXiv:2405.16247Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Comparative Analysis of Open-Source Language Models in Summarizing Medical Text Data 25 May 2024 · 0 repositories · arXiv:2405.16295
-
Confidence Under the Hood: An Investigation into the Confidence-Probability Alignment in Large Language Models 25 May 2024 · 1 repository · arXiv:2405.16282Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified; the one sample that ran constructed an object rather than computing a result (of 6 harvested samples)
-
Dynamic Inhomogeneous Quantum Resource Scheduling with Reinforcement Learning 25 May 2024 · 0 repositories · arXiv:2405.16380
-
Efficient Temporal Action Segmentation via Boundary-aware Query Voting 25 May 2024 · 1 repository · arXiv:2405.15995
-
Evaluating deep learning methods applied to Landsat time series subsequences to detect and classify boreal forest disturbances events: The challenge of partial and progressive disturbances 25 May 2024 · 1 repository
-
GeneAgent: Self-verification Language Agent for Gene Set Knowledge Discovery using Domain Databases 25 May 2024 · 0 repositories · arXiv:2405.16205
-
HETHUB: A Distributed Training System with Heterogeneous Cluster for Large-Scale Models 25 May 2024 · 0 repositories · arXiv:2405.16256
-
Incremental Comprehension of Garden-Path Sentences by Large Language Models: Semantic Interpretation, Syntactic Re-Analysis, and Attention 25 May 2024 · 0 repositories · arXiv:2405.16042
-
Lateralization MLP: A Simple Brain-inspired Architecture for Diffusion 25 May 2024 · 1 repository · arXiv:2405.16098
-
MindStar: Enhancing Math Reasoning in Pre-trained LLMs at Inference Time 25 May 2024 · 0 repositories · arXiv:2405.16265
-
MoEUT: Mixture-of-Experts Universal Transformers 25 May 2024 · 1 repository · arXiv:2405.16039Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples)
-
Picturing Ambiguity: A Visual Twist on the Winograd Schema Challenge 25 May 2024 · 0 repositories · arXiv:2405.16277
-
STRIDE: A Tool-Assisted LLM Agent Framework for Strategic and Interactive Decision-Making 25 May 2024 · 1 repository · arXiv:2405.16376Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Towards Black-Box Membership Inference Attack for Diffusion Models 25 May 2024 · 0 repositories · arXiv:2405.20771
-
Towards Unlocking Insights from Logbooks Using AI 25 May 2024 · 0 repositories · arXiv:2406.12881
-
Transformer Meets Gated Residual Networks To Enhance Photoplethysmogram Artifact Detection Informed by Mutual Information Neural Estimation 25 May 2024 · 0 repositories · arXiv:2405.16177
-
An Evaluation of Estimative Uncertainty in Large Language Models 24 May 2024 · 0 repositories · arXiv:2405.15185
-
Before Generation, Align it! A Novel and Effective Strategy for Mitigating Hallucinations in Text-to-SQL Generation 24 May 2024 · 1 repository · arXiv:2405.15307Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 1 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples) · 10 pointer-only (licence)