Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 15
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 15 of 190: papers 1,401 to 1,500 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
When Do Transformers Outperform Feedforward and Recurrent Networks? A Statistical Perspective 14 Mar 2025 · 1 repository · arXiv:2503.11272
-
A Frustratingly Simple Yet Highly Effective Attack Baseline: Over 90% Success Rate Against the Strong Black-box Models of GPT-4.5/4o/o1 13 Mar 2025 · 1 repository · arXiv:2503.10635Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
A Hybrid Architecture with Efficient Fine Tuning for Abstractive Patent Document Summarization 13 Mar 2025 · 0 repositories · arXiv:2503.10354
-
Advanced Tool Learning and Selection System (ATLASS): A Closed-Loop Framework Using LLM 13 Mar 2025 · 0 repositories · arXiv:2503.10071
-
ARLED: Leveraging LED-based ARMAN Model for Abstractive Summarization of Persian Long Documents 13 Mar 2025 · 0 repositories · arXiv:2503.10233
-
AttentionRAG: Attention-Guided Context Pruning in Retrieval-Augmented Generation 13 Mar 2025 · 0 repositories · arXiv:2503.10720
-
AudioX: Diffusion Transformer for Anything-to-Audio Generation 13 Mar 2025 · 0 repositories · arXiv:2503.10522
-
ChatGPT Encounters Morphing Attack Detection: Zero-Shot MAD with Multi-Modal Large Language Models and General Vision Models 13 Mar 2025 · 0 repositories · arXiv:2503.10937
-
CoCMT: Communication-Efficient Cross-Modal Transformer for Collaborative Perception 13 Mar 2025 · 0 repositories · arXiv:2503.13504
-
Compositional Subspace Representation Fine-tuning for Adaptive Large Language Models 13 Mar 2025 · 0 repositories · arXiv:2503.10617
-
Cosh-DiT: Co-Speech Gesture Video Synthesis via Hybrid Audio-Visual Diffusion Transformers 13 Mar 2025 · 0 repositories · arXiv:2503.09942
-
CountPath: Automating Fragment Counting in Digital Pathology 13 Mar 2025 · 0 repositories · arXiv:2503.10520
-
Do I look like a `cat.n.01` to you? A Taxonomy Image Generation Benchmark 13 Mar 2025 · 0 repositories · arXiv:2503.10357
-
Emotion Recognition with CLIP and Sequential Learning 13 Mar 2025 · 0 repositories · arXiv:2503.09929
-
FG-RAG: Enhancing Query-Focused Summarization with Context-Aware Fine-Grained Graph RAG 13 Mar 2025 · 1 repository · arXiv:2504.07103
-
Fixed-Point RNNs: From Diagonal to Dense in a Few Iterations 13 Mar 2025 · 0 repositories · arXiv:2503.10799
-
Gumiho: A Hybrid Architecture to Prioritize Early Tokens in Speculative Decoding 13 Mar 2025 · 0 repositories · arXiv:2503.10135Syntology 12 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 4 where Syntology's instrument failed) · 3 unverified (of 15 harvested samples) · 8 pointer-only (licence)
-
It is Too Many Options: Pitfalls of Multiple-Choice Questions in Generative AI and Medical Education 13 Mar 2025 · 0 repositories · arXiv:2503.13508
-
KV-Distill: Nearly Lossless Learnable Context Compression for LLMs 13 Mar 2025 · 0 repositories · arXiv:2503.10337
-
Radar: Fast Long-Context Decoding for Any Transformer 13 Mar 2025 · 1 repository · arXiv:2503.10571Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Retrieval-Augmented Generation with Hierarchical Knowledge 13 Mar 2025 · 1 repository · arXiv:2503.10150Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples)
-
Robustness Tokens: Towards Adversarial Robustness of Transformers 13 Mar 2025 · 1 repository · arXiv:2503.10191
-
Tempest: Autonomous Multi-Turn Jailbreaking of Large Language Models with Tree Search 13 Mar 2025 · 0 repositories · arXiv:2503.10619
-
TacticExpert: Spatial-Temporal Graph Language Model for Basketball Tactics 13 Mar 2025 · 0 repositories · arXiv:2503.10722
-
Taxonomic Reasoning for Rare Arthropods: Combining Dense Image Captioning and RAG for Interpretable Classification 13 Mar 2025 · 0 repositories · arXiv:2503.10886
-
TGP: Two-modal occupancy prediction with 3D Gaussian and sparse points for 3D Environment Awareness 13 Mar 2025 · 0 repositories · arXiv:2503.09941
-
Towards Efficient Large Scale Spatial-Temporal Time Series Forecasting via Improved Inverted Transformers 13 Mar 2025 · 0 repositories · arXiv:2503.10858
-
Why Does Your CoT Prompt (Not) Work? Theoretical Analysis of Prompt Space Complexity, its Interaction with Answer Space During CoT Reasoning with LLMs: A Recurrent Perspective 13 Mar 2025 · 0 repositories · arXiv:2503.10084
-
4D-ACFNet: A 4D Attention Mechanism-Based Prognostic Framework for Colorectal Cancer Liver Metastasis Integrating Multimodal Spatiotemporal Features 12 Mar 2025 · 0 repositories · arXiv:2503.09652
-
AgentDAM: Privacy Leakage Evaluation for Autonomous Web Agents 12 Mar 2025 · 1 repository · arXiv:2503.09780Syntology official: no sample here; runs from other or unrecorded repositories · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
An Evaluation of LLMs for Detecting Harmful Computing Terms 12 Mar 2025 · 0 repositories · arXiv:2503.09341
-
ClaimTrust: Propagation Trust Scoring for RAG Systems 12 Mar 2025 · 0 repositories · arXiv:2503.10702
-
Conversational Gold: Evaluating Personalized Conversational Search System using Gold Nuggets 12 Mar 2025 · 1 repository · arXiv:2503.09902
-
Discovering Influential Neuron Path in Vision Transformers 12 Mar 2025 · 0 repositories · arXiv:2503.09046
-
Finding the Muses: Identifying Coresets through Loss Trajectories 12 Mar 2025 · 0 repositories · arXiv:2503.09721
-
How to Protect Yourself from 5G Radiation? Investigating LLM Responses to Implicit Misinformation 12 Mar 2025 · 1 repository · arXiv:2503.09598
-
Memory-enhanced Retrieval Augmentation for Long Video Understanding 12 Mar 2025 · 0 repositories · arXiv:2503.09149
-
Minimal Time Series Transformer 12 Mar 2025 · 1 repository · arXiv:2503.09791
-
MoC: Mixtures of Text Chunking Learners for Retrieval-Augmented Generation System 12 Mar 2025 · 1 repository · arXiv:2503.09600
-
Language-Enhanced Representation Learning for Single-Cell Transcriptomics 12 Mar 2025 · 1 repository · arXiv:2503.09427Syntology official: no sample here; runs from other or unrecorded repositories · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
NAMI: Efficient Image Generation via Progressive Rectified Flow Transformers 12 Mar 2025 · 0 repositories · arXiv:2503.09242
-
Other Vehicle Trajectories Are Also Needed: A Driving World Model Unifies Ego-Other Vehicle Trajectories in Video Latant Space 12 Mar 2025 · 0 repositories · arXiv:2503.09215
-
Post-interactive Multimodal Trajectory Prediction for Autonomous Driving 12 Mar 2025 · 0 repositories · arXiv:2503.09366
-
Rethinking Prompt-based Debiasing in Large Language Models 12 Mar 2025 · 0 repositories · arXiv:2503.09219
-
Robust Multimodal Survival Prediction with the Latent Differentiation Conditional Variational AutoEncoder 12 Mar 2025 · 1 repository · arXiv:2503.09496Syntology official (archive's flag): 5 ran · 6 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
SE(3)-Equivariant Robot Learning and Control: A Tutorial Survey 12 Mar 2025 · 0 repositories · arXiv:2503.09829
-
Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning 12 Mar 2025 · 3 repositories · arXiv:2503.09516Syntology official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 1 pointer-only (licence)
-
TA-V2A: Textually Assisted Video-to-Audio Generation 12 Mar 2025 · 0 repositories · arXiv:2503.10700
-
Un-Straightening Generative AI: How Queer Artists Surface and Challenge the Normativity of Generative AI Models 12 Mar 2025 · 0 repositories · arXiv:2503.09805
-
Unified Locomotion Transformer with Simultaneous Sim-to-Real Transfer for Quadrupeds 12 Mar 2025 · 0 repositories · arXiv:2503.08997
-
VaxGuard: A Multi-Generator, Multi-Type, and Multi-Role Dataset for Detecting LLM-Generated Vaccine Misinformation 12 Mar 2025 · 0 repositories · arXiv:2503.09103
-
VLog: Video-Language Models by Generative Retrieval of Narration Vocabulary 12 Mar 2025 · 1 repository · arXiv:2503.09402Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Who Are You Behind the Screen? Implicit MBTI and Gender Detection Using Artificial Intelligence 12 Mar 2025 · 0 repositories · arXiv:2503.09853
-
A Survey on Knowledge-Oriented Retrieval-Augmented Generation 11 Mar 2025 · 0 repositories · arXiv:2503.10677
-
Accurate INT8 Training Through Dynamic Block-Level Fallback 11 Mar 2025 · 0 repositories · arXiv:2503.08040
-
Context-aware Biases for Length Extrapolation 11 Mar 2025 · 1 repository · arXiv:2503.08067Syntology official: no sample here; runs from other or unrecorded repositories · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample) · 1 pointer-only (licence)
-
EFPC: Towards Efficient and Flexible Prompt Compression 11 Mar 2025 · 0 repositories · arXiv:2503.07956
-
External Knowledge Injection for CLIP-Based Class-Incremental Learning 11 Mar 2025 · 3 repositories · arXiv:2503.08510
-
GPT-PPG: A GPT-based Foundation Model for Photoplethysmography Signals 11 Mar 2025 · 0 repositories · arXiv:2503.08015
-
HOTFormerLoc: Hierarchical Octree Transformer for Versatile Lidar Place Recognition Across Ground and Aerial Views 11 Mar 2025 · 0 repositories · arXiv:2503.08140
-
Interpretable and Robust Dialogue State Tracking via Natural Language Summarization with LLMs 11 Mar 2025 · 0 repositories · arXiv:2503.08857
-
KAN-Mixers: a new deep learning architecture for image classification 11 Mar 2025 · 0 repositories · arXiv:2503.08939
-
LLM-based Corroborating and Refuting Evidence Retrieval for Scientific Claim Verification 11 Mar 2025 · 0 repositories · arXiv:2503.07937
-
Llms, Virtual Users, and Bias: Predicting Any Survey Question Without Human Data 11 Mar 2025 · 0 repositories · arXiv:2503.16498
-
OpenRAG: Optimizing RAG End-to-End via In-Context Retrieval Learning 11 Mar 2025 · 0 repositories · arXiv:2503.08398
-
QUIET-SR: Quantum Image Enhancement Transformer for Single Image Super-Resolution 11 Mar 2025 · 0 repositories · arXiv:2503.08759
-
Seeing What's Not There: Spurious Correlation in Multimodal LLMs 11 Mar 2025 · 0 repositories · arXiv:2503.08884
-
TransECG: Leveraging Transformers for Explainable ECG Re-identification Risk Analysis 11 Mar 2025 · 0 repositories · arXiv:2503.13495
-
A LongFormer-Based Framework for Accurate and Efficient Medical Text Summarization 10 Mar 2025 · 0 repositories · arXiv:2503.06888
-
A LSTM-Transformer Model for pulsation control of pVADs 10 Mar 2025 · 0 repositories · arXiv:2503.07110
-
Bot Wars Evolved: Orchestrating Competing LLMs in a Counterstrike Against Phone Scams 10 Mar 2025 · 0 repositories · arXiv:2503.07036
-
CtrlRAG: Black-box Adversarial Attacks Based on Masked Language Models in Retrieval-Augmented Language Generation 10 Mar 2025 · 0 repositories · arXiv:2503.06950
-
Enhancing Time Series Forecasting via Logic-Inspired Regularization 10 Mar 2025 · 0 repositories · arXiv:2503.06867
-
Exploring Multimodal Perception in Large Language Models Through Perceptual Strength Ratings 10 Mar 2025 · 0 repositories · arXiv:2503.06980
-
Exposure Bias Reduction for Enhancing Diffusion Transformer Feature Caching 10 Mar 2025 · 1 repository · arXiv:2503.07120
-
Filter Images First, Generate Instructions Later: Pre-Instruction Data Selection for Visual Instruction Tuning 10 Mar 2025 · 0 repositories · arXiv:2503.07591
-
From Idea to Implementation: Evaluating the Influence of Large Language Models in Software Development -- An Opinion Paper 10 Mar 2025 · 0 repositories · arXiv:2503.07450
-
Fully Autonomous Programming using Iterative Multi-Agent Debugging with Large Language Models 10 Mar 2025 · 0 repositories · arXiv:2503.07693
-
Implicit Reasoning in Transformers is Reasoning through Shortcuts 10 Mar 2025 · 1 repository · arXiv:2503.07604Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples) · 1 pointer-only (licence)
-
Improving cognitive diagnostics in pathology: a deep learning approach for augmenting perceptional understanding of histopathology images 10 Mar 2025 · 0 repositories · arXiv:2503.06894
-
Large model enhanced computational ghost imaging 10 Mar 2025 · 1 repository · arXiv:2503.08710
-
MambaFlow: A Mamba-Centric Architecture for End-to-End Optical Flow Estimation 10 Mar 2025 · 0 repositories · arXiv:2503.07046
-
MapQA: Open-domain Geospatial Question Answering on Map Data 10 Mar 2025 · 0 repositories · arXiv:2503.07871
-
Post-Training Quantization for Diffusion Transformer via Hierarchical Timestep Grouping 10 Mar 2025 · 0 repositories · arXiv:2503.06930
-
ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual Restoration 10 Mar 2025 · 1 repository · arXiv:2503.06881Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 3 where Syntology's instrument failed) · 5 unverified (of 15 harvested samples) · 1 pointer-only (licence)
-
Roamify: Designing and Evaluating an LLM Based Google Chrome Extension for Personalised Itinerary Planning 10 Mar 2025 · 1 repository · arXiv:2504.10489
-
Taking Notes Brings Focus? Towards Multi-Turn Multimodal Dialogue Learning 10 Mar 2025 · 0 repositories · arXiv:2503.07002
-
Talking to GDELT Through Knowledge Graphs 10 Mar 2025 · 0 repositories · arXiv:2503.07584
-
VACE: All-in-One Video Creation and Editing 10 Mar 2025 · 2 repositories · arXiv:2503.07598Syntology 11 ran (of which 5 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 1 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 8 unverified (of 19 harvested samples)
-
Beyond Decoder-only: Large Language Models Can be Good Encoders for Machine Translation 9 Mar 2025 · 1 repository · arXiv:2503.06594
-
Effectiveness of Zero-shot-CoT in Japanese Prompts 9 Mar 2025 · 0 repositories · arXiv:2503.06765
-
Global-Aware Monocular Semantic Scene Completion with State Space Models 9 Mar 2025 · 0 repositories · arXiv:2503.06569
-
GroMo: Plant Growth Modeling with Multiview Images 9 Mar 2025 · 1 repository · arXiv:2503.06608
-
Human Cognition Inspired RAG with Knowledge Graph for Complex Problem Solving 9 Mar 2025 · 0 repositories · arXiv:2503.06567
-
LSA: Latent Style Augmentation Towards Stain-Agnostic Cervical Cancer Screening 9 Mar 2025 · 0 repositories · arXiv:2503.06563
-
Seeing Delta Parameters as JPEG Images: Data-Free Delta Compression with Discrete Cosine Transform 9 Mar 2025 · 0 repositories · arXiv:2503.06676
-
SKG-LLM: Developing a Mathematical Model for Stroke Knowledge Graph Construction Using Large Language Models 9 Mar 2025 · 0 repositories · arXiv:2503.06475
-
UniGenX: Unified Generation of Sequence and Structure with Autoregressive Diffusion 9 Mar 2025 · 0 repositories · arXiv:2503.06687
-
A Noise-Robust Turn-Taking System for Real-World Dialogue Robots: A Field Experiment 8 Mar 2025 · 1 repository · arXiv:2503.06241
-
End-to-End Action Segmentation Transformer 8 Mar 2025 · 0 repositories · arXiv:2503.06316