Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 2
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 2 of 190: papers 101 to 200 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
In-Context Learning for Gradient-Free Receiver Adaptation: Principles, Applications, and Theory 18 Jun 2025 · 0 repositories · arXiv:2506.15176
-
SignBart -- New approach with the skeleton sequence for Isolated Sign language Recognition 18 Jun 2025 · 1 repository · arXiv:2506.21592
-
Advances in Compliance Detection: Novel Models Using Vision-Based Tactile Sensors 17 Jun 2025 · 0 repositories · arXiv:2506.14980
-
AviationLLM: An LLM-based Knowledge System for Aviation Training 17 Jun 2025 · 0 repositories · arXiv:2506.14336
-
From Bytes to Ideas: Language Modeling with Autoregressive U-Nets 17 Jun 2025 · 1 repository · arXiv:2506.14761Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 16 harvested samples) · 1 pointer-only (licence)
-
ImpliRet: Benchmarking the Implicit Fact Retrieval Challenge 17 Jun 2025 · 1 repository · arXiv:2506.14407Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
InsertRank: LLMs can reason over BM25 scores to Improve Listwise Reranking 17 Jun 2025 · 0 repositories · arXiv:2506.14086
-
Lightweight Relevance Grader in RAG 17 Jun 2025 · 1 repository · arXiv:2506.14084
-
M2BeamLLM: Multimodal Sensing-empowered mmWave Beam Prediction with Large Language Models 17 Jun 2025 · 0 repositories · arXiv:2506.14532
-
PoseGRAF: Geometric-Reinforced Adaptive Fusion for Monocular 3D Human Pose Estimation 17 Jun 2025 · 1 repository · arXiv:2506.14596
-
RAGtifier: Evaluating RAG Generation Approaches of State-of-the-Art RAG Systems for the SIGIR LiveRAG Competition 17 Jun 2025 · 0 repositories · arXiv:2506.14412
-
Sampling from Your Language Model One Byte at a Time 17 Jun 2025 · 1 repository · arXiv:2506.14123Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
Scaling Intelligence: Designing Data Centers for Next-Gen Language Models 17 Jun 2025 · 0 repositories · arXiv:2506.15006
-
Toward a Graph Foundation Model: Pre-Training Transformers With Random Walks 17 Jun 2025 · 0 repositories · arXiv:2506.14098
-
A Gravity-informed Spatiotemporal Transformer for Human Activity Intensity Prediction 16 Jun 2025 · 0 repositories · arXiv:2506.13678
-
Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences 16 Jun 2025 · 2 repositories · arXiv:2506.13996
-
FOAM: A General Frequency-Optimized Anti-Overlapping Framework for Overlapping Object Perception 16 Jun 2025 · 0 repositories · arXiv:2506.13501
-
HierVL: Semi-Supervised Segmentation leveraging Hierarchical Vision-Language Synergy with Dynamic Text-Spatial Query Alignment 16 Jun 2025 · 0 repositories · arXiv:2506.13925
-
LTRR: Learning To Rank Retrievers for LLMs 16 Jun 2025 · 1 repository · arXiv:2506.13743
-
MambaMia: A State-Space-Model-Based Compression for Efficient Video Understanding in Large Multimodal Models 16 Jun 2025 · 0 repositories · arXiv:2506.13564
-
MT-PCR: A Hybrid Mamba-Transformer with Spatial Serialization for Hierarchical Point Cloud Registration 16 Jun 2025 · 0 repositories · arXiv:2506.13183
-
Overcoming Occlusions in the Wild: A Multi-Task Age Head Approach to Age Estimation 16 Jun 2025 · 0 repositories · arXiv:2506.13445
-
PF-LHM: 3D Animatable Avatar Reconstruction from Pose-free Articulated Human Images 16 Jun 2025 · 0 repositories · arXiv:2506.13766
-
Tree-Based Text Retrieval via Hierarchical Clustering in RAGFrameworks: Application on Taiwanese Regulations 16 Jun 2025 · 1 repository · arXiv:2506.13607
-
Alphabet Index Mapping: Jailbreaking LLMs through Semantic Dissimilarity 15 Jun 2025 · 0 repositories · arXiv:2506.12685
-
Cross-architecture universal feature coding via distribution alignment 15 Jun 2025 · 0 repositories · arXiv:2506.12737
-
Evaluating Cell Type Inference in Vision Language Models Under Varying Visual Context 15 Jun 2025 · 1 repository · arXiv:2506.12683
-
Frequency Dynamic Convolutions for Sound Event Detection 15 Jun 2025 · 0 repositories · arXiv:2506.12785
-
GM-LDM: Latent Diffusion Model for Brain Biomarker Identification through Functional Data-Driven Gray Matter Synthesis 15 Jun 2025 · 0 repositories · arXiv:2506.12719
-
Transforming Chatbot Text: A Sequence-to-Sequence Approach 15 Jun 2025 · 0 repositories · arXiv:2506.12843
-
BSA: Ball Sparse Attention for Large-scale Geometries 14 Jun 2025 · 1 repository · arXiv:2506.12541Syntology official (archive's flag): 2 ran · 2 ran (of which 1 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 1 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
FlexRAG: A Flexible and Comprehensive Framework for Retrieval-Augmented Generation 14 Jun 2025 · 1 repository · arXiv:2506.12494
-
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling 14 Jun 2025 · 0 repositories · arXiv:2506.12543
-
Language Models Enable Data-Augmented Synthesis Planning for Inorganic Materials 14 Jun 2025 · 0 repositories · arXiv:2506.12557
-
Parkinson's Disease Freezing of Gait (FoG) Symptom Detection Using Machine Learning from Wearable Sensor Data 14 Jun 2025 · 0 repositories · arXiv:2506.12561
-
Structural feature enhanced transformer for fine-grained image recognition 14 Jun 2025 · 0 repositories
-
Bias Amplification in RAG: Poisoning Knowledge Retrieval to Steer LLMs 13 Jun 2025 · 0 repositories · arXiv:2506.11415
-
Bubble Dynamics Transformer: Microrheology at Ultra-High Strain Rates 13 Jun 2025 · 0 repositories · arXiv:2506.11936
-
Chunk Twice, Embed Once: A Systematic Study of Segmentation and Representation Trade-offs in Chemistry-Aware Retrieval-Augmented Generation 13 Jun 2025 · 0 repositories · arXiv:2506.17277
-
Dr. GPT Will See You Now, but Should It? Exploring the Benefits and Harms of Large Language Models in Medical Diagnosis using Crowdsourced Clinical Cases 13 Jun 2025 · 0 repositories · arXiv:2506.13805
-
GraphGSOcc: Semantic-Geometric Graph Transformer with Dynamic-Static Decoupling for 3D Gaussian Splatting-based Occupancy Prediction 13 Jun 2025 · 0 repositories · arXiv:2506.14825
-
Identifying Helpful Context for LLM-based Vulnerability Repair: A Preliminary Study 13 Jun 2025 · 0 repositories · arXiv:2506.11561
-
Leveraging GPT-4 for Vulnerability-Witnessing Unit Test Generation 13 Jun 2025 · 0 repositories · arXiv:2506.11559
-
Voxel-Level Brain States Prediction Using Swin Transformer 13 Jun 2025 · 0 repositories · arXiv:2506.11455
-
"Check My Work?": Measuring Sycophancy in a Simulated Educational Context 12 Jun 2025 · 1 repository · arXiv:2506.10297
-
Augmenting Large Language Models with Static Code Analysis for Automated Code Quality Improvements 12 Jun 2025 · 0 repositories · arXiv:2506.10330
-
LogiPlan: A Structured Benchmark for Logical Planning and Relational Reasoning in LLMs 12 Jun 2025 · 0 repositories · arXiv:2506.10527
-
Med-URWKV: Pure RWKV With ImageNet Pre-training For Medical Image Segmentation 12 Jun 2025 · 0 repositories · arXiv:2506.10858
-
CIIR@LiveRAG 2025: Optimizing Multi-Agent Retrieval Augmented Generation through Self-Training 12 Jun 2025 · 1 repository · arXiv:2506.10844
-
Constructing and Evaluating Declarative RAG Pipelines in PyTerrier 12 Jun 2025 · 1 repository · arXiv:2506.10802
-
DART: Differentiable Dynamic Adaptive Region Tokenizer for Vision Transformer and Mamba 12 Jun 2025 · 1 repository · arXiv:2506.10390
-
Decomposing MLP Activations into Interpretable Features via Semi-Nonnegative Matrix Factorization 12 Jun 2025 · 1 repository · arXiv:2506.10920
-
Don't Pay Attention 12 Jun 2025 · 0 repositories · arXiv:2506.11305
-
DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers 12 Jun 2025 · 0 repositories · arXiv:2506.10568
-
Fine-Grained Perturbation Guidance via Attention Head Selection 12 Jun 2025 · 0 repositories · arXiv:2506.10978
-
FSATFusion: Frequency-Spatial Attention Transformer for Infrared and Visible Image Fusion 12 Jun 2025 · 1 repository · arXiv:2506.10366
-
HyBiomass: Global Hyperspectral Imagery Benchmark Dataset for Evaluating Geospatial Foundation Models in Forest Aboveground Biomass Estimation 12 Jun 2025 · 0 repositories · arXiv:2506.11314
-
MF2Summ: Multimodal Fusion for Video Summarization with Temporal Alignment 12 Jun 2025 · 0 repositories · arXiv:2506.10430
-
NeuralNexus at BEA 2025 Shared Task: Retrieval-Augmented Prompting for Mistake Identification in AI Tutors 12 Jun 2025 · 1 repository · arXiv:2506.10627
-
Reasoning RAG via System 1 or System 2: A Survey on Reasoning Agentic Retrieval-Augmented Generation for Industry Challenges 12 Jun 2025 · 0 repositories · arXiv:2506.10408
-
Spelling-out is not Straightforward: LLMs' Capability of Tokenization from Token to Characters 12 Jun 2025 · 0 repositories · arXiv:2506.10641
-
SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks 12 Jun 2025 · 1 repository · arXiv:2506.10954
-
TableRAG: A Retrieval Augmented Generation Framework for Heterogeneous Document Reasoning 12 Jun 2025 · 1 repository · arXiv:2506.10380
-
Think before You Simulate: Symbolic Reasoning to Orchestrate Neural Computation for Counterfactual Question Answering 12 Jun 2025 · 1 repository · arXiv:2506.10753
-
Towards Robust Multimodal Emotion Recognition under Missing Modalities and Distribution Shifts 12 Jun 2025 · 1 repository · arXiv:2506.10452
-
Retrieval of Surface Solar Radiation through Implicit Albedo Recovery from Temporal Context 11 Jun 2025 · 1 repository · arXiv:2506.10174
-
A Novel Lightweight Transformer with Edge-Aware Fusion for Remote Sensing Image Captioning 11 Jun 2025 · 0 repositories · arXiv:2506.09429
-
Analysis of Anonymous User Interaction Relationships and Prediction of Advertising Feedback Based on Graph Neural Network 11 Jun 2025 · 0 repositories · arXiv:2506.13787
-
Can LLMs Generate Good Stories? Insights and Challenges from a Narrative Planning Perspective 11 Jun 2025 · 0 repositories · arXiv:2506.10161
-
Large Language Models for Toxic Language Detection in Low-Resource Balkan Languages 11 Jun 2025 · 1 repository · arXiv:2506.09992
-
Latent Multi-Head Attention for Small Language Models 11 Jun 2025 · 0 repositories · arXiv:2506.09342
-
Learning Efficient and Generalizable Graph Retriever for Knowledge-Graph Question Answering 11 Jun 2025 · 1 repository · arXiv:2506.09645
-
Marrying Autoregressive Transformer and Diffusion with Multi-Reference Autoregression 11 Jun 2025 · 1 repository · arXiv:2506.09482
-
Mutual-Supervised Learning for Sequential-to-Parallel Code Translation 11 Jun 2025 · 1 repository · arXiv:2506.11153
-
On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention 11 Jun 2025 · 1 repository · arXiv:2506.09316
-
SparseSSM: Efficient Selective Structured State Space Models Can Be Pruned in One-Shot 11 Jun 2025 · 0 repositories · arXiv:2506.09613
-
TransXSSM: A Hybrid Transformer State Space Model with Unified Rotary Position Embedding 11 Jun 2025 · 0 repositories · arXiv:2506.09507
-
H²GFM: Towards unifying Homogeneity and Heterogeneity on Text-Attributed Graphs 10 Jun 2025 · 0 repositories · arXiv:2506.08298
-
MD-ViSCo: A Unified Model for Multi-Directional Vital Sign Waveform Conversion 10 Jun 2025 · 1 repository · arXiv:2506.08357
-
CALT: A Library for Computer Algebra with Transformer 10 Jun 2025 · 1 repository · arXiv:2506.08600
-
POLARON: Precision-aware On-device Learning and Adaptive Runtime-cONfigurable AI acceleration 10 Jun 2025 · 0 repositories · arXiv:2506.08785
-
Safeguarding Multimodal Knowledge Copyright in the RAG-as-a-Service Environment 10 Jun 2025 · 1 repository · arXiv:2506.10030
-
AraReasoner: Evaluating Reasoning-Based LLMs for Arabic NLP 10 Jun 2025 · 0 repositories · arXiv:2506.08768
-
CC-RAG: Structured Multi-Hop Reasoning via Theme-Based Causal Graphs 10 Jun 2025 · 0 repositories · arXiv:2506.08364
-
CounselBench: A Large-Scale Expert Evaluation and Adversarial Benchmark of Large Language Models in Mental Health Counseling 10 Jun 2025 · 1 repository · arXiv:2506.08584
-
DRAGged into Conflicts: Detecting and Addressing Conflicting Sources in Search-Augmented LLMs 10 Jun 2025 · 1 repository · arXiv:2506.08500
-
ECMNet:Lightweight Semantic Segmentation with Efficient CNN-Mamba Network 10 Jun 2025 · 0 repositories · arXiv:2506.08629
-
Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem Solving 10 Jun 2025 · 1 repository · arXiv:2506.08349Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
FaithfulRAG: Fact-Level Conflict Modeling for Context-Faithful Retrieval-Augmented Generation 10 Jun 2025 · 1 repository · arXiv:2506.08938
-
FedRAG: A Framework for Fine-Tuning Retrieval-Augmented Generation Systems 10 Jun 2025 · 1 repository · arXiv:2506.09200
-
FloorplanMAE:A self-supervised framework for complete floorplan generation from partial inputs 10 Jun 2025 · 0 repositories · arXiv:2506.08363
-
Genetic Transformer-Assisted Quantum Neural Networks for Optimal Circuit Design 10 Jun 2025 · 0 repositories · arXiv:2506.09205
-
HGFormer: A Hierarchical Graph Transformer Framework for Two-Stage Colonel Blotto Games via Reinforcement Learning 10 Jun 2025 · 0 repositories · arXiv:2506.08580
-
Hierarchical Neural Collapse Detection Transformer for Class Incremental Object Detection 10 Jun 2025 · 0 repositories · arXiv:2506.08562
-
Hyperspectral Image Classification via Transformer-based Spectral-Spatial Attention Decoupling and Adaptive Gating 10 Jun 2025 · 1 repository · arXiv:2506.08324
-
JoFormer (Journey-based Transformer): Theory and Empirical Analysis on the Tiny Shakespeare Dataset 10 Jun 2025 · 1 repository · arXiv:2506.08652
-
MedMoE: Modality-Specialized Mixture of Experts for Medical Vision-Language Understanding 10 Jun 2025 · 0 repositories · arXiv:2506.08356
-
PatchGuard: Adversarially Robust Anomaly Detection and Localization through Vision Transformers and Pseudo Anomalies 10 Jun 2025 · 2 repositories · arXiv:2506.09237
-
Robust Visual Localization via Semantic-Guided Multi-Scale Transformer 10 Jun 2025 · 0 repositories · arXiv:2506.08526
-
TACTIC: Translation Agents with Cognitive-Theoretic Interactive Collaboration 10 Jun 2025 · 1 repository · arXiv:2506.08403