Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 9
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 9 of 190: papers 801 to 900 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Retrieval-Augmented Generation in Biomedicine: A Survey of Technologies, Datasets, and Clinical Applications 2 May 2025 · 0 repositories · arXiv:2505.01146
-
Token-free Models for Sarcasm Detection 2 May 2025 · 0 repositories · arXiv:2505.01006
-
Zero-Shot Document-Level Biomedical Relation Extraction via Scenario-based Prompt Design in Two-Stage with LLM 2 May 2025 · 0 repositories · arXiv:2505.01077
-
DARTer: Dynamic Adaptive Representation Tracker for Nighttime UAV Tracking 1 May 2025 · 0 repositories · arXiv:2505.00752
-
A Time-Series Data Augmentation Model through Diffusion and Transformer Integration 1 May 2025 · 0 repositories · arXiv:2505.03790
-
Directly Forecasting Belief for Reinforcement Learning with Delays 1 May 2025 · 1 repository · arXiv:2505.00546
-
Efficient Recommendation with Millions of Items by Dynamic Pruning of Sub-Item Embeddings 1 May 2025 · 0 repositories · arXiv:2505.00560
-
Enhancing Tropical Cyclone Path Forecasting with an Improved Transformer Network 1 May 2025 · 0 repositories · arXiv:2505.00495
-
EnronQA: Towards Personalized RAG over Private Documents 1 May 2025 · 0 repositories · arXiv:2505.00263
-
Gateformer: Advancing Multivariate Time Series Forecasting through Temporal and Variate-Wise Attention with Gated Representations 1 May 2025 · 1 repository · arXiv:2505.00307Syntology official (archive's flag): 2 ran · 2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; every one of the 2 samples that ran constructed an object rather than computing a result (of 2 harvested samples) · 2 pointer-only (licence)
-
Open-Source LLM-Driven Federated Transformer for Predictive IoV Management 1 May 2025 · 0 repositories · arXiv:2505.00651
-
Patchwork: A Unified Framework for RAG Serving 1 May 2025 · 0 repositories · arXiv:2505.07833
-
Unlocking the Potential of Linear Networks for Irregular Multivariate Time Series Forecasting 1 May 2025 · 0 repositories · arXiv:2505.00590
-
Consistency-aware Fake Videos Detection on Short Video Platforms 30 Apr 2025 · 1 repository · arXiv:2504.21495
-
DOPE: Dual Object Perception-Enhancement Network for Vision-and-Language Navigation 30 Apr 2025 · 0 repositories · arXiv:2505.00743
-
Enhancing Security and Strengthening Defenses in Automated Short-Answer Grading Systems 30 Apr 2025 · 0 repositories · arXiv:2505.00061
-
LLM-Empowered Embodied Agent for Memory-Augmented Task Planning in Household Robotics 30 Apr 2025 · 1 repository · arXiv:2504.21716Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Talk Before You Retrieve: Agent-Led Discussions for Better RAG in Medical QA 30 Apr 2025 · 1 repository · arXiv:2504.21252
-
Traceback of Poisoning Attacks to Retrieval-Augmented Generation 30 Apr 2025 · 0 repositories · arXiv:2504.21668
-
Advance Fake Video Detection via Vision Transformers 29 Apr 2025 · 0 repositories · arXiv:2504.20669
-
AlignDiT: Multimodal Aligned Diffusion Transformer for Synchronized Speech Generation 29 Apr 2025 · 0 repositories · arXiv:2504.20629
-
ARCS: Agentic Retrieval-Augmented Code Synthesis with Iterative Refinement 29 Apr 2025 · 0 repositories · arXiv:2504.20434
-
CBM-RAG: Demonstrating Enhanced Interpretability in Radiology Report Generation with Multi-Agent RAG and Concept Bottleneck Models 29 Apr 2025 · 1 repository · arXiv:2504.20898
-
DB-GNN: Dual-Branch Graph Neural Network with Multi-Level Contrastive Learning for Jointly Identifying Within- and Cross-Frequency Coupled Brain Networks 29 Apr 2025 · 0 repositories · arXiv:2504.20744
-
Geolocating Earth Imagery from ISS: Integrating Machine Learning with Astronaut Photography for Enhanced Geographic Mapping 29 Apr 2025 · 1 repository · arXiv:2504.21194
-
Graph RAG for Legal Norms: A Hierarchical and Temporal Approach 29 Apr 2025 · 0 repositories · arXiv:2505.00039
-
In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer 29 Apr 2025 · 0 repositories · arXiv:2504.20690
-
ISDrama: Immersive Spatial Drama Generation through Multimodal Prompting 29 Apr 2025 · 0 repositories · arXiv:2504.20630
-
JaccDiv: A Metric and Benchmark for Quantifying Diversity of Generated Marketing Text in the Music Industry 29 Apr 2025 · 0 repositories · arXiv:2504.20849
-
Multimodal Large Language Models for Medicine: A Comprehensive Survey 29 Apr 2025 · 0 repositories · arXiv:2504.21051
-
ReasonIR: Training Retrievers for Reasoning Tasks 29 Apr 2025 · 1 repository · arXiv:2504.20595Syntology official: no sample here; runs from other or unrecorded repositories · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
SteelBlastQC: Shot-blasted Steel Surface Dataset with Interpretable Detection of Surface Defects 29 Apr 2025 · 1 repository · arXiv:2504.20510
-
Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition 29 Apr 2025 · 1 repository · arXiv:2504.20938
-
UniDetox: Universal Detoxification of Large Language Models via Dataset Distillation 29 Apr 2025 · 1 repository · arXiv:2504.20500Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
UniversalRAG: Retrieval-Augmented Generation over Corpora of Diverse Modalities and Granularities 29 Apr 2025 · 1 repository · arXiv:2504.20734
-
YoChameleon: Personalized Vision and Language Generation 29 Apr 2025 · 0 repositories · arXiv:2504.20998
-
A Transformer-Based Approach for Diagnosing Fault Cases in Optical Fiber Amplifiers 28 Apr 2025 · 0 repositories · arXiv:2505.06245
-
Can LLMs Be Trusted for Evaluating RAG Systems? A Survey of Methods and Datasets 28 Apr 2025 · 0 repositories · arXiv:2504.20119
-
Chatbot Arena Meets Nuggets: Towards Explanations and Diagnostics in the Evaluation of LLM Responses 28 Apr 2025 · 0 repositories · arXiv:2504.20006
-
Enhancing Systematic Reviews with Large Language Models: Using GPT-4 and Kimi 28 Apr 2025 · 0 repositories · arXiv:2504.20276
-
Geometry-Informed Neural Operator Transformer 28 Apr 2025 · 1 repository · arXiv:2504.19452
-
Large Language Models are Qualified Benchmark Builders: Rebuilding Pre-Training Datasets for Advancing Code Intelligence Tasks 28 Apr 2025 · 0 repositories · arXiv:2504.19444
-
m-KAILIN: Knowledge-Driven Agentic Scientific Corpus Distillation Framework for Biomedical Large Language Models Training 28 Apr 2025 · 0 repositories · arXiv:2504.19565
-
Reconstructing Context: Evaluating Advanced Chunking Strategies for Retrieval-Augmented Generation 28 Apr 2025 · 1 repository · arXiv:2504.19754
-
TreeHop: Generate and Filter Next Query Embeddings Efficiently for Multi-hop Question Answering 28 Apr 2025 · 1 repository · arXiv:2504.20114
-
UNet with Axial Transformer : A Neural Weather Model for Precipitation Nowcasting 28 Apr 2025 · 1 repository · arXiv:2504.19408
-
Enhancing Speech-to-Speech Dialogue Modeling with End-to-End Retrieval-Augmented Generation 27 Apr 2025 · 1 repository · arXiv:2505.00028
-
From Inductive to Deductive: LLMs-Based Qualitative Data Analysis in Requirements Engineering 27 Apr 2025 · 1 repository · arXiv:2504.19384
-
LM-MCVT: A Lightweight Multi-modal Multi-view Convolutional-Vision Transformer Approach for 3D Object Recognition 27 Apr 2025 · 0 repositories · arXiv:2504.19256
-
VeriDebug: A Unified LLM for Verilog Debugging via Contrastive Embedding and Guided Correction 27 Apr 2025 · 0 repositories · arXiv:2504.19099
-
TSRM: A Lightweight Temporal Feature Encoding Architecture for Time Series Forecasting and Imputation 26 Apr 2025 · 1 repository · arXiv:2504.18878
-
Why you shouldn't fully trust ChatGPT: A synthesis of this AI tool's error rates across disciplines and the software engineering lifecycle 26 Apr 2025 · 0 repositories · arXiv:2504.18858
-
A model and package for German ColBERT 25 Apr 2025 · 0 repositories · arXiv:2504.20083
-
Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections 25 Apr 2025 · 0 repositories · arXiv:2504.18333
-
MEDIBENG WHISPER TINY: A FINE-TUNED CODE-SWITCHED BENGALI-ENGLISH TRANSLATOR FOR CLINICAL APPLICATIONS 25 Apr 2025 · 1 repository
-
SMARTFinRAG: Interactive Modularized Financial RAG Benchmark 25 Apr 2025 · 1 repository · arXiv:2504.18024
-
A RAG-Based Multi-Agent LLM System for Natural Hazard Resilience and Adaptation 24 Apr 2025 · 1 repository · arXiv:2504.17200
-
A Spatially-Aware Multiple Instance Learning Framework for Digital Pathology 24 Apr 2025 · 1 repository · arXiv:2504.17379
-
Masked strategies for images with small objects 24 Apr 2025 · 0 repositories · arXiv:2504.17935
-
Optimism, Expectation, or Sarcasm? Multi-Class Hope Speech Detection in Spanish and English 24 Apr 2025 · 0 repositories · arXiv:2504.17974
-
The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs 24 Apr 2025 · 0 repositories · arXiv:2504.17768
-
Token-Shuffle: Towards High-Resolution Image Generation with Autoregressive Models 24 Apr 2025 · 0 repositories · arXiv:2504.17789
-
A Novel Hybrid Approach Using an Attention-Based Transformer + GRU Model for Predicting Cryptocurrency Prices 23 Apr 2025 · 0 repositories · arXiv:2504.17079
-
A Survey of Foundation Model-Powered Recommender Systems: From Feature-Based, Generative to Agentic Paradigms 23 Apr 2025 · 0 repositories · arXiv:2504.16420
-
Advanced Chest X-Ray Analysis via Transformer-Based Image Descriptors and Cross-Model Attention Mechanism 23 Apr 2025 · 0 repositories · arXiv:2504.16774
-
Amplified Vulnerabilities: Structured Jailbreak Attacks on LLM-based Multi-Agent Debate 23 Apr 2025 · 0 repositories · arXiv:2504.16489
-
Emo Pillars: Knowledge Distillation to Support Fine-Grained Context-Aware and Context-Less Emotion Classification 23 Apr 2025 · 0 repositories · arXiv:2504.16856
-
From Past to Present: A Survey of Malicious URL Detection Techniques, Datasets and Code Repositories 23 Apr 2025 · 0 repositories · arXiv:2504.16449
-
Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments 23 Apr 2025 · 0 repositories · arXiv:2504.17087
-
Transformers for Complex Query Answering over Knowledge Hypergraphs 23 Apr 2025 · 0 repositories · arXiv:2504.16537
-
A Large-scale Class-level Benchmark Dataset for Code Generation with LLMs 22 Apr 2025 · 0 repositories · arXiv:2504.15564
-
Benchmarking LLM for Code Smells Detection: OpenAI GPT-4.0 vs DeepSeek-V3 22 Apr 2025 · 0 repositories · arXiv:2504.16027
-
CiteFix: Enhancing RAG Accuracy Through Post-Processing Citation Correction 22 Apr 2025 · 0 repositories · arXiv:2504.15629
-
COBRA: Algorithm-Architecture Co-optimized Binary Transformer Accelerator for Edge Inference 22 Apr 2025 · 0 repositories · arXiv:2504.16269
-
DiTPainter: Efficient Video Inpainting with Diffusion Transformers 22 Apr 2025 · 0 repositories · arXiv:2504.15661
-
DSDNet: Raw Domain Demoiréing via Dual Color-Space Synergy 22 Apr 2025 · 0 repositories · arXiv:2504.15756
-
FinDER: Financial Dataset for Question Answering and Evaluating Retrieval-Augmented Generation 22 Apr 2025 · 0 repositories · arXiv:2504.15800
-
Grounded in Context: Retrieval-Based Method for Hallucination Detection 22 Apr 2025 · 0 repositories · arXiv:2504.15771
-
LongMamba: Enhancing Mamba's Long Context Capabilities via Training-Free Receptive Field Enlargement 22 Apr 2025 · 1 repository · arXiv:2504.16053
-
Quantum Doubly Stochastic Transformers 22 Apr 2025 · 0 repositories · arXiv:2504.16275Syntology 3 ran (of which 2 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
Synergizing RAG and Reasoning: A Systematic Review 22 Apr 2025 · 0 repositories · arXiv:2504.15909
-
The Viability of Crowdsourcing for RAG Evaluation 22 Apr 2025 · 1 repository · arXiv:2504.15689
-
Acquire and then Adapt: Squeezing out Text-to-Image Model for Image Restoration 21 Apr 2025 · 0 repositories · arXiv:2504.15159
-
AlignRAG: Leveraging Critique Learning for Evidence-Sensitive Retrieval-Augmented Reasoning 21 Apr 2025 · 1 repository · arXiv:2504.14858Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
An Efficient Aerial Image Detection with Variable Receptive Fields 21 Apr 2025 · 0 repositories · arXiv:2504.15165
-
Automated Measurement of Eczema Severity with Self-Supervised Learning 21 Apr 2025 · 0 repositories · arXiv:2504.15193
-
Distribution-aware Dataset Distillation for Efficient Image Restoration 21 Apr 2025 · 0 repositories · arXiv:2504.14826
-
ECViT: Efficient Convolutional Vision Transformer with Local-Attention and Multi-scale Stages 21 Apr 2025 · 1 repository · arXiv:2504.14825
-
Efficient Pretraining Length Scaling 21 Apr 2025 · 0 repositories · arXiv:2504.14992
-
Impact of Latent Space Dimension on IoT Botnet Detection Performance: VAE-Encoder Versus ViT-Encoder 21 Apr 2025 · 0 repositories · arXiv:2504.14879
-
Insert Anything: Image Insertion via In-Context Editing in DiT 21 Apr 2025 · 0 repositories · arXiv:2504.15009
-
LLMs as Data Annotators: How Close Are We to Human Performance 21 Apr 2025 · 0 repositories · arXiv:2504.15022
-
Mitigating Degree Bias in Graph Representation Learning with Learnable Structural Augmentation and Structural Self-Attention 21 Apr 2025 · 1 repository · arXiv:2504.15075
-
MoE Parallel Folding: Heterogeneous Parallelism Mappings for Efficient Large-Scale MoE Model Training with Megatron Core 21 Apr 2025 · 0 repositories · arXiv:2504.14960
-
POLYRAG: Integrating Polyviews into Retrieval-Augmented Generation for Medical Applications 21 Apr 2025 · 0 repositories · arXiv:2504.14917
-
Retrieval Augmented Generation Evaluation in the Era of Large Language Models: A Comprehensive Survey 21 Apr 2025 · 1 repository · arXiv:2504.14891
-
Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction 21 Apr 2025 · 1 repository · arXiv:2504.15266
-
Structure-guided Diffusion Transformer for Low-Light Image Enhancement 21 Apr 2025 · 0 repositories · arXiv:2504.15054
-
Support Evaluation for the TREC 2024 RAG Track: Comparing Human versus LLM Judges 21 Apr 2025 · 0 repositories · arXiv:2504.15205
-
The 1st EReL@MIR Workshop on Efficient Representation Learning for Multimodal Information Retrieval 21 Apr 2025 · 0 repositories · arXiv:2504.14788