Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 32
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 32 of 190: papers 3,101 to 3,200 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
AllRestorer: All-in-One Transformer for Image Restoration under Composite Degradations 16 Nov 2024 · 0 repositories · arXiv:2411.10708
-
Bag of Design Choices for Inference of High-Resolution Masked Generative Transformer 16 Nov 2024 · 1 repository · arXiv:2411.10781
-
IntentGPT: Few-shot Intent Discovery with Large Language Models 16 Nov 2024 · 0 repositories · arXiv:2411.10670
-
MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map 16 Nov 2024 · 1 repository · arXiv:2411.10741Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
MpoxVLM: A Vision-Language Model for Diagnosing Skin Lesions from Mpox Virus Infection 16 Nov 2024 · 1 repository · arXiv:2411.10888
-
A Multi-Scale Spatial-Temporal Network for Wireless Video Transmission 15 Nov 2024 · 0 repositories · arXiv:2411.09936
-
Building 6G Radio Foundation Models with Transformer Architectures 15 Nov 2024 · 0 repositories · arXiv:2411.09996
-
CMATH: Cross-Modality Augmented Transformer with Hierarchical Variational Distillation for Multimodal Emotion Recognition in Conversation 15 Nov 2024 · 0 repositories · arXiv:2411.10060
-
DiMoDif: Discourse Modality-information Differentiation for Audio-visual Deepfake Detection and Localization 15 Nov 2024 · 0 repositories · arXiv:2411.10193
-
Does Prompt Formatting Have Any Impact on LLM Performance? 15 Nov 2024 · 0 repositories · arXiv:2411.10541
-
Evidential Federated Learning for Skin Lesion Image Classification 15 Nov 2024 · 0 repositories · arXiv:2411.10071
-
LoRA-LiteE: A Computationally Efficient Framework for Chatbot Preference-Tuning 15 Nov 2024 · 0 repositories · arXiv:2411.09947
-
MARS: Unleashing the Power of Variance Reduction for Training Large Models 15 Nov 2024 · 2 repositories · arXiv:2411.10438Syntology official (archive's flag): 2 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
Probabilistic Prior Driven Attention Mechanism Based on Diffusion Model for Imaging Through Atmospheric Turbulence 15 Nov 2024 · 0 repositories · arXiv:2411.10321
-
Prompting and Fine-tuning Large Language Models for Automated Code Review Comment Generation 15 Nov 2024 · 0 repositories · arXiv:2411.10129
-
RETR: Multi-View Radar Detection Transformer for Indoor Perception 15 Nov 2024 · 1 repository · arXiv:2411.10293Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 2 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
P² Law: Scaling Law for Post-Training After Model Pruning 15 Nov 2024 · 0 repositories · arXiv:2411.10272
-
Take Package as Language: Anomaly Detection Using Transformer 15 Nov 2024 · 0 repositories · arXiv:2412.04473
-
ULTra: Unveiling Latent Token Interpretability in Transformer Based Understanding 15 Nov 2024 · 0 repositories · arXiv:2411.12589
-
Xmodel-1.5: An 1B-scale Multilingual LLM 15 Nov 2024 · 1 repository · arXiv:2411.10083
-
Adopting RAG for LLM-Aided Future Vehicle Design 14 Nov 2024 · 0 repositories · arXiv:2411.09590
-
Automating Autograding: Large Language Models as Test Suite Generators for Introductory Programming 14 Nov 2024 · 0 repositories · arXiv:2411.09261
-
BabyLM Challenge: Exploring the Effect of Variation Sets on Language Model Training Efficiency 14 Nov 2024 · 0 repositories · arXiv:2411.09587
-
Beyond Static Tools: Evaluating Large Language Models for Cryptographic Misuse Detection 14 Nov 2024 · 0 repositories · arXiv:2411.09772
-
Comprehensive and Practical Evaluation of Retrieval-Augmented Generation Systems for Medical Question Answering 14 Nov 2024 · 0 repositories · arXiv:2411.09213
-
DSCformer: A Dual-Branch Network Integrating Enhanced Dynamic Snake Convolution and SegFormer for Crack Segmentation 14 Nov 2024 · 0 repositories · arXiv:2411.09371
-
DT-JRD: Deep Transformer based Just Recognizable Difference Prediction Model for Video Coding for Machines 14 Nov 2024 · 0 repositories · arXiv:2411.09308
-
Evaluating Gender Bias in Large Language Models 14 Nov 2024 · 0 repositories · arXiv:2411.09826
-
HateGPT: Unleashing GPT-3.5 Turbo to Combat Hate Speech on X 14 Nov 2024 · 0 repositories · arXiv:2411.09214
-
Initial Nugget Evaluation Results for the TREC 2024 RAG Track with the AutoNuggetizer Framework 14 Nov 2024 · 2 repositories · arXiv:2411.09607
-
Local deployment of large-scale music AI models on commodity hardware 14 Nov 2024 · 0 repositories · arXiv:2411.09625
-
MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs 14 Nov 2024 · 1 repository · arXiv:2411.09492
-
OpenGeMM: A High-Utilization GeMM Accelerator Generator with Lightweight RISC-V Control and Tight Memory Coupling 14 Nov 2024 · 1 repository · arXiv:2411.09543
-
Partial Multi-View Clustering via Meta-Learning and Contrastive Feature Alignment 14 Nov 2024 · 0 repositories · arXiv:2411.09758
-
Re-Parameterization of Lightweight Transformer for On-Device Speech Emotion Recognition 14 Nov 2024 · 0 repositories · arXiv:2411.09339
-
SAG-ViT: A Scale-Aware, High-Fidelity Patching Approach with Graph Attention for Vision Transformers 14 Nov 2024 · 1 repository · arXiv:2411.09420
-
A Large-Scale Study of Relevance Assessments with Large Language Models: An Initial Look 13 Nov 2024 · 1 repository · arXiv:2411.08275
-
A Transformer-Based Visual Piano Transcription Algorithm 13 Nov 2024 · 0 repositories · arXiv:2411.09037
-
Continuous GNN-based Anomaly Detection on Edge using Efficient Adaptive Knowledge Graph Learning 13 Nov 2024 · 0 repositories · arXiv:2411.09072
-
LLMStinger: Jailbreaking LLMs using RL fine-tuned LLMs 13 Nov 2024 · 0 repositories · arXiv:2411.08862
-
Oblique Bayesian additive regression trees 13 Nov 2024 · 0 repositories · arXiv:2411.08849
-
Responsible AI in Construction Safety: Systematic Evaluation of Large Language Models and Prompt Engineering 13 Nov 2024 · 0 repositories · arXiv:2411.08320
-
SAD-TIME: a Spatiotemporal-fused network for depression detection with Automated multi-scale Depth-wise and TIME-interval-related common feature extractor 13 Nov 2024 · 0 repositories · arXiv:2411.08521
-
Theoretical Analysis of Byte-Pair Encoding 13 Nov 2024 · 0 repositories · arXiv:2411.08671
-
Towards Optimizing a Retrieval Augmented Generation using Large Language Model on Academic Data 13 Nov 2024 · 0 repositories · arXiv:2411.08438
-
TRACE: Transformer-based Risk Assessment for Clinical Evaluation 13 Nov 2024 · 1 repository · arXiv:2411.08701
-
UIFormer: A Unified Transformer-based Framework for Incremental Few-Shot Object Detection and Instance Segmentation 13 Nov 2024 · 0 repositories · arXiv:2411.08569
-
VALTEST: Automated Validation of Language Model Generated Test Cases 13 Nov 2024 · 0 repositories · arXiv:2411.08254
-
A Preview of XiYan-SQL: A Multi-Generator Ensemble Framework for Text-to-SQL 13 Nov 2024 · 5 repositories · arXiv:2411.08599
-
Breaking the Low-Rank Dilemma of Linear Attention 12 Nov 2024 · 1 repository · arXiv:2411.07635Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
BudgetMLAgent: A Cost-Effective LLM Multi-Agent system for Automating Machine Learning Tasks 12 Nov 2024 · 0 repositories · arXiv:2411.07464
-
Can adversarial attacks by large language models be attributed? 12 Nov 2024 · 0 repositories · arXiv:2411.08003
-
Circuit Complexity Bounds for RoPE-based Transformer Architecture 12 Nov 2024 · 0 repositories · arXiv:2411.07602
-
Controlled Evaluation of Syntactic Knowledge in Multilingual Language Models 12 Nov 2024 · 1 repository · arXiv:2411.07474
-
Efficient Federated Finetuning of Tiny Transformers with Resource-Constrained Devices 12 Nov 2024 · 0 repositories · arXiv:2411.07826
-
Evaluating ChatGPT-3.5 Efficiency in Solving Coding Problems of Different Complexity Levels: An Empirical Analysis 12 Nov 2024 · 1 repository · arXiv:2411.07529
-
Fair Summarization: Bridging Quality and Diversity in Extractive Summaries 12 Nov 2024 · 1 repository · arXiv:2411.07521Syntology official (archive's flag): 3 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Improving Grapheme-to-Phoneme Conversion through In-Context Knowledge Retrieval with Large Language Models 12 Nov 2024 · 0 repositories · arXiv:2411.07563
-
Large Language Models Can Self-Improve in Long-context Reasoning 12 Nov 2024 · 1 repository · arXiv:2411.08147Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Leveraging Multimodal Models for Enhanced Neuroimaging Diagnostics in Alzheimer's Disease 12 Nov 2024 · 0 repositories · arXiv:2411.07871
-
LLM App Squatting and Cloning 12 Nov 2024 · 0 repositories · arXiv:2411.07518
-
Query Optimization for Parametric Knowledge Refinement in Retrieval-Augmented Large Language Models 12 Nov 2024 · 0 repositories · arXiv:2411.07820
-
Retrieval Augmented Time Series Forecasting 12 Nov 2024 · 1 repository · arXiv:2411.08249Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Trustful LLMs: Customizing and Grounding Text Generation with Knowledge Bases and Dual Decoders 12 Nov 2024 · 0 repositories · arXiv:2411.07870
-
Unraveling the Gradient Descent Dynamics of Transformers 12 Nov 2024 · 0 repositories · arXiv:2411.07538
-
Verbosity ≠ Veracity: Demystify Verbosity Compensation Behavior of Large Language Models 12 Nov 2024 · 1 repository · arXiv:2411.07858
-
Ambient AI Scribing Support: Comparing the Performance of Specialized AI Agentic Architecture to Leading Foundational Models 11 Nov 2024 · 0 repositories · arXiv:2411.06713
-
AssistRAG: Boosting the Potential of Large Language Models with an Intelligent Information Assistant 11 Nov 2024 · 1 repository · arXiv:2411.06805Syntology official (archive's flag): 10 ran · 10 ran (of which 1 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
Autonomous Droplet Microfluidic Design Framework with Large Language Models 11 Nov 2024 · 1 repository · arXiv:2411.06691
-
Cancer-Answer: Empowering Cancer Care with Advanced Large Language Models 11 Nov 2024 · 0 repositories · arXiv:2411.06946
-
Explore the Reasoning Capability of LLMs in the Chess Testbed 11 Nov 2024 · 0 repositories · arXiv:2411.06655
-
Invar-RAG: Invariant LLM-aligned Retrieval for Better Generation 11 Nov 2024 · 0 repositories · arXiv:2411.07021
-
On Active Privacy Auditing in Supervised Fine-tuning for White-Box Language Models 11 Nov 2024 · 0 repositories · arXiv:2411.07070
-
StoryTeller: Improving Long Video Description through Global Audio-Visual Character Identification 11 Nov 2024 · 1 repository · arXiv:2411.07076
-
Toward Optimal Search and Retrieval for RAG 11 Nov 2024 · 1 repository · arXiv:2411.07396
-
TreeCoders: Trees of Transformers 11 Nov 2024 · 0 repositories · arXiv:2411.07218
-
White-Box Diffusion Transformer for single-cell RNA-seq generation 11 Nov 2024 · 1 repository · arXiv:2411.06785
-
Emotion-Aware Interaction Design in Intelligent User Interface Using Multi-Modal Deep Learning 10 Nov 2024 · 0 repositories · arXiv:2411.06326
-
Feature Fusion Transferability Aware Transformer for Unsupervised Domain Adaptation 10 Nov 2024 · 2 repositories · arXiv:2411.07794
-
Few-shot Semantic Learning for Robust Multi-Biome 3D Semantic Mapping in Off-Road Environments 10 Nov 2024 · 0 repositories · arXiv:2411.06632
-
Local Implicit Wavelet Transformer for Arbitrary-Scale Super-Resolution 10 Nov 2024 · 1 repository · arXiv:2411.06442
-
Prompt-Efficient Fine-Tuning for GPT-like Deep Models to Reduce Hallucination and to Improve Reproducibility in Scientific Text Generation Using Stochastic Optimisation Techniques 10 Nov 2024 · 0 repositories · arXiv:2411.06445
-
Region-Aware Text-to-Image Generation via Hard Binding and Soft Refinement 10 Nov 2024 · 1 repository · arXiv:2411.06558Syntology official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 1 pointer-only (licence)
-
AI's Spatial Intelligence: Evaluating AI's Understanding of Spatial Transformations in PSVT:R and Augmented Reality 9 Nov 2024 · 0 repositories · arXiv:2411.06269
-
Clustering Algorithms and RAG Enhancing Semi-Supervised Text Classification with Large LLMs 9 Nov 2024 · 0 repositories · arXiv:2411.06175
-
Detecting Reference Errors in Scientific Literature with Large Language Models 9 Nov 2024 · 1 repository · arXiv:2411.06101
-
Exploring Knowledge Boundaries in Large Language Models for Retrieval Judgment 9 Nov 2024 · 0 repositories · arXiv:2411.06207
-
Leveraging Retrieval-Augmented Generation for Persian University Knowledge Retrieval 9 Nov 2024 · 0 repositories · arXiv:2411.06237
-
NeuReg: Domain-invariant 3D Image Registration on Human and Mouse Brains 9 Nov 2024 · 0 repositories · arXiv:2411.06315
-
Pattern Integration and Enhancement Vision Transformer for Self-Supervised Learning in Remote Sensing 9 Nov 2024 · 0 repositories · arXiv:2411.06091
-
Selective State Space Model for Monaural Speech Enhancement 9 Nov 2024 · 0 repositories · arXiv:2411.06217
-
Sufficient Context: A New Lens on Retrieval Augmented Generation Systems 9 Nov 2024 · 0 repositories · arXiv:2411.06037
-
ViTOC: Vision Transformer and Object-aware Captioner 9 Nov 2024 · 0 repositories · arXiv:2411.07265
-
AgentOps: Enabling Observability of LLM Agents 8 Nov 2024 · 1 repository · arXiv:2411.05285
-
Autoregressive Adaptive Hypergraph Transformer for Skeleton-based Activity Recognition 8 Nov 2024 · 1 repository · arXiv:2411.05692
-
Cascaded Dual Vision Transformer for Accurate Facial Landmark Detection 8 Nov 2024 · 1 repository · arXiv:2411.07167
-
Classification of Adventitious Sounds Combining Cochleogram and Vision Transformers 8 Nov 2024 · 0 repositories · arXiv:2411.05955
-
Efficient Self-Supervised Barlow Twins from Limited Tissue Slide Cohorts for Colonic Pathology Diagnostics 8 Nov 2024 · 1 repository · arXiv:2411.05959
-
Emotional Images: Assessing Emotions in Images and Potential Biases in Generative Models 8 Nov 2024 · 0 repositories · arXiv:2411.05985
-
Enhancing Visual Classification using Comparative Descriptors 8 Nov 2024 · 1 repository · arXiv:2411.05357