Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 71
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 71 of 190: papers 7,001 to 7,100 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
SPMamba: State-space model is all you need in speech separation 2 Apr 2024 · 1 repository · arXiv:2404.02063Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 2 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 2 pointer-only (licence)
-
Toward Informal Language Processing: Knowledge of Slang in Large Language Models 2 Apr 2024 · 1 repository · arXiv:2404.02323
-
WcDT: World-centric Diffusion Transformer for Traffic Scene Generation 2 Apr 2024 · 1 repository · arXiv:2404.02082Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
Action Detection via an Image Diffusion Process 1 Apr 2024 · 0 repositories · arXiv:2404.01051
-
Advancing AI with Integrity: Ethical Challenges and Solutions in Neural Machine Translation 1 Apr 2024 · 0 repositories · arXiv:2404.01070
-
ARAGOG: Advanced RAG Output Grading 1 Apr 2024 · 1 repository · arXiv:2404.01037
-
Artificial Intelligence and the Spatial Documentation of Languages 1 Apr 2024 · 0 repositories · arXiv:2404.01263
-
Automated Assessment of Encouragement and Warmth in Classrooms Leveraging Multimodal Emotional Features and ChatGPT 1 Apr 2024 · 0 repositories · arXiv:2404.15310
-
BERT-Enhanced Retrieval Tool for Homework Plagiarism Detection System 1 Apr 2024 · 0 repositories · arXiv:2404.01582
-
CMT: Cross Modulation Transformer with Hybrid Loss for Pansharpening 1 Apr 2024 · 0 repositories · arXiv:2404.01121
-
FABLES: Evaluating faithfulness and content selection in book-length summarization 1 Apr 2024 · 3 repositories · arXiv:2404.01261Syntology community repositories only · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Flare-Free Vision: Empowering Uformer with Depth Insights 1 Apr 2024 · 1 repository
-
Forklift: An Extensible Neural Lifter 1 Apr 2024 · 0 repositories · arXiv:2404.16041
-
IsoBench: Benchmarking Multimodal Foundation Models on Isomorphic Representations 1 Apr 2024 · 0 repositories · arXiv:2404.01266
-
Large Motion Model for Unified Multi-Modal Motion Generation 1 Apr 2024 · 0 repositories · arXiv:2404.01284
-
LLM-RadJudge: Achieving Radiologist-Level Evaluation for X-Ray Report Generation 1 Apr 2024 · 0 repositories · arXiv:2404.00998
-
NeRF-MAE: Masked AutoEncoders for Self-Supervised 3D Representation Learning for Neural Radiance Fields 1 Apr 2024 · 1 repository · arXiv:2404.01300Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
On the Faithfulness of Vision Transformer Explanations 1 Apr 2024 · 0 repositories · arXiv:2404.01415
-
Prompt Learning for Oriented Power Transmission Tower Detection in High-Resolution SAR Images 1 Apr 2024 · 0 repositories · arXiv:2404.01074
-
Unveiling Divergent Inductive Biases of LLMs on Temporal Data 1 Apr 2024 · 1 repository · arXiv:2404.01453
-
A General and Efficient Training for Transformer via Token Expansion 31 Mar 2024 · 1 repository · arXiv:2404.00672Syntology official (archive's flag): 7 ran · 7 ran (of which 1 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples)
-
A Theory for Length Generalization in Learning to Reason 31 Mar 2024 · 0 repositories · arXiv:2404.00560
-
Algorithmic Collusion by Large Language Models 31 Mar 2024 · 0 repositories · arXiv:2404.00806
-
CHOPS: CHat with custOmer Profile Systems for Customer Service with LLMs 31 Mar 2024 · 1 repository · arXiv:2404.01343
-
CoUDA: Coherence Evaluation via Unified Data Augmentation 31 Mar 2024 · 1 repository · arXiv:2404.00681
-
DMSSN: Distilled Mixed Spectral-Spatial Network for Hyperspectral Salient Object Detection 31 Mar 2024 · 1 repository · arXiv:2404.00694
-
DRCT: Saving Image Super-resolution away from Information Bottleneck 31 Mar 2024 · 1 repository · arXiv:2404.00722Syntology official (archive's flag): 1 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 3 pointer-only (licence)
-
Dual DETRs for Multi-Label Temporal Action Detection 31 Mar 2024 · 0 repositories · arXiv:2404.00653
-
EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-World Code Repositories 31 Mar 2024 · 1 repository · arXiv:2404.00599Syntology official (archive's flag): 15 ran · 15 ran (of which 0 constructed an object rather than computing a result; 14 with no instrument failure: 3 honoured, 0 violated, 11 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 15 harvested samples) · 4 pointer-only (licence)
-
Extracting Social Determinants of Health from Pediatric Patient Notes Using Large Language Models: Novel Corpus and Methods 31 Mar 2024 · 1 repository · arXiv:2404.00826
-
How Much are Large Language Models Contaminated? A Comprehensive Survey and the LLMSanitize Library 31 Mar 2024 · 1 repository · arXiv:2404.00699Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples)
-
KTPFormer: Kinematics and Trajectory Prior Knowledge-Enhanced Transformer for 3D Human Pose Estimation 31 Mar 2024 · 1 repository · arXiv:2404.00658Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 1 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
MugenNet: A Novel Combined Convolution Neural Network and Transformer Network with its Application for Colonic Polyp Image Segmentation 31 Mar 2024 · 0 repositories · arXiv:2404.00726
-
Observations on Building RAG Systems for Technical Documents 31 Mar 2024 · 0 repositories · arXiv:2404.00657
-
On Difficulties of Attention Factorization through Shared Memory 31 Mar 2024 · 1 repository · arXiv:2404.00798Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Revealing Trends in Datasets from the 2022 ACL and EMNLP Conferences 31 Mar 2024 · 0 repositories · arXiv:2404.08666
-
RQ-RAG: Learning to Refine Queries for Retrieval Augmented Generation 31 Mar 2024 · 1 repository · arXiv:2404.00610
-
Training-Free Semantic Segmentation via LLM-Supervision 31 Mar 2024 · 0 repositories · arXiv:2404.00701
-
Transformer based Pluralistic Image Completion with Reduced Information Loss 31 Mar 2024 · 1 repository · arXiv:2404.00513
-
A Comprehensive Study on NLP Data Augmentation for Hate Speech Detection: Legacy Methods, BERT, and LLMs 30 Mar 2024 · 0 repositories · arXiv:2404.00303
-
A Novel Feature Map Enhancement Technique Integrating Residual CNN and Transformer for Alzheimer Diseases Diagnosis 30 Mar 2024 · 0 repositories · arXiv:2405.12986
-
Can LLMs Master Math? Investigating Large Language Models on Math Stack Exchange 30 Mar 2024 · 1 repository · arXiv:2404.00344Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 4 harvested samples)
-
Dependability Evaluation of Stable Diffusion with Soft Errors on the Model Parameters 30 Mar 2024 · 0 repositories · arXiv:2404.00352
-
Edinburgh Clinical NLP at SemEval-2024 Task 2: Fine-tune your model unless you have access to GPT-4 30 Mar 2024 · 1 repository · arXiv:2404.00484
-
Injecting New Knowledge into Large Language Models via Supervised Fine-Tuning 30 Mar 2024 · 0 repositories · arXiv:2404.00213
-
Small Language Models Learn Enhanced Reasoning Skills from Medical Textbooks 30 Mar 2024 · 0 repositories · arXiv:2404.00376
-
Spread Your Wings: A Radial Strip Transformer for Image Deblurring 30 Mar 2024 · 0 repositories · arXiv:2404.00358
-
A Parallel Attention Network for Cattle Face Recognition 29 Mar 2024 · 0 repositories · arXiv:2403.19980
-
A Systematic Analysis of Subwords and Cross-Lingual Transfer in Multilingual Translation 29 Mar 2024 · 0 repositories · arXiv:2403.20157
-
ChatGPT v.s. Media Bias: A Comparative Study of GPT-3.5 and Fine-tuned Language Models 29 Mar 2024 · 0 repositories · arXiv:2403.20158
-
Classifying Conspiratorial Narratives At Scale: False Alarms and Erroneous Connections 29 Mar 2024 · 1 repository · arXiv:2404.00141
-
DataAgent: Evaluating Large Language Models' Ability to Answer Zero-Shot, Natural Language Queries 29 Mar 2024 · 0 repositories · arXiv:2404.00188
-
Decision Mamba: Reinforcement Learning via Sequence Modeling with Selective State Spaces 29 Mar 2024 · 1 repository · arXiv:2403.19925Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
DiJiang: Efficient Large Language Models through Compact Kernelization 29 Mar 2024 · 1 repository · arXiv:2403.19928Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Enhancing the General Agent Capabilities of Low-Parameter LLMs through Tuning and Multi-Branch Reasoning 29 Mar 2024 · 1 repository · arXiv:2403.19962Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Latxa: An Open Language Model and Evaluation Suite for Basque 29 Mar 2024 · 1 repository · arXiv:2403.20266
-
Localising the Seizure Onset Zone from Single-Pulse Electrical Stimulation Responses with a CNN Transformer 29 Mar 2024 · 1 repository · arXiv:2403.20324
-
MANGO: A Benchmark for Evaluating Mapping and Navigation Abilities of Large Language Models 29 Mar 2024 · 1 repository · arXiv:2403.19913Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
On-the-fly Definition Augmentation of LLMs for Biomedical NER 29 Mar 2024 · 1 repository · arXiv:2404.00152Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples)
-
ReALM: Reference Resolution As Language Modeling 29 Mar 2024 · 0 repositories · arXiv:2403.20329
-
SceneTracker: Long-term Scene Flow Estimation Network 29 Mar 2024 · 1 repository · arXiv:2403.19924Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
A Novel Stochastic Transformer-based Approach for Post-Traumatic Stress Disorder Detection using Audio Recording of Clinical Interviews 28 Mar 2024 · 0 repositories · arXiv:2403.19441
-
A Review of Multi-Modal Large Language and Vision Models 28 Mar 2024 · 0 repositories · arXiv:2404.01322
-
AAPMT: AGI Assessment Through Prompt and Metric Transformer 28 Mar 2024 · 1 repository · arXiv:2403.19101
-
Are Large Language Models Good at Utility Judgments? 28 Mar 2024 · 1 repository · arXiv:2403.19216Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Checkpoint Merging via Bayesian Optimization in LLM Pretraining 28 Mar 2024 · 0 repositories · arXiv:2403.19390
-
Code Comparison Tuning for Code Large Language Models 28 Mar 2024 · 0 repositories · arXiv:2403.19121
-
DenseNets Reloaded: Paradigm Shift Beyond ResNets and ViTs 28 Mar 2024 · 3 repositories · arXiv:2403.19588Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 3 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Enhancing Efficiency in Vision Transformer Networks: Design Techniques and Insights 28 Mar 2024 · 0 repositories · arXiv:2403.19882
-
FACTOID: FACtual enTailment fOr hallucInation Detection 28 Mar 2024 · 0 repositories · arXiv:2403.19113
-
Generating Multi-Aspect Queries for Conversational Search 28 Mar 2024 · 0 repositories · arXiv:2403.19302
-
Genetic Quantization-Aware Approximation for Non-Linear Operations in Transformers 28 Mar 2024 · 1 repository · arXiv:2403.19591
-
Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models 28 Mar 2024 · 1 repository · arXiv:2403.19521Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 14 harvested samples) · 14 pointer-only (licence)
-
Jamba: A Hybrid Transformer-Mamba Language Model 28 Mar 2024 · 3 repositories · arXiv:2403.19887
-
Just-DNA-Seq, open-source personal genomics platform: longevity science for everyone 28 Mar 2024 · 0 repositories · arXiv:2403.19087
-
Keypoint Action Tokens Enable In-Context Imitation Learning in Robotics 28 Mar 2024 · 0 repositories · arXiv:2403.19578
-
MATEval: A Multi-Agent Discussion Framework for Advancing Open-Ended Text Evaluation 28 Mar 2024 · 1 repository · arXiv:2403.19305
-
Mitigating Misleading Chain-of-Thought Reasoning with Selective Filtering 28 Mar 2024 · 1 repository · arXiv:2403.19167
-
Single-Shared Network with Prior-Inspired Loss for Parameter-Efficient Multi-Modal Imaging Skin Lesion Classification 28 Mar 2024 · 0 repositories · arXiv:2403.19203
-
A Survey on Large Language Models from Concept to Implementation 27 Mar 2024 · 0 repositories · arXiv:2403.18969
-
Attention-aware semantic relevance predicting Chinese sentence reading 27 Mar 2024 · 0 repositories · arXiv:2403.18542
-
BioMedLM: A 2.7B Parameter Language Model Trained On Biomedical Text 27 Mar 2024 · 1 repository · arXiv:2403.18421Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
BLADE: Enhancing Black-box Large Language Models with Small Domain-Specific Models 27 Mar 2024 · 0 repositories · arXiv:2403.18365
-
Boosting Conversational Question Answering with Fine-Grained Retrieval-Augmentation and Self-Check 27 Mar 2024 · 0 repositories · arXiv:2403.18243
-
CPR: Retrieval Augmented Generation for Copyright Protection 27 Mar 2024 · 0 repositories · arXiv:2403.18920
-
Cross-domain Fiber Cluster Shape Analysis for Language Performance Cognitive Score Prediction 27 Mar 2024 · 0 repositories · arXiv:2403.19001
-
Evaluating Large Language Models for Health-Related Text Classification Tasks with Public Social Media Data 27 Mar 2024 · 0 repositories · arXiv:2403.19031
-
Faster Convergence for Transformer Fine-tuning with Line Search Methods 27 Mar 2024 · 1 repository · arXiv:2403.18506
-
Cross-System Categorization of Abnormal Traces in Microservice-Based Systems via Meta-Learning 27 Mar 2024 · 0 repositories · arXiv:2403.18998
-
Fourier or Wavelet bases as counterpart self-attention in spikformer for efficient visual classification 27 Mar 2024 · 0 repositories · arXiv:2403.18228
-
Illicit object detection in X-ray images using Vision Transformers 27 Mar 2024 · 0 repositories · arXiv:2403.19043
-
LLMs in HCI Data Work: Bridging the Gap Between Information Retrieval and Responsible Research Practices 27 Mar 2024 · 0 repositories · arXiv:2403.18173
-
Long-form factuality in large language models 27 Mar 2024 · 3 repositories · arXiv:2403.18802Syntology official (archive's flag): 4 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 4 pointer-only (licence)
-
Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models 27 Mar 2024 · 2 repositories · arXiv:2403.18814Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 1 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples) · 1 pointer-only (licence)
-
ParCo: Part-Coordinating Text-to-Motion Synthesis 27 Mar 2024 · 1 repository · arXiv:2403.18512
-
RankMamba: Benchmarking Mamba's Document Ranking Performance in the Era of Transformers 27 Mar 2024 · 1 repository · arXiv:2403.18276
-
Reshaping Free-Text Radiology Notes Into Structured Reports With Generative Transformers 27 Mar 2024 · 1 repository · arXiv:2403.18938
-
ViTAR: Vision Transformer with Any Resolution 27 Mar 2024 · 0 repositories · arXiv:2403.18361
-
Vulnerability Detection with Code Language Models: How Far Are We? 27 Mar 2024 · 1 repository · arXiv:2403.18624Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples)
-
ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching 26 Mar 2024 · 0 repositories · arXiv:2403.17312