Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 34
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 34 of 190: papers 3,301 to 3,400 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
HiMemFormer: Hierarchical Memory-Aware Transformer for Multi-Agent Action Anticipation 3 Nov 2024 · 0 repositories · arXiv:2411.01455
-
Integration of Large Vision Language Models for Efficient Post-disaster Damage Assessment and Reporting 3 Nov 2024 · 0 repositories · arXiv:2411.01511
-
LinRec: Linear Attention Mechanism for Long-term Sequential Recommender Systems 3 Nov 2024 · 1 repository · arXiv:2411.01537Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
UniGuard: Towards Universal Safety Guardrails for Jailbreak Attacks on Multimodal Large Language Models 3 Nov 2024 · 0 repositories · arXiv:2411.01703
-
Can Large Language Model Predict Employee Attrition? 2 Nov 2024 · 0 repositories · arXiv:2411.01353
-
Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders 2 Nov 2024 · 1 repository · arXiv:2411.01220Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Few-Class Arena: A Benchmark for Efficient Selection of Vision Models and Dataset Difficulty Measurement 2 Nov 2024 · 1 repository · arXiv:2411.01099
-
Reasoning Limitations of Multimodal Large Language Models. A case study of Bongard Problems 2 Nov 2024 · 0 repositories · arXiv:2411.01173
-
Task-Aware Harmony Multi-Task Decision Transformer for Offline Reinforcement Learning 2 Nov 2024 · 1 repository · arXiv:2411.01146Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 11 harvested samples) · 2 pointer-only (licence)
-
A Lorentz-Equivariant Transformer for All of the LHC 1 Nov 2024 · 1 repository · arXiv:2411.00446Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
AttackQA: Development and Adoption of a Dataset for Assisting Cybersecurity Operations using Fine-tuned and Open-Source LLMs 1 Nov 2024 · 0 repositories · arXiv:2411.01073
-
CORAG: A Cost-Constrained Retrieval Optimization System for Retrieval-Augmented Generation 1 Nov 2024 · 0 repositories · arXiv:2411.00744
-
Cross-Fundus Transformer for Multi-modal Diabetic Retinopathy Grading with Cataract 1 Nov 2024 · 0 repositories · arXiv:2411.00726
-
Evaluating the Impact of Lab Test Results on Large Language Models Generated Differential Diagnoses from Clinical Case Vignettes 1 Nov 2024 · 0 repositories · arXiv:2411.02523
-
LLM-Ref: Enhancing Reference Handling in Technical Writing with Large Language Models 1 Nov 2024 · 0 repositories · arXiv:2411.00294
-
LLMs: A Game-Changer for Software Engineers? 1 Nov 2024 · 0 repositories · arXiv:2411.00932
-
Provenance: A Light-weight Fact-checker for Retrieval Augmented LLM Generation Output 1 Nov 2024 · 0 repositories · arXiv:2411.01022
-
Rationale-Guided Retrieval Augmented Generation for Medical Question Answering 1 Nov 2024 · 1 repository · arXiv:2411.00300Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Self-Evolved Reward Learning for LLMs 1 Nov 2024 · 1 repository · arXiv:2411.00418Syntology 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
STAA: Spatio-Temporal Attention Attribution for Real-Time Interpreting Transformer-based Video Models 1 Nov 2024 · 1 repository · arXiv:2411.00630
-
Target-Guided Adversarial Point Cloud Transformer Towards Recognition Against Real-world Corruptions 1 Nov 2024 · 1 repository · arXiv:2411.00462Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Towards High-fidelity Head Blending with Chroma Keying for Industrial Applications 1 Nov 2024 · 0 repositories · arXiv:2411.00652
-
Towards Multi-Source Retrieval-Augmented Generation via Synergizing Reasoning and Preference-Driven Retrieval 1 Nov 2024 · 0 repositories · arXiv:2411.00689
-
Ada-MSHyper: Adaptive Multi-Scale Hypergraph Transformer for Time Series Forecasting 31 Oct 2024 · 1 repository · arXiv:2410.23992Syntology official: harvested, nothing ran · 0 ran · 4 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Aerial Flood Scene Classification Using Fine-Tuned Attention-based Architecture for Flood-Prone Countries in South Asia 31 Oct 2024 · 0 repositories · arXiv:2411.00169
-
Analyzing & Reducing the Need for Learning Rate Warmup in GPT Training 31 Oct 2024 · 0 repositories · arXiv:2410.23922
-
Automating Quantum Software Maintenance: Flakiness Detection and Root Cause Analysis 31 Oct 2024 · 0 repositories · arXiv:2410.23578
-
Deep Learning in Long-Short Stock Portfolio Allocation: An Empirical Study 31 Oct 2024 · 0 repositories · arXiv:2411.13555
-
Desert Camels and Oil Sheikhs: Arab-Centric Red Teaming of Frontier LLMs 31 Oct 2024 · 0 repositories · arXiv:2410.24049
-
EDT: An Efficient Diffusion Transformer Framework Inspired by Human-like Sketching 31 Oct 2024 · 1 repository · arXiv:2410.23788Syntology official (archive's flag): 10 ran · 12 ran (of which 7 constructed an object rather than computing a result; 9 with no instrument failure: 1 honoured, 0 violated, 8 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (of 15 harvested samples) · 2 pointer-only (licence)
-
Enhancing Brain Tumor Classification Using TrAdaBoost and Multi-Classifier Deep Learning Approaches 31 Oct 2024 · 0 repositories · arXiv:2411.00875
-
Handwriting Recognition in Historical Documents with Multimodal LLM 31 Oct 2024 · 0 repositories · arXiv:2410.24034
-
Improbable Bigrams Expose Vulnerabilities of Incomplete Tokens in Byte-Level Tokenizers 31 Oct 2024 · 0 repositories · arXiv:2410.23684
-
IO Transformer: Evaluating SwinV2-Based Reward Models for Computer Vision 31 Oct 2024 · 0 repositories · arXiv:2411.00252
-
JEMA: A Joint Embedding Framework for Scalable Co-Learning with Multimodal Alignment 31 Oct 2024 · 0 repositories · arXiv:2410.23988
-
JudgeRank: Leveraging Large Language Models for Reasoning-Intensive Reranking 31 Oct 2024 · 0 repositories · arXiv:2411.00142
-
Large Language Models for Patient Comments Multi-Label Classification 31 Oct 2024 · 0 repositories · arXiv:2410.23528
-
LEAF: Learning and Evaluation Augmented by Fact-Checking to Improve Factualness in Large Language Models 31 Oct 2024 · 0 repositories · arXiv:2410.23526
-
LSEAttention is All You Need for Time Series Forecasting 31 Oct 2024 · 0 repositories · arXiv:2410.23749
-
Morphological Typology in BPE Subword Productivity and Language Modeling 31 Oct 2024 · 0 repositories · arXiv:2410.23656
-
Reinforcement Learning Gradients as Vitamin for Online Finetuning Decision Transformers 31 Oct 2024 · 1 repository · arXiv:2410.24108Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Responsible Retrieval Augmented Generation for Climate Decision Making from Documents 31 Oct 2024 · 0 repositories · arXiv:2410.23902
-
RSL-SQL: Robust Schema Linking in Text-to-SQL Generation 31 Oct 2024 · 1 repository · arXiv:2411.00073Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
SelfCodeAlign: Self-Alignment for Code Generation 31 Oct 2024 · 2 repositories · arXiv:2410.24198Syntology official (archive's flag): 9 ran · 30 ran (of which 3 constructed an object rather than computing a result; 22 with no instrument failure: 1 honoured, 0 violated, 21 with no contract checked; 8 where Syntology's instrument failed) · 7 unverified (of 37 harvested samples)
-
A Comprehensive Study on Quantization Techniques for Large Language Models 30 Oct 2024 · 0 repositories · arXiv:2411.02530
-
A Transformer Model for Segmentation, Classification, and Caller Identification of Marmoset Vocalization 30 Oct 2024 · 0 repositories · arXiv:2410.23279
-
CORAL: Benchmarking Multi-turn Conversational Retrieval-Augmentation Generation 30 Oct 2024 · 1 repository · arXiv:2410.23090
-
Danoliteracy of Generative, Large Language Models 30 Oct 2024 · 0 repositories · arXiv:2410.22839
-
Eliciting Critical Reasoning in Retrieval-Augmented Language Models via Contrastive Explanations 30 Oct 2024 · 0 repositories · arXiv:2410.22874
-
Emergence of meta-stable clustering in mean-field transformer models 30 Oct 2024 · 0 repositories · arXiv:2410.23228
-
Emotional RAG: Enhancing Role-Playing Agents through Emotional Retrieval 30 Oct 2024 · 1 repository · arXiv:2410.23041
-
Epipolar-Free 3D Gaussian Splatting for Generalizable Novel View Synthesis 30 Oct 2024 · 0 repositories · arXiv:2410.22817
-
EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations 30 Oct 2024 · 0 repositories · arXiv:2410.22821Syntology 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 3 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
High-Fidelity Document Stain Removal via A Large-Scale Real-World Dataset and A Memory-Augmented Transformer 30 Oct 2024 · 1 repository · arXiv:2410.22922
-
Higher-order Cross-structural Embedding Model for Time Series Analysis 30 Oct 2024 · 0 repositories · arXiv:2410.22984
-
HijackRAG: Hijacking Attacks against Retrieval-Augmented Large Language Models 30 Oct 2024 · 0 repositories · arXiv:2410.22832
-
Learning to Achieve Goals with Belief State Transformers 30 Oct 2024 · 0 repositories · arXiv:2410.23506
-
LoFLAT: Local Feature Matching using Focused Linear Attention Transformer 30 Oct 2024 · 0 repositories · arXiv:2410.22710
-
ProTransformer: Robustify Transformers via Plug-and-Play Paradigm 30 Oct 2024 · 1 repository · arXiv:2410.23182
-
Retrieval-Augmented Generation with Estimation of Source Reliability 30 Oct 2024 · 0 repositories · arXiv:2410.22954
-
Return Augmented Decision Transformer for Off-Dynamics Reinforcement Learning 30 Oct 2024 · 0 repositories · arXiv:2410.23450
-
SciPIP: An LLM-based Scientific Paper Idea Proposer 30 Oct 2024 · 1 repository · arXiv:2410.23166Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Semantic Enrichment of the Quantum Cascade Laser Properties in Text- A Knowledge Graph Generation Approach 30 Oct 2024 · 1 repository · arXiv:2410.22996
-
st-DTPM: Spatial-Temporal Guided Diffusion Transformer Probabilistic Model for Delayed Scan PET Image Prediction 30 Oct 2024 · 0 repositories · arXiv:2410.22732
-
Long²RAG: Evaluating Long-Context & Long-Form Retrieval-Augmented Generation with Key Point Recall 30 Oct 2024 · 0 repositories · arXiv:2410.23000
-
Very fast Bayesian Additive Regression Trees on GPU 30 Oct 2024 · 1 repository · arXiv:2410.23244
-
A Large Recurrent Action Model: xLSTM enables Fast Inference for Robotics Tasks 29 Oct 2024 · 1 repository · arXiv:2410.22391Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
AmpleGCG-Plus: A Strong Generative Model of Adversarial Suffixes to Jailbreak LLMs with Higher Success Rates in Fewer Attempts 29 Oct 2024 · 1 repository · arXiv:2410.22143Syntology 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Beyond Text: Optimizing RAG with Multimodal Inputs for Industrial Applications 29 Oct 2024 · 1 repository · arXiv:2410.21943Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
CFSafety: Comprehensive Fine-grained Safety Assessment for LLMs 29 Oct 2024 · 0 repositories · arXiv:2410.21695
-
Coupling quantum-like cognition with the neuronal networks within generalized probability theory 29 Oct 2024 · 0 repositories · arXiv:2411.00036
-
DINeuro: Distilling Knowledge from 2D Natural Images via Deformable Tubular Transferring Strategy for 3D Neuron Reconstruction 29 Oct 2024 · 0 repositories · arXiv:2410.22078
-
Dual Conditional Diffusion Models for Sequential Recommendation 29 Oct 2024 · 0 repositories · arXiv:2410.21967
-
Efficient Machine Translation with a BiLSTM-Attention Approach 29 Oct 2024 · 2 repositories · arXiv:2410.22335
-
Emotion-Guided Image to Music Generation 29 Oct 2024 · 0 repositories · arXiv:2410.22299
-
ET-Flow: Equivariant Flow-Matching for Molecular Conformer Generation 29 Oct 2024 · 1 repository · arXiv:2410.22388Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 12 harvested samples)
-
Evaluating K-Fold Cross Validation for Transformer Based Symbolic Regression Models 29 Oct 2024 · 0 repositories · arXiv:2410.21896
-
FactBench: A Dynamic Benchmark for In-the-Wild Language Model Factuality Evaluation 29 Oct 2024 · 0 repositories · arXiv:2410.22257
-
Fourier Head: Helping Large Language Models Learn Complex Probability Distributions 29 Oct 2024 · 0 repositories · arXiv:2410.22269
-
Leveraging User History with Transformers for News Clicking: The DArgk Approach 29 Oct 2024 · 1 repository
-
Long-context Protein Language Modeling Using Bidirectional Mamba with Shared Projection Layers 29 Oct 2024 · 1 repository · arXiv:2411.08909
-
Multi-step feature fusion for natural disaster damage assessment on satellite images 29 Oct 2024 · 1 repository · arXiv:2410.21901
-
On the Role of Depth and Looping for In-Context Learning with Task Diversity 29 Oct 2024 · 0 repositories · arXiv:2410.21698
-
SAM-Swin: SAM-Driven Dual-Swin Transformers with Adaptive Lesion Enhancement for Laryngo-Pharyngeal Tumor Detection 29 Oct 2024 · 1 repository · arXiv:2410.21813
-
Self-Preference Bias in LLM-as-a-Judge 29 Oct 2024 · 0 repositories · arXiv:2410.21819
-
Sequential choice in ordered bundles 29 Oct 2024 · 0 repositories · arXiv:2410.21670
-
Spatio-temporal Transformers for Action Unit Classification with Event Cameras 29 Oct 2024 · 0 repositories · arXiv:2410.21958
-
Topic-Conversation Relevance (TCR) Dataset and Benchmarks 29 Oct 2024 · 1 repository · arXiv:2411.00038Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples)
-
A Simple Yet Effective Corpus Construction Framework for Indonesian Grammatical Error Correction 28 Oct 2024 · 1 repository · arXiv:2410.20838
-
AutoRAG: Automated Framework for optimization of Retrieval Augmented Generation Pipeline 28 Oct 2024 · 2 repositories · arXiv:2410.20878
-
Belief in the Machine: Investigating Epistemological Blind Spots of Language Models 28 Oct 2024 · 1 repository · arXiv:2410.21195
-
BLAST: Block-Level Adaptive Structured Matrices for Efficient Deep Neural Network Inference 28 Oct 2024 · 1 repository · arXiv:2410.21262
-
ByteNet: Rethinking Multimedia File Fragment Classification through Visual Perspectives 28 Oct 2024 · 1 repository · arXiv:2410.20855
-
Calibrated Decision-Making through LLM-Assisted Retrieval 28 Oct 2024 · 0 repositories · arXiv:2411.08891
-
Causal Interventions on Causal Paths: Mapping GPT-2's Reasoning From Syntax to Semantics 28 Oct 2024 · 0 repositories · arXiv:2410.21353
-
Combining Domain-Specific Models and LLMs for Automated Disease Phenotyping from Survey Data 28 Oct 2024 · 0 repositories · arXiv:2410.20695
-
CRAT: A Multi-Agent Framework for Causality-Enhanced Reflective and Retrieval-Augmented Translation with Large Language Models 28 Oct 2024 · 0 repositories · arXiv:2410.21067
-
CT2C-QA: Multimodal Question Answering over Chinese Text, Table and Chart 28 Oct 2024 · 0 repositories · arXiv:2410.21414
-
Deep Learning for Medical Text Processing: BERT Model Fine-Tuning and Comparative Study 28 Oct 2024 · 0 repositories · arXiv:2410.20792
-
Efficient Bilinear Attention-based Fusion for Medical Visual Question Answering 28 Oct 2024 · 0 repositories · arXiv:2410.21000