Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 26
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 26 of 190: papers 2,501 to 2,600 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT 27 Dec 2024 · 1 repository · arXiv:2412.19505Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 1 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples) · 2 pointer-only (licence)
-
Feature Alignment-Based Knowledge Distillation for Efficient Compression of Large Language Models 27 Dec 2024 · 0 repositories · arXiv:2412.19449
-
Generative Pretrained Embedding and Hierarchical Irregular Time Series Representation for Daily Living Activity Recognition 27 Dec 2024 · 1 repository · arXiv:2412.19732
-
Hidformer: Transformer-Style Neural Network in Stock Price Forecasting 27 Dec 2024 · 0 repositories · arXiv:2412.19932
-
Long Context vs. RAG for LLMs: An Evaluation and Revisits 27 Dec 2024 · 1 repository · arXiv:2501.01880
-
Optimizing Local-Global Dependencies for Accurate 3D Human Pose Estimation 27 Dec 2024 · 1 repository · arXiv:2412.19676
-
Revisiting PCA for time series reduction in temporal dimension 27 Dec 2024 · 1 repository · arXiv:2412.19423
-
Toward Adaptive Reasoning in Large Language Models with Thought Rollback 27 Dec 2024 · 1 repository · arXiv:2412.19707Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Context-Aware Deep Learning for Multi Modal Depression Detection 26 Dec 2024 · 1 repository · arXiv:2412.19209
-
DAPoinTr: Domain Adaptive Point Transformer for Point Cloud Completion 26 Dec 2024 · 1 repository · arXiv:2412.19062
-
Dual Channel Multi-Attention in ViT for Biometric Authentication using Forehead Subcutaneous Vein Pattern and Periocular Pattern 26 Dec 2024 · 0 repositories · arXiv:2412.19160
-
MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes 26 Dec 2024 · 1 repository · arXiv:2412.19260
-
On the Expressiveness and Length Generalization of Selective State-Space Models on Regular Languages 26 Dec 2024 · 1 repository · arXiv:2412.19350
-
RAG with Differential Privacy 26 Dec 2024 · 1 repository · arXiv:2412.19291
-
Reversed in Time: A Novel Temporal-Emphasized Benchmark for Cross-Modal Video-Text Retrieval 26 Dec 2024 · 1 repository · arXiv:2412.19178
-
Sentiment trading with large language models 26 Dec 2024 · 0 repositories · arXiv:2412.19245
-
SpectralKD: A Unified Framework for Interpreting and Distilling Vision Transformers via Spectral Analysis 26 Dec 2024 · 1 repository · arXiv:2412.19055
-
Transformer-Based Wireless Capsule Endoscopy Bleeding Tissue Detection and Classification 26 Dec 2024 · 1 repository · arXiv:2412.19218
-
Adopting Trustworthy AI for Sleep Disorder Prediction: Deep Time Series Analysis with Temporal Attention Mechanism and Counterfactual Explanations 25 Dec 2024 · 0 repositories · arXiv:2412.18971
-
DCIS: Efficient Length Extrapolation of LLMs via Divide-and-Conquer Scaling Factor Search 25 Dec 2024 · 1 repository · arXiv:2412.18811
-
Distortion-Aware Adversarial Attacks on Bounding Boxes of Object Detectors 25 Dec 2024 · 1 repository · arXiv:2412.18815
-
EC-Diffuser: Multi-Object Manipulation via Entity-Centric Behavior Generation 25 Dec 2024 · 0 repositories · arXiv:2412.18907
-
Evaluating the Adversarial Robustness of Detection Transformers 25 Dec 2024 · 0 repositories · arXiv:2412.18718
-
HAND: Hierarchical Attention Network for Multi-Scale Handwritten Document Recognition and Layout Analysis 25 Dec 2024 · 0 repositories · arXiv:2412.18981
-
Ister: Inverted Seasonal-Trend Decomposition Transformer for Explainable Multivariate Time Series Forecasting 25 Dec 2024 · 0 repositories · arXiv:2412.18798
-
MTCAE-DFER: Multi-Task Cascaded Autoencoder for Dynamic Facial Expression Recognition 25 Dec 2024 · 1 repository · arXiv:2412.18988
-
Optimizing Large Language Models with an Enhanced LoRA Fine-Tuning Algorithm for Efficiency and Robustness in NLP Tasks 25 Dec 2024 · 0 repositories · arXiv:2412.18729
-
Position-aware Graph Transformer for Recommendation 25 Dec 2024 · 0 repositories · arXiv:2412.18731
-
SAFLITE: Fuzzing Autonomous Systems via Large Language Models 25 Dec 2024 · 0 repositories · arXiv:2412.18727
-
UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation 25 Dec 2024 · 0 repositories · arXiv:2412.18928
-
Using Large Language Models for Automated Grading of Student Writing about Science 25 Dec 2024 · 0 repositories · arXiv:2412.18719
-
Whose Morality Do They Speak? Unraveling Cultural Bias in Multilingual Language Models 25 Dec 2024 · 0 repositories · arXiv:2412.18863
-
Advancing Explainability in Neural Machine Translation: Analytical Metrics for Attention and Alignment Consistency 24 Dec 2024 · 0 repositories · arXiv:2412.18669
-
AutoSculpt: A Pattern-based Model Auto-pruning Framework Using Reinforcement Learning and Graph Learning 24 Dec 2024 · 0 repositories · arXiv:2412.18091
-
Decentralized Intelligence in GameFi: Embodied AI Agents and the Convergence of DeFi and Virtual Ecosystems 24 Dec 2024 · 1 repository · arXiv:2412.18601
-
DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation 24 Dec 2024 · 1 repository · arXiv:2412.18597Syntology official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Do Language Models Understand the Cognitive Tasks Given to Them? Investigations with the N-Back Paradigm 24 Dec 2024 · 0 repositories · arXiv:2412.18120
-
EvoPat: A Multi-LLM-based Patents Summarization and Analysis Agent 24 Dec 2024 · 0 repositories · arXiv:2412.18100
-
GeAR: Graph-enhanced Agent for Retrieval-augmented Generation 24 Dec 2024 · 0 repositories · arXiv:2412.18431
-
HTR-JAND: Handwritten Text Recognition with Joint Attention Network and Knowledge Distillation 24 Dec 2024 · 0 repositories · arXiv:2412.18524
-
Improving Factuality with Explicit Working Memory 24 Dec 2024 · 0 repositories · arXiv:2412.18069
-
Leveraging Convolutional Neural Network-Transformer Synergy for Predictive Modeling in Risk-Based Applications 24 Dec 2024 · 0 repositories · arXiv:2412.18222
-
Molly: Making Large Language Model Agents Solve Python Problem More Logically 24 Dec 2024 · 0 repositories · arXiv:2412.18093
-
Multilingual Mathematical Reasoning: Advancing Open-Source LLMs in Hindi and English 24 Dec 2024 · 1 repository · arXiv:2412.18415
-
Pirates of the RAG: Adaptively Attacking LLMs to Leak Knowledge Bases 24 Dec 2024 · 0 repositories · arXiv:2412.18295
-
Segment-Based Attention Masking for GPTs 24 Dec 2024 · 1 repository · arXiv:2412.18487
-
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models 24 Dec 2024 · 1 repository · arXiv:2412.18675
-
TimelyLLM: Segmented LLM Serving System for Time-sensitive Robotic Applications 24 Dec 2024 · 0 repositories · arXiv:2412.18695
-
A Survey of Query Optimization in Large Language Models 23 Dec 2024 · 0 repositories · arXiv:2412.17558
-
An Experimental Evaluation of Japanese Tokenizers for Sentiment-Based Text Classification 23 Dec 2024 · 1 repository · arXiv:2412.17361
-
CiteBART: Learning to Generate Citations for Local Citation Recommendation 23 Dec 2024 · 1 repository · arXiv:2412.17534Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 16 harvested samples)
-
DiffFormer: a Differential Spatial-Spectral Transformer for Hyperspectral Image Classification 23 Dec 2024 · 1 repository · arXiv:2412.17350
-
Edge-AI for Agriculture: Lightweight Vision Models for Disease Detection in Resource-Limited Settings 23 Dec 2024 · 0 repositories · arXiv:2412.18635
-
Efficient fine-tuning methodology of text embedding models for information retrieval: contrastive learning penalty (clp) 23 Dec 2024 · 1 repository · arXiv:2412.17364
-
Fast Gradient Computation for RoPE Attention in Almost Linear Time 23 Dec 2024 · 0 repositories · arXiv:2412.17316
-
LayerDropBack: A Universally Applicable Approach for Accelerating Training of Deep Networks 23 Dec 2024 · 1 repository · arXiv:2412.18027
-
Multimodal Preference Data Synthetic Alignment with Reward Model 23 Dec 2024 · 1 repository · arXiv:2412.17417
-
STeInFormer: Spatial-Temporal Interaction Transformer Architecture for Remote Sensing Change Detection 23 Dec 2024 · 1 repository · arXiv:2412.17247
-
Theoretical Constraints on the Expressive Power of RoPE-based Tensor Attention Transformers 23 Dec 2024 · 0 repositories · arXiv:2412.18040
-
Token Statistics Transformer: Linear-Time Attention via Variational Rate Reduction 23 Dec 2024 · 1 repository · arXiv:2412.17810
-
URoadNet: Dual Sparse Attentive U-Net for Multiscale Road Network Extraction 23 Dec 2024 · 0 repositories · arXiv:2412.17573
-
A Reality Check on Context Utilisation for Retrieval-Augmented Generation 22 Dec 2024 · 1 repository · arXiv:2412.17031
-
An OpenMind for 3D medical vision self-supervised learning 22 Dec 2024 · 1 repository · arXiv:2412.17041
-
Bridging Auditory Perception and Language Comprehension through MEG-Driven Encoding Models 22 Dec 2024 · 0 repositories · arXiv:2501.03246
-
DR-Encoder: Encode Low-rank Gradients with Random Prior for Large Language Models Differentially Privately 22 Dec 2024 · 0 repositories · arXiv:2412.17053
-
Multifaceted User Modeling in Recommendation: A Federated Foundation Models Approach 22 Dec 2024 · 1 repository · arXiv:2412.16969
-
On Fusing ChatGPT and Ensemble Learning in Discon-tinuous Named Entity Recognition in Health Corpora 22 Dec 2024 · 0 repositories · arXiv:2412.16976
-
PsychAdapter: Adapting LLM Transformers to Reflect Traits, Personality and Mental Health 22 Dec 2024 · 1 repository · arXiv:2412.16882
-
Reconsidering SMT Over NMT for Closely Related Languages: A Case Study of Persian-Hindi Pair 22 Dec 2024 · 0 repositories · arXiv:2412.16877
-
Robustness of Large Language Models Against Adversarial Attacks 22 Dec 2024 · 0 repositories · arXiv:2412.17011
-
SubstationAI: Multimodal Large Model-Based Approaches for Analyzing Substation Equipment Faults 22 Dec 2024 · 0 repositories · arXiv:2412.17077
-
Survey on Abstractive Text Summarization: Dataset, Models, and Metrics 22 Dec 2024 · 2 repositories · arXiv:2412.17165
-
TAR3D: Creating High-Quality 3D Assets via Next-Part Prediction 22 Dec 2024 · 0 repositories · arXiv:2412.16919
-
AlzheimerRAG: Multimodal Retrieval Augmented Generation for PubMed articles 21 Dec 2024 · 0 repositories · arXiv:2412.16701
-
Assessing Social Alignment: Do Personality-Prompted Large Language Models Behave Like Humans? 21 Dec 2024 · 0 repositories · arXiv:2412.16772
-
Evaluating the Performance of Large Language Models in Scientific Claim Detection and Classification 21 Dec 2024 · 0 repositories · arXiv:2412.16486
-
Flash3D: Super-scaling Point Transformers through Joint Hardware-Geometry Locality 21 Dec 2024 · 1 repository · arXiv:2412.16481
-
Formal Language Knowledge Corpus for Retrieval Augmented Generation 21 Dec 2024 · 0 repositories · arXiv:2412.16689
-
From Histopathology Images to Cell Clouds: Learning Slide Representations with Hierarchical Cell Transformer 21 Dec 2024 · 0 repositories · arXiv:2412.16715
-
Identifying Cyberbullying Roles in Social Media 21 Dec 2024 · 0 repositories · arXiv:2412.16417
-
Improving FIM Code Completions via Context & Curriculum Based Learning 21 Dec 2024 · 0 repositories · arXiv:2412.16589
-
Lillama: Large Language Models Compression via Low-Rank Feature Distillation 21 Dec 2024 · 0 repositories · arXiv:2412.16719
-
Object Detection Approaches to Identifying Hand Images with High Forensic Values 21 Dec 2024 · 0 repositories · arXiv:2412.16431
-
Paraformer: Parameterization of Sub-grid Scale Processes Using Transformers 21 Dec 2024 · 0 repositories · arXiv:2412.16763
-
STKDRec: Spatial-Temporal Knowledge Distillation for Takeaway Recommendation 21 Dec 2024 · 1 repository · arXiv:2412.16502
-
TimeRAG: BOOSTING LLM Time Series Forecasting via Retrieval-Augmented Generation 21 Dec 2024 · 0 repositories · arXiv:2412.16643
-
Towards More Robust Retrieval-Augmented Generation: Evaluating RAG Under Adversarial Poisoning Attacks 21 Dec 2024 · 1 repository · arXiv:2412.16708
-
VSFormer: Value and Shape-Aware Transformer with Prior-Enhanced Self-Attention for Multivariate Time Series Classification 21 Dec 2024 · 0 repositories · arXiv:2412.16515
-
Adaptable and Precise: Enterprise-Scenario LLM Function-Calling Capability Training Pipeline 20 Dec 2024 · 0 repositories · arXiv:2412.15660
-
Adversarial Robustness through Dynamic Ensemble Learning 20 Dec 2024 · 0 repositories · arXiv:2412.16254
-
Benchmarking LLMs and SLMs for patient reported outcomes 20 Dec 2024 · 0 repositories · arXiv:2412.16291
-
Can LLMs Obfuscate Code? A Systematic Analysis of Large Language Models into Assembly Code Obfuscation 20 Dec 2024 · 0 repositories · arXiv:2412.16135
-
Demystifying the Potential of ChatGPT-4 Vision for Construction Progress Monitoring 20 Dec 2024 · 0 repositories · arXiv:2412.16108
-
Don't Do RAG: When Cache-Augmented Generation is All You Need for Knowledge Tasks 20 Dec 2024 · 1 repository · arXiv:2412.15605Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Explainable AI for Multivariate Time Series Pattern Exploration: Latent Space Visual Analytics with Temporal Fusion Transformer and Variational Autoencoders in Power Grid Event Diagnosis 20 Dec 2024 · 0 repositories · arXiv:2412.16098
-
Human-Readable Adversarial Prompts: An Investigation into LLM Vulnerabilities Using Situational Context 20 Dec 2024 · 0 repositories · arXiv:2412.16359
-
Humanlike Cognitive Patterns as Emergent Phenomena in Large Language Models 20 Dec 2024 · 0 repositories · arXiv:2412.15501
-
HybGRAG: Hybrid Retrieval-Augmented Generation on Textual and Relational Knowledge Bases 20 Dec 2024 · 0 repositories · arXiv:2412.16311
-
Linguistic Features Extracted by GPT-4 Improve Alzheimer's Disease Detection based on Spontaneous Speech 20 Dec 2024 · 1 repository · arXiv:2412.15772
-
Multi-dimensional Visual Prompt Enhanced Image Restoration via Mamba-Transformer Aggregation 20 Dec 2024 · 1 repository · arXiv:2412.15845