Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 75
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 75 of 190: papers 7,401 to 7,500 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Narrating Causal Graphs with Large Language Models 11 Mar 2024 · 0 repositories · arXiv:2403.07118
-
SMART: Automatically Scaling Down Language Models with Accuracy Guarantees for Reduced Processing Fees 11 Mar 2024 · 1 repository · arXiv:2403.13835
-
The pitfalls of next-token prediction 11 Mar 2024 · 1 repository · arXiv:2403.06963Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Unraveling the Mystery of Scaling Laws: Part I 11 Mar 2024 · 0 repositories · arXiv:2403.06563
-
Attacking Transformers with Feature Diversity Adversarial Perturbation 10 Mar 2024 · 0 repositories · arXiv:2403.07942
-
Finding Visual Saliency in Continuous Spike Stream 10 Mar 2024 · 1 repository · arXiv:2403.06233Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 2 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
FrameQuant: Flexible Low-Bit Quantization for Transformers 10 Mar 2024 · 1 repository · arXiv:2403.06082Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 1 where Syntology's instrument failed) · 6 unverified (of 18 harvested samples) · 18 pointer-only (licence)
-
Target-constrained Bidirectional Planning for Generation of Target-oriented Proactive Dialogue 10 Mar 2024 · 1 repository · arXiv:2403.06063
-
Towards In-Vehicle Multi-Task Facial Attribute Recognition: Investigating Synthetic Data and Vision Foundation Models 10 Mar 2024 · 0 repositories · arXiv:2403.06088
-
Unpacking Tokenization: Evaluating Text Compression and its Correlation with Model Performance 10 Mar 2024 · 0 repositories · arXiv:2403.06265
-
AutoEval Done Right: Using Synthetic Data for Model Evaluation 9 Mar 2024 · 1 repository · arXiv:2403.07008Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
ClinicalMamba: A Generative Clinical Language Model on Longitudinal Clinical Notes 9 Mar 2024 · 1 repository · arXiv:2403.05795Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 1 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples)
-
Long-term Frame-Event Visual Tracking: Benchmark Dataset and Baseline 9 Mar 2024 · 4 repositories · arXiv:2403.05839
-
Segmentation Guided Sparse Transformer for Under-Display Camera Image Restoration 9 Mar 2024 · 0 repositories · arXiv:2403.05906
-
A Dataset and Benchmark for Hospital Course Summarization with Adapted Large Language Models 8 Mar 2024 · 1 repository · arXiv:2403.05720
-
A Novel Nuanced Conversation Evaluation Framework for Large Language Models in Mental Health 8 Mar 2024 · 0 repositories · arXiv:2403.09705
-
ActFormer: Scalable Collaborative Perception via Active Queries 8 Mar 2024 · 0 repositories · arXiv:2403.04968
-
An In-depth Evaluation of GPT-4 in Sentence Simplification with Error-based Human Assessment 8 Mar 2024 · 0 repositories · arXiv:2403.04963
-
Are Large Language Models Aligned with People's Social Intuitions for Human-Robot Interactions? 8 Mar 2024 · 1 repository · arXiv:2403.05701
-
Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought 8 Mar 2024 · 1 repository · arXiv:2403.05518Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 1 pointer-only (licence)
-
Can't Remember Details in Long Documents? You Need Some R&R 8 Mar 2024 · 1 repository · arXiv:2403.05004Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
CommitBench: A Benchmark for Commit Message Generation 8 Mar 2024 · 1 repository · arXiv:2403.05188
-
Considering Nonstationary within Multivariate Time Series with Variational Hierarchical Transformer for Forecasting 8 Mar 2024 · 1 repository · arXiv:2403.05406Syntology official (archive's flag): 12 ran · 12 ran (of which 10 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 17 harvested samples) · 17 pointer-only (licence)
-
Cost-Performance Optimization for Processing Low-Resource Language Tasks Using Commercial LLMs 8 Mar 2024 · 0 repositories · arXiv:2403.05434
-
How Well Do Multi-modal LLMs Interpret CT Scans? An Auto-Evaluation Framework for Analyses 8 Mar 2024 · 0 repositories · arXiv:2403.05680
-
Denoising Autoregressive Representation Learning 8 Mar 2024 · 0 repositories · arXiv:2403.05196
-
DualBEV: Unifying Dual View Transformation with Probabilistic Correspondences 8 Mar 2024 · 1 repository · arXiv:2403.05402Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Enhancing Automatic Modulation Recognition for IoT Applications Using Transformers 8 Mar 2024 · 0 repositories · arXiv:2403.15417
-
ERBench: An Entity-Relationship based Automatically Verifiable Hallucination Benchmark for Large Language Models 8 Mar 2024 · 1 repository · arXiv:2403.05266Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context 8 Mar 2024 · 1 repository · arXiv:2403.05530
-
Inverse Design of Photonic Crystal Surface Emitting Lasers is a Sequence Modeling Problem 8 Mar 2024 · 0 repositories · arXiv:2403.05149
-
JointMotion: Joint Self-Supervision for Joint Motion Prediction 8 Mar 2024 · 1 repository · arXiv:2403.05489
-
LightM-UNet: Mamba Assists in Lightweight UNet for Medical Image Segmentation 8 Mar 2024 · 1 repository · arXiv:2403.05246Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 2 pointer-only (licence)
-
LLM4Decompile: Decompiling Binary Code with Large Language Models 8 Mar 2024 · 1 repository · arXiv:2403.05286Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 13 harvested samples) · 1 pointer-only (licence)
-
MamMIL: Multiple Instance Learning for Whole Slide Images with State Space Models 8 Mar 2024 · 1 repository · arXiv:2403.05160
-
Med3DInsight: Enhancing 3D Medical Image Understanding with 2D Multi-Modal Large Language Models 8 Mar 2024 · 0 repositories · arXiv:2403.05141
-
PipeRAG: Fast Retrieval-Augmented Generation via Algorithm-System Co-design 8 Mar 2024 · 0 repositories · arXiv:2403.05676
-
RAT: Retrieval Augmented Thoughts Elicit Context-Aware Reasoning in Long-Horizon Generation 8 Mar 2024 · 1 repository · arXiv:2403.05313
-
Rule-driven News Captioning 8 Mar 2024 · 0 repositories · arXiv:2403.05101
-
SIRST-5K: Exploring Massive Negatives Synthesis with Self-supervised Learning for Robust Infrared Small Target Detection 8 Mar 2024 · 1 repository · arXiv:2403.05416
-
Spatial-aware Transformer-GRU Framework for Enhanced Glaucoma Diagnosis from 3D OCT Imaging 8 Mar 2024 · 1 repository · arXiv:2403.05702
-
Will GPT-4 Run DOOM? 8 Mar 2024 · 0 repositories · arXiv:2403.05468
-
Aligning GPTRec with Beyond-Accuracy Goals with Reinforcement Learning 7 Mar 2024 · 1 repository · arXiv:2403.04875
-
AO-DETR: Anti-Overlapping DETR for X-Ray Prohibited Items Detection 7 Mar 2024 · 1 repository · arXiv:2403.04309
-
AUFormer: Vision Transformers are Parameter-Efficient Facial Action Unit Detectors 7 Mar 2024 · 1 repository · arXiv:2403.04697
-
Disentangled Diffusion-Based 3D Human Pose Estimation with Hierarchical Spatial and Temporal Denoiser 7 Mar 2024 · 1 repository · arXiv:2403.04444
-
Federated Recommendation via Hybrid Retrieval Augmented Generation 7 Mar 2024 · 1 repository · arXiv:2403.04256
-
Feedback-Generation for Programming Exercises With GPT-4 7 Mar 2024 · 0 repositories · arXiv:2403.04449
-
HaluEval-Wild: Evaluating Hallucinations of Language Models in the Wild 7 Mar 2024 · 1 repository · arXiv:2403.04307
-
LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error 7 Mar 2024 · 1 repository · arXiv:2403.04746Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 7 unverified (of 13 harvested samples) · 1 pointer-only (licence)
-
PixArt-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation 7 Mar 2024 · 2 repositories · arXiv:2403.04692Syntology official (archive's flag): 5 ran · 6 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples) · 6 pointer-only (licence)
-
RATSF: Empowering Customer Service Volume Management through Retrieval-Augmented Time-Series Forecasting 7 Mar 2024 · 0 repositories · arXiv:2403.04180
-
Speech Emotion Recognition Via CNN-Transformer and Multidimensional Attention Mechanism 7 Mar 2024 · 1 repository · arXiv:2403.04743
-
Telecom Language Models: Must They Be Large? 7 Mar 2024 · 0 repositories · arXiv:2403.04666
-
Assessing the Aesthetic Evaluation Capabilities of GPT-4 with Vision: Insights from Group and Individual Assessments 6 Mar 2024 · 0 repositories · arXiv:2403.03594
-
Benchmarking Hallucination in Large Language Models based on Unanswerable Math Word Problem 6 Mar 2024 · 1 repository · arXiv:2403.03558
-
Can Large Language Models do Analytical Reasoning? 6 Mar 2024 · 0 repositories · arXiv:2403.04031
-
Design of an Open-Source Architecture for Neural Machine Translation 6 Mar 2024 · 0 repositories · arXiv:2403.03582
-
Designing Informative Metrics for Few-Shot Example Selection 6 Mar 2024 · 0 repositories · arXiv:2403.03861
-
Enhancing Price Prediction in Cryptocurrency Using Transformer Neural Network and Technical Indicators 6 Mar 2024 · 0 repositories · arXiv:2403.03606
-
FaaF: Facts as a Function for the evaluation of generated text 6 Mar 2024 · 1 repository · arXiv:2403.03888
-
General2Specialized LLMs Translation for E-commerce 6 Mar 2024 · 0 repositories · arXiv:2403.03689
-
Guiding Enumerative Program Synthesis with Large Language Models 6 Mar 2024 · 0 repositories · arXiv:2403.03997
-
Inverse-Free Fast Natural Gradient Descent Method for Deep Learning 6 Mar 2024 · 0 repositories · arXiv:2403.03473
-
Japanese-English Sentence Translation Exercises Dataset for Automatic Grading 6 Mar 2024 · 0 repositories · arXiv:2403.03396
-
Joint multi-task learning improves weakly-supervised biomarker prediction in computational pathology 6 Mar 2024 · 1 repository · arXiv:2403.03891
-
Multi-modal Deep Learning 6 Mar 2024 · 0 repositories · arXiv:2403.03385
-
PPTC-R benchmark: Towards Evaluating the Robustness of Large Language Models for PowerPoint Task Completion 6 Mar 2024 · 1 repository · arXiv:2403.03788
-
Rapidly Developing High-quality Instruction Data and Evaluation Benchmark for Large Language Models with Minimal Human Effort: A Case Study on Japanese 6 Mar 2024 · 2 repositories · arXiv:2403.03690
-
AI Insights: A Case Study on Utilizing ChatGPT Intelligence for Research Paper Analysis 5 Mar 2024 · 0 repositories · arXiv:2403.03293
-
An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4 5 Mar 2024 · 1 repository · arXiv:2403.02839Syntology official (archive's flag): 15 ran · 15 ran (of which 0 constructed an object rather than computing a result; 15 with no instrument failure: 0 honoured, 0 violated, 15 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 18 harvested samples) · 18 pointer-only (licence)
-
ARNN: Attentive Recurrent Neural Network for Multi-channel EEG Signals to Identify Epileptic Seizures 5 Mar 2024 · 1 repository · arXiv:2403.03276
-
Behavior Generation with Latent Actions 5 Mar 2024 · 2 repositories · arXiv:2403.03181Syntology official (archive's flag): 7 ran · 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 1 honoured, 2 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 6 unverified (of 14 harvested samples) · 2 pointer-only (licence)
-
CLEVR-POC: Reasoning-Intensive Visual Question Answering in Partially Observable Environments 5 Mar 2024 · 0 repositories · arXiv:2403.03203
-
Drug Resistance Predictions Based on a Directed Flag Transformer 5 Mar 2024 · 0 repositories · arXiv:2403.02603
-
Emerging Synergies Between Large Language Models and Machine Learning in Ecommerce Recommendations 5 Mar 2024 · 0 repositories · arXiv:2403.02760
-
Enhancing Weakly Supervised 3D Medical Image Segmentation through Probabilistic-aware Learning 5 Mar 2024 · 1 repository · arXiv:2403.02566
-
Evaluating and Optimizing Educational Content with Large Language Model Judgments 5 Mar 2024 · 1 repository · arXiv:2403.02795Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Evolution Transformer: In-Context Evolutionary Optimization 5 Mar 2024 · 1 repository · arXiv:2403.02985
-
Exploring Naive Approaches to Tell Apart LLMs Productions from Human-written Text 5 Mar 2024 · 1 repository
-
FAR: Flexible, Accurate and Robust 6DoF Relative Camera Pose Estimation 5 Mar 2024 · 0 repositories · arXiv:2403.03221
-
How Well Can Transformers Emulate In-context Newton's Method? 5 Mar 2024 · 0 repositories · arXiv:2403.03183
-
Improving Event Definition Following For Zero-Shot Event Detection 5 Mar 2024 · 0 repositories · arXiv:2403.02586
-
InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents 5 Mar 2024 · 2 repositories · arXiv:2403.02691Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
InjectTST: A Transformer Method of Injecting Global Information into Independent Channels for Long Time Series Forecasting 5 Mar 2024 · 0 repositories · arXiv:2403.02814
-
JMI at SemEval 2024 Task 3: Two-step approach for multimodal ECAC using in-context learning with GPT and instruction-tuned Llama models 5 Mar 2024 · 1 repository · arXiv:2403.04798Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 7 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
Knowledge Graphs as Context Sources for LLM-Based Explanations of Learning Recommendations 5 Mar 2024 · 0 repositories · arXiv:2403.03008
-
Language Guided Exploration for RL Agents in Text Environments 5 Mar 2024 · 0 repositories · arXiv:2403.03141
-
Learning without Exact Guidance: Updating Large-scale High-resolution Land Cover Maps from Low-resolution Historical Labels 5 Mar 2024 · 3 repositories · arXiv:2403.02746Syntology official (archive's flag): 2 ran · 5 ran (of which 2 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 1 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 1 pointer-only (licence)
-
MathScale: Scaling Instruction Tuning for Mathematical Reasoning 5 Mar 2024 · 1 repository · arXiv:2403.02884
-
MiKASA: Multi-Key-Anchor & Scene-Aware Transformer for 3D Visual Grounding 5 Mar 2024 · 1 repository · arXiv:2403.03077Syntology official (archive's flag): 11 ran · 11 ran (of which 10 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 16 harvested samples) · 16 pointer-only (licence)
-
PARADISE: Evaluating Implicit Planning Skills of Language Models with Procedural Warnings and Tips Dataset 5 Mar 2024 · 1 repository · arXiv:2403.03167
-
MeanCache: User-Centric Semantic Caching for LLM Web Services 5 Mar 2024 · 0 repositories · arXiv:2403.02694
-
Scope of Large Language Models for Mining Emerging Opinions in Online Health Discourse 5 Mar 2024 · 0 repositories · arXiv:2403.03336
-
SNIFFER: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection 5 Mar 2024 · 0 repositories · arXiv:2403.03170
-
Towards Democratized Flood Risk Management: An Advanced AI Assistant Enabled by GPT-4 for Enhanced Interpretability and Public Engagement 5 Mar 2024 · 2 repositories · arXiv:2403.03188
-
Towards Training A Chinese Large Language Model for Anesthesiology 5 Mar 2024 · 0 repositories · arXiv:2403.02742
-
Zero-Shot Cross-Lingual Document-Level Event Causality Identification with Heterogeneous Graph Contrastive Transfer Learning 5 Mar 2024 · 0 repositories · arXiv:2403.02893
-
A Spatio-temporal Aligned SUNet Model for Low-light Video Enhancement 4 Mar 2024 · 0 repositories · arXiv:2403.02408
-
adaptNMT: an open-source, language-agnostic development environment for Neural Machine Translation 4 Mar 2024 · 0 repositories · arXiv:2403.02367