Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 46
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 46 of 190: papers 4,501 to 4,600 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Alfie: Democratising RGBA Image Generation With No $$$ 27 Aug 2024 · 2 repositories · arXiv:2408.14826
-
Applying ViT in Generalized Few-shot Semantic Segmentation 27 Aug 2024 · 1 repository · arXiv:2408.14957
-
FastTextSpotter: A High-Efficiency Transformer for Multilingual Scene Text Spotting 27 Aug 2024 · 1 repository · arXiv:2408.14998
-
GSIFN: A Graph-Structured and Interlaced-Masked Multimodal Transformer-based Fusion Network for Multimodal Sentiment Analysis 27 Aug 2024 · 1 repository · arXiv:2408.14809
-
Into the Unknown Unknowns: Engaged Human Learning through Participation in Language Model Agent Conversations 27 Aug 2024 · 1 repository · arXiv:2408.15232
-
MMASD+: A Novel Dataset for Privacy-Preserving Behavior Analysis of Children with Autism Spectrum Disorder 27 Aug 2024 · 1 repository · arXiv:2408.15077
-
Multi-Modal Instruction-Tuning Small-Scale Language-and-Vision Assistant for Semiconductor Electron Micrograph Analysis 27 Aug 2024 · 0 repositories · arXiv:2409.07463
-
Negation Blindness in Large Language Models: Unveiling the NO Syndrome in Image Generation 27 Aug 2024 · 0 repositories · arXiv:2409.00105
-
Parameter-Efficient Quantized Mixture-of-Experts Meets Vision-Language Instruction Tuning for Semiconductor Electron Micrograph Analysis 27 Aug 2024 · 0 repositories · arXiv:2408.15305
-
Strategic Optimization and Challenges of Large Language Models in Object-Oriented Programming 27 Aug 2024 · 0 repositories · arXiv:2408.14834
-
The Mamba in the Llama: Distilling and Accelerating Hybrid Models 27 Aug 2024 · 2 repositories · arXiv:2408.15237Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 2 where Syntology's instrument failed) · 5 unverified (of 14 harvested samples) · 5 pointer-only (licence)
-
The Uniqueness of LLaMA3-70B Series with Per-Channel Quantization 27 Aug 2024 · 0 repositories · arXiv:2408.15301
-
TourSynbio: A Multi-Modal Large Model and Agent Framework to Bridge Text and Protein Sequences for Protein Engineering 27 Aug 2024 · 1 repository · arXiv:2408.15299
-
Zero-Shot Visual Reasoning by Vision-Language Models: Benchmarking and Analysis 27 Aug 2024 · 0 repositories · arXiv:2409.00106
-
Beyond Detection: Leveraging Large Language Models for Cyber Attack Prediction in IoT Networks 26 Aug 2024 · 0 repositories · arXiv:2408.14045
-
BreakNet: Discontinuity-Resilient Multi-Scale Transformer Segmentation of Retinal Layers 26 Aug 2024 · 0 repositories · arXiv:2408.14606
-
CHARTOM: A Visual Theory-of-Mind Benchmark for Multimodal Large Language Models 26 Aug 2024 · 1 repository · arXiv:2408.14419
-
MSFMamba: Multi-Scale Feature Fusion State Space Model for Multi-Source Remote Sensing Image Classification 26 Aug 2024 · 2 repositories · arXiv:2408.14255
-
Probing Causality Manipulation of Large Language Models 26 Aug 2024 · 0 repositories · arXiv:2408.14380
-
Rate-Distortion-Perception Controllable Joint Source-Channel Coding for High-Fidelity Generative Communications 26 Aug 2024 · 0 repositories · arXiv:2408.14127
-
3D-RCNet: Learning from Transformer to Build a 3D Relational ConvNet for Hyperspectral Image Classification 25 Aug 2024 · 1 repository · arXiv:2408.13728
-
Bidirectional Awareness Induction in Autoregressive Seq2Seq Models 25 Aug 2024 · 0 repositories · arXiv:2408.13959
-
CNN-Transformer Rectified Collaborative Learning for Medical Image Segmentation 25 Aug 2024 · 0 repositories · arXiv:2408.13698
-
CodeGraph: Enhancing Graph Reasoning of LLMs with Code 25 Aug 2024 · 1 repository · arXiv:2408.13863
-
Extremely Fine-Grained Visual Classification over Resembling Glyphs in the Wild 25 Aug 2024 · 1 repository · arXiv:2408.13774
-
AlphaViT: A Flexible Game-Playing AI for Multiple Games and Variable Board Sizes 25 Aug 2024 · 1 repository · arXiv:2408.13871
-
LowCLIP: Adapting the CLIP Model Architecture for Low-Resource Languages in Multimodal Image Retrieval Task 25 Aug 2024 · 0 repositories · arXiv:2408.13909
-
Shifted Window Fourier Transform And Retention For Image Captioning 25 Aug 2024 · 0 repositories · arXiv:2408.13963
-
Vision-Language and Large Language Model Performance in Gastroenterology: GPT, Claude, Llama, Phi, Mistral, Gemma, and Quantized Models 25 Aug 2024 · 1 repository · arXiv:2409.00084
-
A Law of Next-Token Prediction in Large Language Models 24 Aug 2024 · 1 repository · arXiv:2408.13442Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Prompt-Matcher: Leveraging Large Models to Reduce Uncertainty in Schema Matching Results 24 Aug 2024 · 0 repositories · arXiv:2408.14507
-
Integrating Multi-Head Convolutional Encoders with Cross-Attention for Improved SPARQL Query Translation 24 Aug 2024 · 0 repositories · arXiv:2408.13432
-
IQA-EVAL: Automatic Evaluation of Human-Model Interactive Question Answering 24 Aug 2024 · 0 repositories · arXiv:2408.13545
-
Pandora's Box or Aladdin's Lamp: A Comprehensive Analysis Revealing the Role of RAG Noise in Large Language Models 24 Aug 2024 · 1 repository · arXiv:2408.13533Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Probing the Robustness of Vision-Language Pretrained Models: A Multimodal Adversarial Attack Approach 24 Aug 2024 · 0 repositories · arXiv:2408.13461
-
Rethinking Video Deblurring with Wavelet-Aware Dynamic Transformer and Diffusion Model 24 Aug 2024 · 1 repository · arXiv:2408.13459Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 3 pointer-only (licence)
-
Topological GCN for Improving Detection of Hip Landmarks from B-Mode Ultrasound Images 24 Aug 2024 · 0 repositories · arXiv:2408.13495
-
Utilizing Large Language Models for Named Entity Recognition in Traditional Chinese Medicine against COVID-19 Literature: Comparative Study 24 Aug 2024 · 0 repositories · arXiv:2408.13501
-
Accuracy Improvement of Cell Image Segmentation Using Feedback Former 23 Aug 2024 · 0 repositories · arXiv:2408.12974
-
An In-Depth Investigation of Data Collection in LLM App Ecosystems 23 Aug 2024 · 0 repositories · arXiv:2408.13247
-
EAViT: External Attention Vision Transformer for Audio Classification 23 Aug 2024 · 0 repositories · arXiv:2408.13201
-
Knowledge Graph Modeling-Driven Large Language Model Operating System (LLM OS) for Task Automation in Process Engineering Problem-Solving 23 Aug 2024 · 0 repositories · arXiv:2408.14494
-
QD-VMR: Query Debiasing with Contextual Understanding Enhancement for Video Moment Retrieval 23 Aug 2024 · 0 repositories · arXiv:2408.12981
-
State-of-the-Art Fails in the Art of Damage Detection 23 Aug 2024 · 0 repositories · arXiv:2408.12953
-
Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates 23 Aug 2024 · 1 repository · arXiv:2408.13006Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
VFM-Det: Towards High-Performance Vehicle Detection via Large Foundation Models 23 Aug 2024 · 1 repository · arXiv:2408.13031
-
PDDFormer: Pairwise Distance Distribution Graph Transformer for Crystal Material Property Prediction 23 Aug 2024 · 0 repositories · arXiv:2408.12984
-
AI-driven Transformer Model for Fault Prediction in Non-Linear Dynamic Automotive System 22 Aug 2024 · 0 repositories · arXiv:2408.12638
-
Can LLMs Understand Social Norms in Autonomous Driving Games? 22 Aug 2024 · 0 repositories · arXiv:2408.12680
-
Enhanced Infield Agriculture with Interpretable Machine Learning Approaches for Crop Classification 22 Aug 2024 · 0 repositories · arXiv:2408.12426
-
Enhancing Automated Program Repair with Solution Design 22 Aug 2024 · 0 repositories · arXiv:2408.12056
-
Enhancing Multi-hop Reasoning through Knowledge Erasure in Large Language Model Editing 22 Aug 2024 · 0 repositories · arXiv:2408.12456
-
GRATR: Zero-Shot Evidence Graph Retrieval-Augmented Trustworthiness Reasoning 22 Aug 2024 · 1 repository · arXiv:2408.12333
-
Jamba-1.5: Hybrid Transformer-Mamba Models at Scale 22 Aug 2024 · 2 repositories · arXiv:2408.12570
-
Large Language Models as Foundations for Next-Gen Dense Retrieval: A Comprehensive Empirical Assessment 22 Aug 2024 · 0 repositories · arXiv:2408.12194
-
LLMs are not Zero-Shot Reasoners for Biomedical Information Extraction 22 Aug 2024 · 0 repositories · arXiv:2408.12249
-
MedDiT: A Knowledge-Controlled Diffusion Transformer Framework for Dynamic Medical Image Generation in Virtual Simulated Patient 22 Aug 2024 · 0 repositories · arXiv:2408.12236
-
Optimizing Performance: How Compact Models Match or Exceed GPT's Classification Capabilities through Fine-Tuning 22 Aug 2024 · 0 repositories · arXiv:2409.11408
-
RuleAlign: Making Large Language Models Better Physicians with Diagnostic Rule Alignment 22 Aug 2024 · 0 repositories · arXiv:2408.12579
-
Towards Evaluating and Building Versatile Large Language Models for Medicine 22 Aug 2024 · 1 repository · arXiv:2408.12547Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Transformers As Approximations of Solomonoff Induction 22 Aug 2024 · 0 repositories · arXiv:2408.12065
-
xGen-VideoSyn-1: High-fidelity Text-to-Video Synthesis with Compressed Representations 22 Aug 2024 · 0 repositories · arXiv:2408.12590
-
A Quick, trustworthy spectral knowledge Q&A system leveraging retrieval-augmented generation on LLM 21 Aug 2024 · 1 repository · arXiv:2408.11557
-
Ancient Wisdom, Modern Tools: Exploring Retrieval-Augmented LLMs for Ancient Indian Philosophy 21 Aug 2024 · 1 repository · arXiv:2408.11903
-
Applying and Evaluating Large Language Models in Mental Health Care: A Scoping Review of Human-Assessed Generative Tasks 21 Aug 2024 · 0 repositories · arXiv:2408.11288
-
BURExtract-Llama: An LLM for Clinical Concept Extraction in Breast Ultrasound Reports 21 Aug 2024 · 0 repositories · arXiv:2408.11334
-
D-RMGPT: Robot-assisted collaborative tasks driven by large multimodal models 21 Aug 2024 · 0 repositories · arXiv:2408.11761
-
A Benchmark for AI-based Weather Data Assimilation 21 Aug 2024 · 1 repository · arXiv:2408.11438
-
Distributional Properties of Subword Regularization 21 Aug 2024 · 0 repositories · arXiv:2408.11443
-
DocTabQA: Answering Questions from Long Documents Using Tables 21 Aug 2024 · 1 repository · arXiv:2408.11490
-
Energy Estimation of Last Mile Electric Vehicle Routes 21 Aug 2024 · 0 repositories · arXiv:2408.12006
-
Exploring Large Language Models for Feature Selection: A Data-centric Perspective 21 Aug 2024 · 0 repositories · arXiv:2408.12025
-
FATE: Focal-modulated Attention Encoder for Temperature Prediction 21 Aug 2024 · 1 repository · arXiv:2408.11336
-
HMT-UNet: A hybird Mamba-Transformer Vision UNet for Medical Image Segmentation 21 Aug 2024 · 1 repository · arXiv:2408.11289
-
Leveraging Fine-Tuned Retrieval-Augmented Generation with Long-Context Support: For 3GPP Standards 21 Aug 2024 · 1 repository · arXiv:2408.11775
-
Macformer: Transformer with Random Maclaurin Feature Attention 21 Aug 2024 · 0 repositories · arXiv:2408.11656
-
Mixed Sparsity Training: Achieving 4× FLOP Reduction for Transformer Pretraining 21 Aug 2024 · 0 repositories · arXiv:2408.11746
-
OAPT: Offset-Aware Partition Transformer for Double JPEG Artifacts Removal 21 Aug 2024 · 1 repository · arXiv:2408.11480Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
WeQA: A Benchmark for Retrieval Augmented Generation in Wind Energy Domain 21 Aug 2024 · 0 repositories · arXiv:2408.11800
-
Positional Prompt Tuning for Efficient 3D Representation Learning 21 Aug 2024 · 1 repository · arXiv:2408.11567
-
RAG-Optimized Tibetan Tourism LLMs: Enhancing Accuracy and Personalization 21 Aug 2024 · 0 repositories · arXiv:2408.12003
-
RAGLAB: A Modular and Research-Oriented Unified Framework for Retrieval-Augmented Generation 21 Aug 2024 · 1 repository · arXiv:2408.11381
-
SarcasmBench: Towards Evaluating Large Language Models on Sarcasm Understanding 21 Aug 2024 · 0 repositories · arXiv:2408.11319
-
Unlocking Adversarial Suffix Optimization Without Affirmative Phrases: Efficient Black-box Jailbreaking via LLM as Optimizer 21 Aug 2024 · 1 repository · arXiv:2408.11313Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Classification of Endoscopy and Video Capsule Images using CNN-Transformer Model 20 Aug 2024 · 0 repositories · arXiv:2408.10733
-
Crafting Tomorrow's Headlines: Neural News Generation and Detection in English, Turkish, Hungarian, and Persian 20 Aug 2024 · 0 repositories · arXiv:2408.10724
-
CTP-LLM: Clinical Trial Phase Transition Prediction Using Large Language Models 20 Aug 2024 · 0 repositories · arXiv:2408.10995
-
Dr.Academy: A Benchmark for Evaluating Questioning Capability in Education for Large Language Models 20 Aug 2024 · 0 repositories · arXiv:2408.10947
-
EdgeNAT: Transformer for Efficient Edge Detection 20 Aug 2024 · 1 repository · arXiv:2408.10527
-
How Well Do Large Language Models Serve as End-to-End Secure Code Agents for Python? 20 Aug 2024 · 0 repositories · arXiv:2408.10495
-
Integrating Multi-Modal Input Token Mixer Into Mamba-Based Decision Models: Decision MetaMamba 20 Aug 2024 · 0 repositories · arXiv:2408.10517
-
Language Modeling on Tabular Data: A Survey of Foundations, Techniques and Evolution 20 Aug 2024 · 1 repository · arXiv:2408.10548
-
MambaEVT: Event Stream based Visual Object Tracking using State Space Model 20 Aug 2024 · 1 repository · arXiv:2408.10487Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 1 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 10 harvested samples)
-
Navigating Spatio-Temporal Heterogeneity: A Graph Transformer Approach for Traffic Forecasting 20 Aug 2024 · 1 repository · arXiv:2408.10822
-
Open-FinLLMs: Open Multimodal Large Language Models for Financial Applications 20 Aug 2024 · 0 repositories · arXiv:2408.11878
-
Out-of-Distribution Detection with Attention Head Masking for Multimodal Document Classification 20 Aug 2024 · 1 repository · arXiv:2408.11237
-
PRformer: Pyramidal Recurrent Transformer for Multivariate Time Series Forecasting 20 Aug 2024 · 1 repository · arXiv:2408.10483Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Quantum Inverse Contextual Vision Transformers (Q-ICVT): A New Frontier in 3D Object Detection for AVs 20 Aug 2024 · 1 repository · arXiv:2408.11207
-
Reading with Intent 20 Aug 2024 · 0 repositories · arXiv:2408.11189
-
Reconciling Methodological Paradigms: Employing Large Language Models as Novice Qualitative Research Assistants in Talent Management Research 20 Aug 2024 · 0 repositories · arXiv:2408.11043