Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 90
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 90 of 190: papers 8,901 to 9,000 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Making Translators Privacy-aware on the User's Side 7 Dec 2023 · 0 repositories · arXiv:2312.04068
-
On Sarcasm Detection with OpenAI GPT-based Models 7 Dec 2023 · 0 repositories · arXiv:2312.04642
-
Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models 7 Dec 2023 · 0 repositories · arXiv:2312.04724
-
Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos 7 Dec 2023 · 2 repositories · arXiv:2312.04746Syntology 11 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 1 violated, 7 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 12 harvested samples)
-
A Text-to-Text Model for Multilingual Offensive Language Identification 6 Dec 2023 · 0 repositories · arXiv:2312.03379
-
Automatic Transcription of Handwritten Old Occitan Language 6 Dec 2023 · 1 repository
-
Blueprinting the Future: Automatic Item Categorization using Hierarchical Zero-Shot and Few-Shot Classifiers 6 Dec 2023 · 0 repositories · arXiv:2312.03561
-
Comparative Analysis of Multilingual Text Classification & Identification through Deep Learning and Embedding Visualization 6 Dec 2023 · 0 repositories · arXiv:2312.03789
-
Compressed Context Memory For Online Language Model Interaction 6 Dec 2023 · 1 repository · arXiv:2312.03414Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 4 where Syntology's instrument failed) · 4 unverified (of 13 harvested samples) · 1 pointer-only (licence)
-
GPT-4 Enhanced Multimodal Grounding for Autonomous Driving: Leveraging Cross-Modal Attention with Large Language Models 6 Dec 2023 · 1 repository · arXiv:2312.03543
-
Holmes: Towards Distributed Training Across Clusters with Heterogeneous NIC Environment 6 Dec 2023 · 0 repositories · arXiv:2312.03549
-
Interpretability Illusions in the Generalization of Simplified Models 6 Dec 2023 · 0 repositories · arXiv:2312.03656
-
Lite-Mind: Towards Efficient and Robust Brain Representation Network 6 Dec 2023 · 1 repository · arXiv:2312.03781Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Memory-Efficient Optical Flow via Radius-Distribution Orthogonal Cost Volume 6 Dec 2023 · 1 repository · arXiv:2312.03790
-
Exploring the Reversal Curse and Other Deductive Logical Reasoning in BERT and GPT-Based Large Language Models 6 Dec 2023 · 1 repository · arXiv:2312.03633
-
Parameter-Efficient Transfer Learning of Audio Spectrogram Transformers 6 Dec 2023 · 1 repository · arXiv:2312.03694
-
When an Image is Worth 1,024 x 1,024 Words: A Case Study in Computational Pathology 6 Dec 2023 · 0 repositories · arXiv:2312.03558
-
XAIQA: Explainer-Based Data Augmentation for Extractive Question Answering 6 Dec 2023 · 0 repositories · arXiv:2312.03567
-
A Comparative Study of AI-Generated (GPT-4) and Human-crafted MCQs in Programming Education 5 Dec 2023 · 0 repositories · arXiv:2312.03173
-
A Hardware Evaluation Framework for Large Language Model Inference 5 Dec 2023 · 0 repositories · arXiv:2312.03134
-
C3: High-performance and low-complexity neural compression from a single image or video 5 Dec 2023 · 1 repository · arXiv:2312.02753
-
Compositional Generalization for Data-to-Text Generation 5 Dec 2023 · 0 repositories · arXiv:2312.02748
-
DRAFT: Dense Retrieval Augmented Few-shot Topic classifier Framework 5 Dec 2023 · 1 repository · arXiv:2312.02532
-
GPT vs Human for Scientific Reviews: A Dual Source Review on Applications of ChatGPT in Science 5 Dec 2023 · 0 repositories · arXiv:2312.03769
-
Let the LLMs Talk: Simulating Human-to-Human Conversational QA via Zero-Shot LLM-to-LLM Interactions 5 Dec 2023 · 1 repository · arXiv:2312.02913
-
MEMTO: Memory-guided Transformer for Multivariate Time Series Anomaly Detection 5 Dec 2023 · 1 repository · arXiv:2312.02530Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
MIMONets: Multiple-Input-Multiple-Output Neural Networks Exploiting Computation in Superposition 5 Dec 2023 · 1 repository · arXiv:2312.02829Syntology official (archive's flag): 12 ran · 12 ran (of which 8 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 16 harvested samples) · 2 pointer-only (licence)
-
Rank-without-GPT: Building GPT-Independent Listwise Rerankers on Open-Source Large Language Models 5 Dec 2023 · 0 repositories · arXiv:2312.02969
-
RankZephyr: Effective and Robust Zero-Shot Listwise Reranking is a Breeze! 5 Dec 2023 · 2 repositories · arXiv:2312.02724
-
RotaTR: Detection Transformer for Dense and Rotated Object 5 Dec 2023 · 0 repositories · arXiv:2312.02821
-
Sample-based Dynamic Hierarchical Transformer with Layer and Head Flexibility via Contextual Bandit 5 Dec 2023 · 0 repositories · arXiv:2312.03038
-
UPOCR: Towards Unified Pixel-Level OCR Interface 5 Dec 2023 · 0 repositories · arXiv:2312.02694
-
A Comprehensive Literature Review on Sweet Orange Leaf Diseases 4 Dec 2023 · 0 repositories · arXiv:2312.01756
-
A Glitch in the Matrix? Locating and Detecting Language Model Grounding with Fakepedia 4 Dec 2023 · 1 repository · arXiv:2312.02073Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
A Machine Learning Approach Towards SKILL Code Autocompletion 4 Dec 2023 · 0 repositories · arXiv:2312.01921
-
A Survey on Large Language Model (LLM) Security and Privacy: The Good, the Bad, and the Ugly 4 Dec 2023 · 0 repositories · arXiv:2312.02003
-
Adapting Short-Term Transformers for Action Detection in Untrimmed Videos 4 Dec 2023 · 1 repository · arXiv:2312.01897
-
BEVNeXt: Reviving Dense BEV Frameworks for 3D Object Detection 4 Dec 2023 · 1 repository · arXiv:2312.01696Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Competition-Level Problems are Effective LLM Evaluators 4 Dec 2023 · 0 repositories · arXiv:2312.02143
-
CILF-CIAE: CLIP-driven Image-Language Fusion for Correcting Inverse Age Estimation 4 Dec 2023 · 0 repositories · arXiv:2312.01758
-
DiffiT: Diffusion Vision Transformers for Image Generation 4 Dec 2023 · 1 repository · arXiv:2312.02139Syntology official (archive's flag): 20 ran · 20 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 3 honoured, 1 violated, 9 with no contract checked; 7 where Syntology's instrument failed) · 5 unverified (of 25 harvested samples) · 25 pointer-only (licence)
-
EdgeConvFormer: Dynamic Graph CNN and Transformer based Anomaly Detection in Multivariate Time Series 4 Dec 2023 · 0 repositories · arXiv:2312.01729
-
Explore, Select, Derive, and Recall: Augmenting LLM with Human-like Memory for Mobile Task Automation 4 Dec 2023 · 0 repositories · arXiv:2312.03003
-
FaultFormer: Pretraining Transformers for Adaptable Bearing Fault Classification 4 Dec 2023 · 1 repository · arXiv:2312.02380
-
Fine-Tuning Language Models for Context-Specific SQL Query Generation 4 Dec 2023 · 0 repositories · arXiv:2312.02251
-
ImputeFormer: Low Rankness-Induced Transformers for Generalizable Spatiotemporal Imputation 4 Dec 2023 · 2 repositories · arXiv:2312.01728Syntology official: no sample here; runs from other or unrecorded repositories · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
InstructTA: Instruction-Tuned Targeted Attack for Large Vision-Language Models 4 Dec 2023 · 1 repository · arXiv:2312.01886Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 1 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
Jellyfish: A Large Language Model for Data Preprocessing 4 Dec 2023 · 0 repositories · arXiv:2312.01678
-
MobileUtr: Revisiting the relationship between light-weight CNN and Transformer for efficient medical image segmentation 4 Dec 2023 · 1 repository · arXiv:2312.01740
-
Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation 4 Dec 2023 · 4 repositories · arXiv:2312.02145Syntology official (archive's flag): 9 ran · 22 ran (of which 0 constructed an object rather than computing a result; 19 with no instrument failure: 0 honoured, 3 violated, 16 with no contract checked; 3 where Syntology's instrument failed) · 4 unverified (of 26 harvested samples) · 5 pointer-only (licence)
-
Rethinking Urban Mobility Prediction: A Super-Multivariate Time Series Forecasting Approach 4 Dec 2023 · 2 repositories · arXiv:2312.01699
-
Retrieval-augmented Multi-modal Chain-of-Thoughts Reasoning for Large Language Models 4 Dec 2023 · 0 repositories · arXiv:2312.01714
-
SARA-RT: Scaling up Robotics Transformers with Self-Adaptive Robust Attention 4 Dec 2023 · 0 repositories · arXiv:2312.01990
-
SequencePAR: Understanding Pedestrian Attributes via A Sequence Generation Paradigm 4 Dec 2023 · 2 repositories · arXiv:2312.01640
-
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically 4 Dec 2023 · 2 repositories · arXiv:2312.02119Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
VaQuitA: Enhancing Alignment in LLM-Assisted Video Understanding 4 Dec 2023 · 0 repositories · arXiv:2312.02310
-
Automatic Report Generation for Histopathology images using pre-trained Vision Transformers and BERT 3 Dec 2023 · 1 repository · arXiv:2312.01435
-
D-Bot: Database Diagnosis System using Large Language Models 3 Dec 2023 · 1 repository · arXiv:2312.01454Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
ESTformer: Transformer Utilizing Spatiotemporal Dependencies for Electroencaphalogram Super-resolution 3 Dec 2023 · 0 repositories · arXiv:2312.10052
-
MABViT -- Modified Attention Block Enhances Vision Transformers 3 Dec 2023 · 0 repositories · arXiv:2312.01324
-
NLEBench+NorGLM: A Comprehensive Empirical Analysis and Benchmark Dataset for Generative Language Models in Norwegian 3 Dec 2023 · 1 repository · arXiv:2312.01314Syntology official: harvested, nothing ran · 0 ran · 4 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Transformers are uninterpretable with myopic methods: a case study with bounded Dyck grammars 3 Dec 2023 · 0 repositories · arXiv:2312.01429
-
A ripple in time: a discontinuity in American history 2 Dec 2023 · 1 repository · arXiv:2312.01185
-
Axiomatic Preference Modeling for Longform Question Answering 2 Dec 2023 · 0 repositories · arXiv:2312.02206
-
Bootstrapping Interactive Image-Text Alignment for Remote Sensing Image Captioning 2 Dec 2023 · 1 repository · arXiv:2312.01191
-
English to Arabic machine translation of mathematical documents 2 Dec 2023 · 0 repositories · arXiv:2312.03753
-
From Voices to Validity: Leveraging Large Language Models (LLMs) for Textual Analysis of Policy Stakeholder Interviews 2 Dec 2023 · 0 repositories · arXiv:2312.01202
-
Harnessing the Power of Prompt-based Techniques for Generating School-Level Questions using Large Language Models 2 Dec 2023 · 1 repository · arXiv:2312.01032
-
IDPL-PFOD2: A New Large-Scale Dataset for Printed Farsi Optical Character Recognition 2 Dec 2023 · 1 repository · arXiv:2312.01177
-
Large Language Models Are Zero-Shot Text Classifiers 2 Dec 2023 · 1 repository · arXiv:2312.01044
-
Towards leveraging LLMs for Conditional QA 2 Dec 2023 · 0 repositories · arXiv:2312.01143
-
A Bayesian approach for prompt optimization in pre-trained language models 1 Dec 2023 · 0 repositories · arXiv:2312.00471
-
BCN: Batch Channel Normalization for Image Classification 1 Dec 2023 · 1 repository · arXiv:2312.00596
-
Deep Unlearning: Fast and Efficient Gradient-free Approach to Class Forgetting 1 Dec 2023 · 1 repository · arXiv:2312.00761Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 7 unverified (of 17 harvested samples) · 17 pointer-only (licence)
-
EfficientSAM: Leveraged Masked Image Pretraining for Efficient Segment Anything 1 Dec 2023 · 1 repository · arXiv:2312.00863
-
Event Recognition in Laparoscopic Gynecology Videos with Hybrid Transformers 1 Dec 2023 · 0 repositories · arXiv:2312.00593
-
Generative Parameter-Efficient Fine-Tuning 1 Dec 2023 · 1 repository · arXiv:2312.00700Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 16 harvested samples) · 10 pointer-only (licence)
-
Learning to Estimate Critical Gait Parameters from Single-View RGB Videos with Transformer-Based Attention Network 1 Dec 2023 · 1 repository · arXiv:2312.00398
-
Mamba: Linear-Time Sequence Modeling with Selective State Spaces 1 Dec 2023 · 35 repositories · arXiv:2312.00752Syntology community repositories only · 28 ran (of which 7 constructed an object rather than computing a result; 21 with no instrument failure: 0 honoured, 0 violated, 21 with no contract checked; 7 where Syntology's instrument failed) · 34 unverified (of 62 harvested samples) · 29 pointer-only (licence)
-
Mitigating Over-smoothing in Transformers via Regularized Nonlocal Functionals 1 Dec 2023 · 0 repositories · arXiv:2312.00751
-
Nonparametric Variational Regularisation of Pretrained Transformers 1 Dec 2023 · 0 repositories · arXiv:2312.00662
-
Quick Back-Translation for Unsupervised Machine Translation 1 Dec 2023 · 1 repository · arXiv:2312.00912
-
Spatiotemporal Transformer for Imputing Sparse Data: A Deep Learning Approach 1 Dec 2023 · 0 repositories · arXiv:2312.00963
-
SynFundus-1M: A High-quality Million-scale Synthetic fundus images Dataset with Fifteen Types of Annotation 1 Dec 2023 · 1 repository · arXiv:2312.00377
-
A Lightweight Clustering Framework for Unsupervised Semantic Segmentation 30 Nov 2023 · 0 repositories · arXiv:2311.18628
-
Applying Large Language Models and Chain-of-Thought for Automatic Scoring 30 Nov 2023 · 0 repositories · arXiv:2312.03748
-
BAM-DETR: Boundary-Aligned Moment Detection Transformer for Temporal Sentence Grounding in Videos 30 Nov 2023 · 1 repository · arXiv:2312.00083Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 1 honoured, 0 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
Brainformer: Mimic Human Visual Brain Functions to Machine Vision Models via fMRI 30 Nov 2023 · 0 repositories · arXiv:2312.00236
-
Categorical Traffic Transformer: Interpretable and Diverse Behavior Prediction with Tokenized Latent 30 Nov 2023 · 0 repositories · arXiv:2311.18307
-
CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation 30 Nov 2023 · 2 repositories · arXiv:2311.18702Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 3 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Diffusion Models Without Attention 30 Nov 2023 · 0 repositories · arXiv:2311.18257
-
HOT: Higher-Order Dynamic Graph Representation Learning with Efficient Transformers 30 Nov 2023 · 0 repositories · arXiv:2311.18526
-
IAG: Induction-Augmented Generation Framework for Answering Reasoning Questions 30 Nov 2023 · 0 repositories · arXiv:2311.18397
-
MultiResFormer: Transformer with Adaptive Multi-Resolution Modeling for General Time Series Forecasting 30 Nov 2023 · 0 repositories · arXiv:2311.18780
-
OmniMotionGPT: Animal Motion Generation with Limited Data 30 Nov 2023 · 0 repositories · arXiv:2311.18303
-
Relevance-guided Neural Machine Translation 30 Nov 2023 · 0 repositories · arXiv:2312.00214
-
Robust Concept Erasure via Kernelized Rate-Distortion Maximization 30 Nov 2023 · 1 repository · arXiv:2312.00194Syntology official (archive's flag): 20 ran · 20 ran (of which 0 constructed an object rather than computing a result; 18 with no instrument failure: 0 honoured, 0 violated, 18 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 21 harvested samples) · 2 pointer-only (licence)
-
Semantic-Aware Frame-Event Fusion based Pattern Recognition via Large Vision-Language Models 30 Nov 2023 · 1 repository · arXiv:2311.18592
-
TOP-Former: A Multi-Agent Transformer Approach for the Team Orienteering Problem 30 Nov 2023 · 1 repository · arXiv:2311.18662
-
Unnatural Error Correction: GPT-4 Can Almost Perfectly Handle Unnatural Scrambled Text 30 Nov 2023 · 1 repository · arXiv:2311.18805Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples) · 6 pointer-only (licence)