Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 126
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 126 of 190: papers 12,501 to 12,600 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
How Robust is GPT-3.5 to Predecessors? A Comprehensive Study on Language Understanding Tasks 1 Mar 2023 · 0 repositories · arXiv:2303.00293
-
N-best T5: Robust ASR Error Correction using Multiple Input Hypotheses and Constrained Decoding Space 1 Mar 2023 · 0 repositories · arXiv:2303.00456
-
Progressive Scale-aware Network for Remote sensing Image Change Captioning 1 Mar 2023 · 1 repository · arXiv:2303.00355
-
ToxVis: Enabling Interpretability of Implicit vs. Explicit Toxicity Detection Models with Interactive Visualization 1 Mar 2023 · 0 repositories · arXiv:2303.09402
-
A Comprehensive Study on Robustness of Image Classification Models: Benchmarking and Rethinking 28 Feb 2023 · 0 repositories · arXiv:2302.14301
-
A Survey on Long Text Modeling with Transformers 28 Feb 2023 · 0 repositories · arXiv:2302.14502
-
Are Character-level Translations Worth the Wait? Comparing ByT5 and mT5 for Machine Translation 28 Feb 2023 · 1 repository · arXiv:2302.14220
-
BrainBERT: Self-supervised representation learning for intracranial recordings 28 Feb 2023 · 1 repository · arXiv:2302.14367Syntology official (archive's flag): 5 ran · 5 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Zero-Shot Cross-Lingual Summarization via Large Language Models 28 Feb 2023 · 0 repositories · arXiv:2302.14229
-
DC-Former: Diverse and Compact Transformer for Person Re-Identification 28 Feb 2023 · 1 repository · arXiv:2302.14335Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
H-AES: Towards Automated Essay Scoring for Hindi 28 Feb 2023 · 1 repository · arXiv:2302.14635
-
Information-Restricted Neural Language Models Reveal Different Brain Regions' Sensitivity to Semantics, Syntax and Context 28 Feb 2023 · 1 repository · arXiv:2302.14389Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Large Language Models Are State-of-the-Art Evaluators of Translation Quality 28 Feb 2023 · 4 repositories · arXiv:2302.14520
-
Read Pointer Meters in complex environments based on a Human-like Alignment and Recognition Algorithm 28 Feb 2023 · 2 repositories · arXiv:2302.14323
-
Remote Sensing Scene Classification with Masked Image Modeling (MIM) 28 Feb 2023 · 0 repositories · arXiv:2302.14256
-
Ultra-low Precision Multiplication-free Training for Deep Neural Networks 28 Feb 2023 · 0 repositories · arXiv:2302.14458
-
Contrastive Video Question Answering via Video Graph Transformer 27 Feb 2023 · 1 repository · arXiv:2302.13668Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Elementwise Language Representation 27 Feb 2023 · 0 repositories · arXiv:2302.13475
-
Full Stack Optimization of Transformer Inference: a Survey 27 Feb 2023 · 0 repositories · arXiv:2302.14017
-
Inseq: An Interpretability Toolkit for Sequence Generation Models 27 Feb 2023 · 2 repositories · arXiv:2302.13942
-
Let's have a chat! A Conversation with ChatGPT: Technology, Applications, and Limitations 27 Feb 2023 · 0 repositories · arXiv:2302.13817
-
LLaMA: Open and Efficient Foundation Language Models 27 Feb 2023 · 57 repositories · arXiv:2302.13971Syntology official: no sample here; runs from other or unrecorded repositories · 37 ran (of which 9 constructed an object rather than computing a result; 25 with no instrument failure: 3 honoured, 0 violated, 22 with no contract checked; 12 where Syntology's instrument failed) · 21 unverified (of 58 harvested samples) · 4 pointer-only (licence)
-
Reward Design with Language Models 27 Feb 2023 · 1 repository · arXiv:2303.00001
-
SpeechFormer++: A Hierarchical Efficient Framework for Paralinguistic Speech Processing 27 Feb 2023 · 1 repository · arXiv:2302.14638
-
Structured Pruning of Self-Supervised Pre-trained Models for Speech Recognition and Understanding 27 Feb 2023 · 1 repository · arXiv:2302.14132
-
Systematic Rectification of Language Models via Dead-end Analysis 27 Feb 2023 · 1 repository · arXiv:2302.14003Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
Target-Aware Tracking with Long-term Context Attention 27 Feb 2023 · 1 repository · arXiv:2302.13840Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples)
-
Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning 27 Feb 2023 · 3 repositories · arXiv:2302.14115
-
Comparing Sentence-Level Suggestions to Message-Level Suggestions in AI-Mediated Communication 26 Feb 2023 · 0 repositories · arXiv:2302.13382
-
Fast Attention Requires Bounded Entries 26 Feb 2023 · 0 repositories · arXiv:2302.13214
-
Choice Fusion as Knowledge for Zero-Shot Dialogue State Tracking 25 Feb 2023 · 1 repository · arXiv:2302.13013
-
Introducing Depth into Transformer-based 3D Object Detection 25 Feb 2023 · 0 repositories · arXiv:2302.13002
-
Human-in-the-Loop Schema Induction 25 Feb 2023 · 0 repositories · arXiv:2302.13048
-
Prompt-based Learning for Text Readability Assessment 25 Feb 2023 · 1 repository · arXiv:2302.13139
-
Sequential Query Encoding For Complex Query Answering on Knowledge Graphs 25 Feb 2023 · 1 repository · arXiv:2302.13114
-
TBFormer: Two-Branch Transformer for Image Forgery Localization 25 Feb 2023 · 1 repository · arXiv:2302.13004
-
A Convolutional Vision Transformer for Semantic Segmentation of Side-Scan Sonar Data 24 Feb 2023 · 1 repository · arXiv:2302.12416
-
A Deep Neural Network Based Reverse Radio Spectrogram Search Algorithm 24 Feb 2023 · 0 repositories · arXiv:2302.13854
-
HULAT at SemEval-2023 Task 10: Data augmentation for pre-trained transformers applied to the detection of sexism in social media 24 Feb 2023 · 1 repository · arXiv:2302.12840
-
Retrieved Sequence Augmentation for Protein Representation Learning 24 Feb 2023 · 1 repository · arXiv:2302.12563Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Spanish Built Factual Freectianary (Spanish-BFF): the first AI-generated free dictionary 24 Feb 2023 · 0 repositories · arXiv:2302.12746
-
TrafFormer: A Transformer Model for Predicting Long-term Traffic 24 Feb 2023 · 1 repository · arXiv:2302.12388
-
Window transformer for dialogue document: a joint framework for causal emotion entailment 24 Feb 2023 · 0 repositories
-
Does Deep Learning Learn to Abstract? A Systematic Probing Framework 23 Feb 2023 · 1 repository · arXiv:2302.11978
-
LightCTS: A Lightweight Framework for Correlated Time Series Forecasting 23 Feb 2023 · 1 repository · arXiv:2302.11974
-
MossFormer: Pushing the Performance Limit of Monaural Speech Separation using Gated Single-Head Transformer with Convolution-Augmented Joint Self-Attentions 23 Feb 2023 · 2 repositories · arXiv:2302.11824Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
On the Generalization Ability of Retrieval-Enhanced Transformers 23 Feb 2023 · 0 repositories · arXiv:2302.12128
-
Patch Network for medical image Segmentation 23 Feb 2023 · 0 repositories · arXiv:2302.11802
-
One Fits All:Power General Time Series Analysis by Pretrained LM 23 Feb 2023 · 3 repositories · arXiv:2302.11939Syntology official (archive's flag): 1 ran · 6 ran (of which 6 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified; every one of the 6 samples that ran constructed an object rather than computing a result (of 8 harvested samples) · 8 pointer-only (licence)
-
StudyFormer : Attention-Based and Dynamic Multi View Classifier for X-ray images 23 Feb 2023 · 0 repositories · arXiv:2302.11840
-
Teacher Intervention: Improving Convergence of Quantization Aware Training for Ultra-Low Precision Transformers 23 Feb 2023 · 1 repository · arXiv:2302.11812Syntology official: no sample here; runs from other or unrecorded repositories · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Testing AI on language comprehension tasks reveals insensitivity to underlying meaning 23 Feb 2023 · 0 repositories · arXiv:2302.12313
-
Transformers in Single Object Tracking: An Experimental Survey 23 Feb 2023 · 0 repositories · arXiv:2302.11867
-
EVJVQA Challenge: Multilingual Visual Question Answering 23 Feb 2023 · 0 repositories · arXiv:2302.11752
-
What makes a language easy to deep-learn? Deep neural networks and humans similarly benefit from compositional structure 23 Feb 2023 · 1 repository · arXiv:2302.12239Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth 23 Feb 2023 · 6 repositories · arXiv:2302.12288Syntology official (archive's flag): 1 ran · 12 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 7 where Syntology's instrument failed) · 4 unverified (of 16 harvested samples) · 10 pointer-only (licence)
-
BB-GCN: A Bi-modal Bridged Graph Convolutional Network for Multi-label Chest X-Ray Recognition 22 Feb 2023 · 0 repositories · arXiv:2302.11082
-
Cross-modal Audio-visual Co-learning for Text-independent Speaker Verification 22 Feb 2023 · 1 repository · arXiv:2302.11254
-
HINormer: Representation Learning On Heterogeneous Information Networks with Graph Transformer 22 Feb 2023 · 1 repository · arXiv:2302.11329
-
KS-DETR: Knowledge Sharing in Attention Learning for Detection Transformer 22 Feb 2023 · 1 repository · arXiv:2302.11208
-
Video-SwinUNet: Spatio-temporal Deep Learning Framework for VFSS Instance Segmentation 22 Feb 2023 · 2 repositories · arXiv:2302.11325
-
Bokeh Rendering Based on Adaptive Depth Calibration Network 21 Feb 2023 · 0 repositories · arXiv:2302.10808
-
ChatGPT: Jack of all trades, master of none 21 Feb 2023 · 1 repository · arXiv:2302.10724Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Co-Driven Recognition of Semantic Consistency via the Fusion of Transformer and HowNet Sememes Knowledge 21 Feb 2023 · 1 repository · arXiv:2302.10570
-
Edgeformers: Graph-Empowered Transformers for Representation Learning on Textual-Edge Networks 21 Feb 2023 · 1 repository · arXiv:2302.11050Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 7 harvested samples)
-
Hyena Hierarchy: Towards Larger Convolutional Language Models 21 Feb 2023 · 7 repositories · arXiv:2302.10866Syntology official (archive's flag): 2 ran · 5 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 4 pointer-only (licence)
-
kNN-Adapter: Efficient Domain Adaptation for Black-Box Language Models 21 Feb 2023 · 0 repositories · arXiv:2302.10879
-
Label Information Enhanced Fraud Detection against Low Homophily in Graphs 21 Feb 2023 · 1 repository · arXiv:2302.10407
-
Lightweight Real-time Semantic Segmentation Network with Efficient Transformer and CNN 21 Feb 2023 · 1 repository · arXiv:2302.10484
-
MulGT: Multi-task Graph-Transformer with Task-aware Knowledge Injection and Domain Knowledge-driven Pooling for Whole Slide Image Analysis 21 Feb 2023 · 0 repositories · arXiv:2302.10574
-
MVMTnet: A Multi-variate Multi-modal Transformer for Multi-class Classification of Cardiac Irregularities Using ECG Waveforms and Clinical Notes 21 Feb 2023 · 1 repository · arXiv:2302.11021
-
Time to Embrace Natural Language Processing (NLP)-based Digital Pathology: Benchmarking NLP- and Convolutional Neural Network-based Deep Learning Pipelines 21 Feb 2023 · 0 repositories · arXiv:2302.10406
-
Exploring the Advantages of Transformers for High-Frequency Trading 20 Feb 2023 · 1 repository · arXiv:2302.13850
-
Friend Ranking in Online Games via Pre-training Edge Transformers 20 Feb 2023 · 2 repositories · arXiv:2302.10043
-
GlocalFuse-Depth: Fusing Transformers and CNNs for All-day Self-supervised Monocular Depth Estimation 20 Feb 2023 · 0 repositories · arXiv:2302.09884
-
Large-scale Multi-Modal Pre-trained Models: A Comprehensive Survey 20 Feb 2023 · 1 repository · arXiv:2302.10035
-
Optical Transformers 20 Feb 2023 · 0 repositories · arXiv:2302.10360
-
STB-VMM: Swin Transformer Based Video Motion Magnification 20 Feb 2023 · 1 repository · arXiv:2302.10001
-
ChatIE: Zero-Shot Information Extraction via Chatting with ChatGPT 20 Feb 2023 · 1 repository · arXiv:2302.10205
-
MedViT: A Robust Vision Transformer for Generalized Medical Image Classification 19 Feb 2023 · 1 repository · arXiv:2302.09462Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
A Comprehensive Survey on Pretrained Foundation Models: A History from BERT to ChatGPT 18 Feb 2023 · 0 repositories · arXiv:2302.09419
-
Bag of Tricks for Effective Language Model Pretraining and Downstream Adaptation: A Case Study on GLUE 18 Feb 2023 · 0 repositories · arXiv:2302.09268
-
BBT-Fin: Comprehensive Construction of Chinese Financial Domain Pre-trained Language Model, Corpus and Benchmark 18 Feb 2023 · 2 repositories · arXiv:2302.09432
-
How Good Are GPT Models at Machine Translation? A Comprehensive Evaluation 18 Feb 2023 · 1 repository · arXiv:2302.09210Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Hyneter: Hybrid Network Transformer for Object Detection 18 Feb 2023 · 0 repositories · arXiv:2302.09365
-
Neural Attention Memory 18 Feb 2023 · 0 repositories · arXiv:2302.09422
-
Bounding the Capabilities of Large Language Models in Open Text Generation with Prompt Constraints 17 Feb 2023 · 1 repository · arXiv:2302.09185
-
Conveying the Predicted Future to Users: A Case Study of Story Plot Prediction 17 Feb 2023 · 1 repository · arXiv:2302.09122
-
DTAAD: Dual Tcn-Attention Networks for Anomaly Detection in Multivariate Time Series Data 17 Feb 2023 · 1 repository · arXiv:2302.10753
-
GPT4MIA: Utilizing Generative Pre-trained Transformer (GPT-3) as A Plug-and-Play Transductive Model for Medical Image Analysis 17 Feb 2023 · 0 repositories · arXiv:2302.08722
-
Improving Transformer-based Networks With Locality For Automatic Speaker Verification 17 Feb 2023 · 0 repositories · arXiv:2302.08639
-
Like a Good Nearest Neighbor: Practical Content Moderation and Text Classification 17 Feb 2023 · 1 repository · arXiv:2302.08957
-
EnfoMax: Domain Entropy and Mutual Information Maximization for Domain Generalized Face Anti-spoofing 17 Feb 2023 · 0 repositories · arXiv:2302.08674
-
Multiresolution Graph Transformers and Wavelet Positional Encoding for Learning Hierarchical Structures 17 Feb 2023 · 2 repositories · arXiv:2302.08647Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 2 pointer-only (licence)
-
PAC Prediction Sets for Large Language Models of Code 17 Feb 2023 · 1 repository · arXiv:2302.08703
-
Prompting Large Language Models With the Socratic Method 17 Feb 2023 · 0 repositories · arXiv:2303.08769
-
Transformer-based Generative Adversarial Networks in Computer Vision: A Comprehensive Survey 17 Feb 2023 · 0 repositories · arXiv:2302.08641
-
ViTA: A Vision Transformer Inference Accelerator for Edge Applications 17 Feb 2023 · 0 repositories · arXiv:2302.09108
-
A Transformer-based Deep Learning Algorithm to Auto-record Undocumented Clinical One-Lung Ventilation Events 16 Feb 2023 · 0 repositories · arXiv:2302.12713
-
Document Flattening: Beyond Concatenating Context for Document-Level Neural Machine Translation 16 Feb 2023 · 0 repositories · arXiv:2302.08079