Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 120
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 120 of 190: papers 11,901 to 12,000 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Exploring vision transformer layer choosing for semantic segmentation 2 May 2023 · 0 repositories · arXiv:2305.01279
-
FreeLM: Fine-Tuning-Free Language Model 2 May 2023 · 0 repositories · arXiv:2305.01616
-
How to Unleash the Power of Large Language Models for Few-shot Relation Extraction? 2 May 2023 · 2 repositories · arXiv:2305.01555
-
Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation 2 May 2023 · 1 repository · arXiv:2305.01210Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
A Paradigm Shift: The Future of Machine Translation Lies with Large Language Models 2 May 2023 · 0 repositories · arXiv:2305.01181
-
The Pipeline System of ASR and NLU with MLM-based Data Augmentation toward STOP Low-resource Challenge 2 May 2023 · 0 repositories · arXiv:2305.01194
-
Unlimiformer: Long-Range Transformers with Unlimited Length Input 2 May 2023 · 2 repositories · arXiv:2305.01625Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
Vision Meets Definitions: Unsupervised Visual Word Sense Disambiguation Incorporating Gloss Information 2 May 2023 · 1 repository · arXiv:2305.01788
-
Automated Paper Screening for Clinical Reviews Using Large Language Models 1 May 2023 · 0 repositories · arXiv:2305.00844
-
End-to-End Lane detection with One-to-Several Transformer 1 May 2023 · 3 repositories · arXiv:2305.00675
-
Large Linguistic Models: Investigating LLMs' metalinguistic abilities 1 May 2023 · 0 repositories · arXiv:2305.00948
-
LCAUnet: A skin lesion segmentation network with enhanced edge and body fusion 1 May 2023 · 0 repositories · arXiv:2305.00837
-
Multi-scale Transformer-based Network for Emotion Recognition from Multi Physiological Signals 1 May 2023 · 1 repository · arXiv:2305.00769
-
Neural Machine Translation Models with Attention-Based Dropout Layer 1 May 2023 · 1 repository
-
Online Portfolio Management via Deep Reinforcement Learning with High-Frequency Data 1 May 2023 · 1 repository
-
Point Cloud Semantic Segmentation 1 May 2023 · 0 repositories · arXiv:2305.00773
-
Rethinking Boundary Detection in Deep Learning Models for Medical Image Segmentation 1 May 2023 · 1 repository · arXiv:2305.00678
-
Beyond Classification: Financial Reasoning in State-of-the-Art Language Models 30 Apr 2023 · 1 repository · arXiv:2305.01505
-
Discriminative Co-Saliency and Background Mining Transformer for Co-Salient Object Detection 30 Apr 2023 · 1 repository · arXiv:2305.00514
-
Using Large Language Models to Generate JUnit Tests: An Empirical Study 30 Apr 2023 · 1 repository · arXiv:2305.00418
-
How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model 30 Apr 2023 · 3 repositories · arXiv:2305.00586
-
Multimodal Graph Transformer for Multimodal Question Answering 30 Apr 2023 · 0 repositories · arXiv:2305.00581
-
MTC: A Multi-Task Model for Encrypted Network Traffic Classification Based on Transformer and 1D-CNN 29 Apr 2023 · 1 repository
-
3D Brainformer: 3D Fusion Transformer for Brain Tumor Segmentation 28 Apr 2023 · 0 repositories · arXiv:2304.14508
-
A positive feedback method based on F-measure value for Salient Object Detection 28 Apr 2023 · 1 repository · arXiv:2304.14619
-
Causal Reasoning and Large Language Models: Opening a New Frontier for Causality 28 Apr 2023 · 1 repository · arXiv:2305.00050Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
FlowTransformer: A Transformer Framework for Flow-based Network Intrusion Detection Systems 28 Apr 2023 · 1 repository · arXiv:2304.14746
-
LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model 28 Apr 2023 · 3 repositories · arXiv:2304.15010Syntology official: not harvested · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
MASK-CNN-Transformer For Real-Time Multi-Label Weather Recognition 28 Apr 2023 · 0 repositories · arXiv:2304.14857
-
MUDiff: Unified Diffusion for Complete Molecule Generation 28 Apr 2023 · 0 repositories · arXiv:2304.14621
-
ResiDual: Transformer with Dual Residual Connections 28 Apr 2023 · 1 repository · arXiv:2304.14802
-
Speak, Memory: An Archaeology of Books Known to ChatGPT/GPT-4 28 Apr 2023 · 1 repository · arXiv:2305.00118Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Towards Automated Circuit Discovery for Mechanistic Interpretability 28 Apr 2023 · 4 repositories · arXiv:2304.14997Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 5 harvested samples)
-
Training and Evaluation of a Multilingual Tokenizer for GPT-SW3 28 Apr 2023 · 0 repositories · arXiv:2304.14780
-
Boosting Big Brother: Attacking Search Engines with Encodings 27 Apr 2023 · 1 repository · arXiv:2304.14031
-
ChatGPT as an Attack Tool: Stealthy Textual Backdoor Attack via Blackbox Generative Model Trigger 27 Apr 2023 · 0 repositories · arXiv:2304.14475
-
CONSCENDI: A Contrastive and Scenario-Guided Distillation Approach to Guardrail Models for Virtual Assistants 27 Apr 2023 · 0 repositories · arXiv:2304.14364
-
DataComp: In search of the next generation of multimodal datasets 27 Apr 2023 · 3 repositories · arXiv:2304.14108Syntology community repositories only · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Deeply-Coupled Convolution-Transformer with Spatial-temporal Complementary Learning for Video-based Person Re-identification 27 Apr 2023 · 1 repository · arXiv:2304.14122
-
Distinguishing a planetary transit from false positives: a Transformer-based classification for planetary transit signals 27 Apr 2023 · 0 repositories · arXiv:2304.14283
-
Exploiting Inductive Bias in Transformer for Point Cloud Classification and Segmentation 27 Apr 2023 · 1 repository · arXiv:2304.14124
-
Framing the News:From Human Perception to Large Language Model Inferences 27 Apr 2023 · 0 repositories · arXiv:2304.14456
-
ICE-Score: Instructing Large Language Models to Evaluate Code 27 Apr 2023 · 2 repositories · arXiv:2304.14317Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 2 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Lightweight, Pre-trained Transformers for Remote Sensing Timeseries 27 Apr 2023 · 1 repository · arXiv:2304.14065Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified (of 13 harvested samples)
-
Neural Keyphrase Generation: Analysis and Evaluation 27 Apr 2023 · 0 repositories · arXiv:2304.13883
-
Optimization-Inspired Cross-Attention Transformer for Compressive Sensing 27 Apr 2023 · 1 repository · arXiv:2304.13986
-
Origin Tracing and Detecting of LLMs 27 Apr 2023 · 0 repositories · arXiv:2304.14072
-
SweCTRL-Mini: a data-transparent Transformer-based large language model for controllable text generation in Swedish 27 Apr 2023 · 1 repository · arXiv:2304.13994
-
TempEE: Temporal-Spatial Parallel Transformer for Radar Echo Extrapolation Beyond Auto-Regression 27 Apr 2023 · 0 repositories · arXiv:2304.14131
-
We're Afraid Language Models Aren't Modeling Ambiguity 27 Apr 2023 · 1 repository · arXiv:2304.14399
-
Prompting GPT-3.5 for Text-to-SQL with De-semanticization and Skeleton Retrieval 26 Apr 2023 · 0 repositories · arXiv:2304.13301
-
Evaluation of GPT-3.5 and GPT-4 for supporting real-world information needs in healthcare delivery 26 Apr 2023 · 0 repositories · arXiv:2304.13714
-
Exploring the Curious Case of Code Prompts 26 Apr 2023 · 1 repository · arXiv:2304.13250Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 6 harvested samples)
-
Extracting Structured Seed-Mediated Gold Nanorod Growth Procedures from Literature with GPT-3 26 Apr 2023 · 0 repositories · arXiv:2304.13846
-
Technical Report: Impact of Position Bias on Language Models in Token Classification 26 Apr 2023 · 2 repositories · arXiv:2304.13567
-
The Parrot Dilemma: Human-Labeled vs. LLM-augmented Data in Classification Tasks 26 Apr 2023 · 2 repositories · arXiv:2304.13861Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
ScatterFormer: Locally-Invariant Scattering Transformer for Patient-Independent Multispectral Detection of Epileptiform Discharges 26 Apr 2023 · 1 repository · arXiv:2304.14919Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
SIMARA: a database for key-value information extraction from full pages 26 Apr 2023 · 0 repositories · arXiv:2304.13606
-
The Closeness of In-Context Learning and Weight Shifting for Softmax Regression 26 Apr 2023 · 0 repositories · arXiv:2304.13276
-
Towards Multi-Modal DBMSs for Seamless Querying of Texts and Tables 26 Apr 2023 · 0 repositories · arXiv:2304.13559
-
AI-assisted coding: Experiments with GPT-4 25 Apr 2023 · 1 repository · arXiv:2304.13187
-
Application of Transformers for Nonlinear Channel Compensation in Optical Systems 25 Apr 2023 · 0 repositories · arXiv:2304.13119
-
CompletionFormer: Depth Completion with Convolutions and Vision Transformers 25 Apr 2023 · 1 repository · arXiv:2304.13030
-
Depth-Relative Self Attention for Monocular Depth Estimation 25 Apr 2023 · 0 repositories · arXiv:2304.12849
-
DuETT: Dual Event Time Transformer for Electronic Health Records 25 Apr 2023 · 1 repository · arXiv:2304.13017Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Escaping the sentence-level paradigm in machine translation 25 Apr 2023 · 1 repository · arXiv:2304.12959
-
Introducing MBIB -- the first Media Bias Identification Benchmark Task and Dataset Collection 25 Apr 2023 · 1 repository · arXiv:2304.13148
-
LEMaRT: Label-Efficient Masked Region Transform for Image Harmonization 25 Apr 2023 · 0 repositories · arXiv:2304.13166
-
Measuring Massive Multitask Chinese Understanding 25 Apr 2023 · 2 repositories · arXiv:2304.12986Syntology official (archive's flag): 2 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
SDSC-UNet: Dual Skip Connection ViT-based U-shaped Model for Building Extraction 25 Apr 2023 · 1 repository
-
Semantic Compression With Large Language Models 25 Apr 2023 · 0 repositories · arXiv:2304.12512
-
State Spaces Aren't Enough: Machine Translation Needs Attention 25 Apr 2023 · 0 repositories · arXiv:2304.12776
-
STM-UNet: An Efficient U-shaped Architecture Based on Swin Transformer and Multi-scale MLP for Medical Image Segmentation 25 Apr 2023 · 0 repositories · arXiv:2304.12615
-
SwinFSR: Stereo Image Super-Resolution using SwinIR and Frequency Domain Knowledge 25 Apr 2023 · 0 repositories · arXiv:2304.12556
-
The Potential of Visual ChatGPT For Remote Sensing 25 Apr 2023 · 0 repositories · arXiv:2304.13009
-
Theory of Posterior Concentration for Generalized Bayesian Additive Regression Trees 25 Apr 2023 · 0 repositories · arXiv:2304.12505
-
AGI: Artificial General Intelligence for Education 24 Apr 2023 · 0 repositories · arXiv:2304.12479
-
Better Question-Answering Models on a Budget 24 Apr 2023 · 1 repository · arXiv:2304.12370
-
Directed Acyclic Transformer Pre-training for High-quality Non-autoregressive Text Generation 24 Apr 2023 · 1 repository · arXiv:2304.11791Syntology official (archive's flag): 3 ran · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Explicit Correspondence Matching for Generalizable Neural Radiance Fields 24 Apr 2023 · 1 repository · arXiv:2304.12294Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples)
-
Extreme Classification for Answer Type Prediction in Question Answering 24 Apr 2023 · 0 repositories · arXiv:2304.12395
-
Generation-driven Contrastive Self-training for Zero-shot Text Classification with Instruction-following LLM 24 Apr 2023 · 1 repository · arXiv:2304.11872
-
IRNeXt: Rethinking Convolutional Network Design for Image Restoration 24 Apr 2023 · 1 repository
-
Master: Meta Style Transformer for Controllable Zero-Shot and Few-Shot Artistic Style Transfer 24 Apr 2023 · 0 repositories · arXiv:2304.11818
-
MixPro: Data Augmentation with MaskMix and Progressive Attention Labeling for Vision Transformer 24 Apr 2023 · 1 repository · arXiv:2304.12043Syntology official (archive's flag): 16 ran · 16 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 7 where Syntology's instrument failed) · 2 unverified (of 18 harvested samples) · 8 pointer-only (licence)
-
NoiseTrans: Point Cloud Denoising with Transformers 24 Apr 2023 · 0 repositories · arXiv:2304.11812
-
Once Detected, Never Lost: Surpassing Human Performance in Offline LiDAR based 3D Object Detection 24 Apr 2023 · 2 repositories · arXiv:2304.12315Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples)
-
Rank Flow Embedding for Unsupervised and Semi-Supervised Manifold Learning 24 Apr 2023 · 1 repository · arXiv:2304.12448
-
Vision-based Estimation of Fatigue and Engagement in Cognitive Training Sessions 24 Apr 2023 · 1 repository · arXiv:2304.12470
-
Self-regularised Minimum Latency Training for Streaming Transformer-based Speech Recognition 24 Apr 2023 · 0 repositories · arXiv:2304.11985
-
Semantic Tokenizer for Enhanced Natural Language Processing 24 Apr 2023 · 0 repositories · arXiv:2304.12404
-
Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model 24 Apr 2023 · 1 repository · arXiv:2304.13731
-
Transformer-based stereo-aware 3D object detection from binocular images 24 Apr 2023 · 0 repositories · arXiv:2304.11906
-
WizardLM: Empowering Large Language Models to Follow Complex Instructions 24 Apr 2023 · 4 repositories · arXiv:2304.12244Syntology official (archive's flag): 1 ran · 2 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Boosting Theory-of-Mind Performance in Large Language Models via Prompting 22 Apr 2023 · 1 repository · arXiv:2304.11490
-
Dilated-UNet: A Fast and Accurate Medical Image Segmentation Approach using a Dilated Transformer and U-Net Architecture 22 Apr 2023 · 1 repository · arXiv:2304.11450
-
Incomplete Multimodal Learning for Remote Sensing Data Fusion 22 Apr 2023 · 0 repositories · arXiv:2304.11381
-
Can GPT-4 Perform Neural Architecture Search? 21 Apr 2023 · 1 repository · arXiv:2304.10970
-
DeformableFormer: Classification of Endoscopic Ultrasound Guided Fine Needle Biopsy in Pancreatic Diseases 21 Apr 2023 · 0 repositories · arXiv:2304.10791
-
Evaluating Transformer Language Models on Arithmetic Operations Using Number Decomposition 21 Apr 2023 · 1 repository · arXiv:2304.10977