Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 118
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 118 of 190: papers 11,701 to 11,800 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
LIMA: Less Is More for Alignment 18 May 2023 · 5 repositories · arXiv:2305.11206
-
mLongT5: A Multilingual and Efficient Text-To-Text Transformer for Longer Sequences 18 May 2023 · 1 repository · arXiv:2305.11129
-
MolXPT: Wrapping Molecules with Text for Generative Pre-training 18 May 2023 · 1 repository · arXiv:2305.10688
-
Brain Imaging-to-Graph Generation using Adversarial Hierarchical Diffusion Models for MCI Causality Analysis 18 May 2023 · 0 repositories · arXiv:2305.10754
-
Selecting Learnable Training Samples is All DETRs Need in Crowded Pedestrian Detection 18 May 2023 · 0 repositories · arXiv:2305.10801
-
Support for Stock Trend Prediction Using Transformers and Sentiment Analysis 18 May 2023 · 0 repositories · arXiv:2305.14368
-
TextDiffuser: Diffusion Models as Text Painters 18 May 2023 · 0 repositories · arXiv:2305.10855
-
Vaxformer: Antigenicity-controlled Transformer for Vaccine Design Against SARS-CoV-2 18 May 2023 · 1 repository · arXiv:2305.11194
-
A quantitative study of NLP approaches to question difficulty estimation 17 May 2023 · 1 repository · arXiv:2305.10236
-
CageViT: Convolutional Activation Guided Efficient Vision Transformer 17 May 2023 · 0 repositories · arXiv:2305.09924
-
CoEdIT: Text Editing by Task-Specific Instruction Tuning 17 May 2023 · 1 repository · arXiv:2305.09857
-
CostFormer:Cost Transformer for Cost Aggregation in Multi-view Stereo 17 May 2023 · 0 repositories · arXiv:2305.10320
-
EENED: End-to-End Neural Epilepsy Detection based on Convolutional Transformer 17 May 2023 · 0 repositories · arXiv:2305.10502
-
EfficientSCI: Densely Connected Network with Space-time Factorization for Large-scale Video Snapshot Compressive Imaging 17 May 2023 · 1 repository · arXiv:2305.10006Syntology official (archive's flag): 4 ran · 4 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; every one of the 4 samples that ran constructed an object rather than computing a result (of 5 harvested samples) · 5 pointer-only (licence)
-
From chocolate bunny to chocolate crocodile: Do Language Models Understand Noun Compounds? 17 May 2023 · 0 repositories · arXiv:2305.10568
-
G-Adapter: Towards Structure-Aware Parameter-Efficient Transfer Learning for Graph Transformer Networks 17 May 2023 · 0 repositories · arXiv:2305.10329
-
Improving Speaker Verification with Self-Pretrained Transformer Models 17 May 2023 · 0 repositories · arXiv:2305.10517
-
Instruction Tuned Models are Quick Learners 17 May 2023 · 1 repository · arXiv:2306.05539Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Interactive Learning of Hierarchical Tasks from Dialog with GPT 17 May 2023 · 0 repositories · arXiv:2305.10349
-
Knowledge Graph Completion Models are Few-shot Learners: An Empirical Study of Relation Labeling in E-commerce with LLMs 17 May 2023 · 0 repositories · arXiv:2305.09858
-
Large-Scale Text Analysis Using Generative Language Models: A Case Study in Discovering Public Value Expressions in AI Patents 17 May 2023 · 0 repositories · arXiv:2305.10383
-
M3KE: A Massive Multi-Level Multi-Subject Knowledge Evaluation Benchmark for Chinese Large Language Models 17 May 2023 · 1 repository · arXiv:2305.10263
-
Large Language Models Leverage External Knowledge to Extend Clinical Insight Beyond Language Boundaries 17 May 2023 · 0 repositories · arXiv:2305.10163
-
Rethinking Data Augmentation for Tabular Data in Deep Learning 17 May 2023 · 1 repository · arXiv:2305.10308Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Short-Term Electricity Load Forecasting Using the Temporal Fusion Transformer: Effect of Grid Hierarchies and Data Sources 17 May 2023 · 0 repositories · arXiv:2305.10559
-
Smaller Language Models are Better Black-box Machine-Generated Text Detectors 17 May 2023 · 1 repository · arXiv:2305.09859
-
Token-wise Decomposition of Autoregressive Language Model Hidden States for Analyzing Model Predictions 17 May 2023 · 1 repository · arXiv:2305.10614Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 12 harvested samples) · 7 pointer-only (licence)
-
Tree of Thoughts: Deliberate Problem Solving with Large Language Models 17 May 2023 · 6 repositories · arXiv:2305.10601Syntology official (archive's flag): 2 ran · 19 ran (of which 6 constructed an object rather than computing a result; 18 with no instrument failure: 0 honoured, 0 violated, 18 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (of 24 harvested samples) · 1 pointer-only (licence)
-
When Gradient Descent Meets Derivative-Free Optimization: A Match Made in Black-Box Scenario 17 May 2023 · 0 repositories · arXiv:2305.10013
-
A Preliminary Analysis on the Code Generation Capabilities of GPT-3.5 and Bard AI Models for Java Functions 16 May 2023 · 0 repositories · arXiv:2305.09402
-
Blind Image Quality Assessment via Transformer Predicted Error Map and Perceptual Quality Token 16 May 2023 · 1 repository · arXiv:2305.09353
-
CB-HVTNet: A channel-boosted hybrid vision transformer network for lymphocyte assessment in histopathological images 16 May 2023 · 0 repositories · arXiv:2305.09211
-
Cooperation Is All You Need 16 May 2023 · 0 repositories · arXiv:2305.10449
-
Exploring the Impact of Layer Normalization for Zero-shot Neural Machine Translation 16 May 2023 · 0 repositories · arXiv:2305.09312
-
Generative Table Pre-training Empowers Models for Tabular Prediction 16 May 2023 · 1 repository · arXiv:2305.09696
-
GIFT: Graph-Induced Fine-Tuning for Multi-Party Conversation Understanding 16 May 2023 · 1 repository · arXiv:2305.09360
-
Life of PII -- A PII Obfuscation Transformer 16 May 2023 · 0 repositories · arXiv:2305.09550
-
PanelNet: Understanding 360 Indoor Environment via Panel Representation 16 May 2023 · 0 repositories · arXiv:2305.09078
-
AdamR at SemEval-2023 Task 10: Solving the Class Imbalance Problem in Sexism Detection with Ensemble Learning 15 May 2023 · 0 repositories · arXiv:2305.08636
-
C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models 15 May 2023 · 1 repository · arXiv:2305.08322Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Continual Multimodal Knowledge Graph Construction 15 May 2023 · 1 repository · arXiv:2305.08698
-
Document Understanding Dataset and Evaluation (DUDE) 15 May 2023 · 1 repository · arXiv:2305.08455Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
GeoMAE: Masked Geometric Target Prediction for Self-supervised Point Cloud Pre-Training 15 May 2023 · 1 repository · arXiv:2305.08808
-
Keras GPT Copilot: Integrating the Power of Large Language Models in Deep Learning Model Development 15 May 2023 · 1 repository
-
Knowledge Rumination for Pre-trained Language Models 15 May 2023 · 1 repository · arXiv:2305.08732
-
LoViT: Long Video Transformer for Surgical Phase Recognition 15 May 2023 · 1 repository · arXiv:2305.08989
-
Masked Collaborative Contrast for Weakly Supervised Semantic Segmentation 15 May 2023 · 1 repository · arXiv:2305.08491
-
MaxViT-UNet: Multi-Axis Attention for Medical Image Segmentation 15 May 2023 · 2 repositories · arXiv:2305.08396
-
Physics Informed Token Transformer for Solving Partial Differential Equations 15 May 2023 · 1 repository · arXiv:2305.08757
-
RL4F: Generating Natural Language Feedback with Reinforcement Learning for Repairing Model Outputs 15 May 2023 · 1 repository · arXiv:2305.08844Syntology official (archive's flag): 1 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 1 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples) · 6 pointer-only (licence)
-
Schema-adaptable Knowledge Graph Construction 15 May 2023 · 1 repository · arXiv:2305.08703
-
Sensitivity and Robustness of Large Language Models to Prompt Template in Japanese Text Classification Tasks 15 May 2023 · 0 repositories · arXiv:2305.08714
-
Similarity-weighted Construction of Contextualized Commonsense Knowledge Graphs for Knowledge-intense Argumentation Tasks 15 May 2023 · 1 repository · arXiv:2305.08495
-
Small Models are Valuable Plug-ins for Large Language Models 15 May 2023 · 1 repository · arXiv:2305.08848Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Text Classification via Large Language Models 15 May 2023 · 1 repository · arXiv:2305.08377Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Assessing the potential of LLM-assisted annotation for corpus-based pragmatics and discourse analysis: The case of apology 15 May 2023 · 0 repositories · arXiv:2305.08339
-
Mobile-Env: Building Qualified Evaluation Benchmarks for LLM-GUI Interaction 14 May 2023 · 2 repositories · arXiv:2305.08144Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified (of 17 harvested samples)
-
Self-supervised Neural Factor Analysis for Disentangling Utterance-level Speech Representations 14 May 2023 · 0 repositories · arXiv:2305.08099
-
TSGN: Temporal Scene Graph Neural Networks with Projected Vectorized Representation for Multi-Agent Motion Prediction 14 May 2023 · 0 repositories · arXiv:2305.08190
-
Bridging History with AI A Comparative Evaluation of GPT 3.5, GPT4, and GoogleBARD in Predictive Accuracy and Fact Checking 13 May 2023 · 0 repositories · arXiv:2305.07868
-
CEMFormer: Learning to Predict Driver Intentions from In-Cabin and External Cameras via Spatial-Temporal Transformers 13 May 2023 · 0 repositories · arXiv:2305.07840
-
GPT-Sentinel: Distinguishing Human and ChatGPT Generated Content 13 May 2023 · 2 repositories · arXiv:2305.07969
-
GSB: Group Superposition Binarization for Vision Transformer with Limited Training Samples 13 May 2023 · 1 repository · arXiv:2305.07931
-
The Machine Psychology of Cooperation: Can GPT models operationalise prompts for altruism, cooperation, competitiveness and selfishness in economic games? 13 May 2023 · 2 repositories · arXiv:2305.07970Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Meta-Polyp: a baseline for efficient Polyp segmentation 13 May 2023 · 2 repositories · arXiv:2305.07848
-
PESTS: Persian_English Cross Lingual Corpus for Semantic Textual Similarity 13 May 2023 · 0 repositories · arXiv:2305.07893
-
Stackelberg Decision Transformer for Asynchronous Action Coordination in Multi-Agent Systems 13 May 2023 · 0 repositories · arXiv:2305.07856
-
AGFormer: Efficient Graph Representation with Anchor-Graph Transformer 12 May 2023 · 0 repositories · arXiv:2305.07521
-
ArtGPT-4: Towards Artistic-understanding Large Vision-Language Models with Enhanced Adapter 12 May 2023 · 1 repository · arXiv:2305.07490
-
CLIP-Count: Towards Text-Guided Zero-Shot Object Counting 12 May 2023 · 1 repository · arXiv:2305.07304Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Improving Small Language Models on PubMedQA via Generative Data Augmentation 12 May 2023 · 0 repositories · arXiv:2305.07804
-
Learning to Reason over Scene Graphs: A Case Study of Finetuning GPT-2 into a Robot Language Model for Grounded Task Planning 12 May 2023 · 0 repositories · arXiv:2305.07716
-
NL2TL: Transforming Natural Languages to Temporal Logics using Large Language Models 12 May 2023 · 3 repositories · arXiv:2305.07766Syntology official (archive's flag): 3 ran · 5 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 4 unverified (of 9 harvested samples) · 8 pointer-only (licence)
-
Hausdorff Distance Matching with Adaptive Query Denoising for Rotated Detection Transformer 12 May 2023 · 1 repository · arXiv:2305.07598
-
TinyStories: How Small Can Language Models Be and Still Speak Coherent English? 12 May 2023 · 8 repositories · arXiv:2305.07759Syntology 10 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 8 unverified (of 18 harvested samples)
-
When Giant Language Brains Just Aren't Enough! Domain Pizzazz with Knowledge Sparkle Dust 12 May 2023 · 0 repositories · arXiv:2305.07230
-
COCKATIEL: COntinuous Concept ranKed ATtribution with Interpretable ELements for explaining neural net classifiers on NLP tasks 11 May 2023 · 1 repository · arXiv:2305.06754
-
Generative Pre-trained Transformer: A Comprehensive Review on Enabling Technologies, Potential Applications, Emerging Challenges, and Future Directions 11 May 2023 · 0 repositories · arXiv:2305.10435
-
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning 11 May 2023 · 4 repositories · arXiv:2305.06500
-
Spear Phishing With Large Language Models 11 May 2023 · 0 repositories · arXiv:2305.06972
-
Overinformative Question Answering by Humans and Machines 11 May 2023 · 0 repositories · arXiv:2305.07151
-
PVT-SSD: Single-Stage 3D Object Detector with Point-Voxel Transformer 11 May 2023 · 1 repository · arXiv:2305.06621
-
Recommendation as Instruction Following: A Large Language Model Empowered Recommendation Approach 11 May 2023 · 0 repositories · arXiv:2305.07001
-
Salient Mask-Guided Vision Transformer for Fine-Grained Classification 11 May 2023 · 1 repository · arXiv:2305.07102
-
Subword Segmental Machine Translation: Unifying Segmentation and Target Sentence Generation 11 May 2023 · 1 repository · arXiv:2305.07005
-
The ConceptARC Benchmark: Evaluating Understanding and Generalization in the ARC Domain 11 May 2023 · 1 repository · arXiv:2305.07141
-
Transformers for CT Reconstruction From Monoplanar and Biplanar Radiographs 11 May 2023 · 0 repositories · arXiv:2305.06965
-
Undercover Deepfakes: Detecting Fake Segments in Videos 11 May 2023 · 2 repositories · arXiv:2305.06564
-
A Method to Automate the Discharge Summary Hospital Course for Neurology Patients 10 May 2023 · 0 repositories · arXiv:2305.06416
-
Adapter-TST: A Parameter Efficient Method for Multiple-Attribute Text Style Transfer 10 May 2023 · 0 repositories · arXiv:2305.05945
-
Alternating Gradient Descent and Mixture-of-Experts for Integrated Multimodal Perception 10 May 2023 · 0 repositories · arXiv:2305.06324
-
Are ChatGPT and GPT-4 General-Purpose Solvers for Financial Text Analytics? A Study on Several Typical Tasks 10 May 2023 · 0 repositories · arXiv:2305.05862
-
Autonomous GIS: the next-generation AI-powered GIS 10 May 2023 · 1 repository · arXiv:2305.06453
-
BIOT: Cross-data Biosignal Learning in the Wild 10 May 2023 · 1 repository · arXiv:2305.10351
-
Bits of Grass: Does GPT already know how to write like Whitman? 10 May 2023 · 0 repositories · arXiv:2305.11064
-
Bot or Human? Detecting ChatGPT Imposters with A Single Question 10 May 2023 · 1 repository · arXiv:2305.06424
-
Davinci the Dualist: the mind-body divide in large language models and in human learners 10 May 2023 · 0 repositories · arXiv:2305.07667
-
Generating medically-accurate summaries of patient-provider dialogue: A multi-stage approach using large language models 10 May 2023 · 0 repositories · arXiv:2305.05982
-
Benchmarking large language models for biomedical natural language processing applications and recommendations 10 May 2023 · 1 repository · arXiv:2305.16326
-
MMoT: Mixture-of-Modality-Tokens Transformer for Composed Multimodal Conditional Image Synthesis 10 May 2023 · 0 repositories · arXiv:2305.05992