Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 115
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 115 of 190: papers 11,401 to 11,500 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
How Effective Are Neural Networks for Fixing Security Vulnerabilities 29 May 2023 · 1 repository · arXiv:2305.18607Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Large Language Models are not Fair Evaluators 29 May 2023 · 1 repository · arXiv:2305.17926Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
LM-CPPF: Paraphrasing-Guided Data Augmentation for Contrastive Prompt-Based Few-Shot Fine-Tuning 29 May 2023 · 1 repository · arXiv:2305.18169
-
Marked Personas: Using Natural Language Prompts to Measure Stereotypes in Language Models 29 May 2023 · 1 repository · arXiv:2305.18189Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
ProcessGPT: Transforming Business Process Management with Generative Artificial Intelligence 29 May 2023 · 0 repositories · arXiv:2306.01771
-
Syntax and Semantics Meet in the "Middle": Probing the Syntax-Semantics Interface of LMs Through Agentivity 29 May 2023 · 1 repository · arXiv:2305.18185
-
Test-Time Training on Nearest Neighbors for Large Language Models 29 May 2023 · 1 repository · arXiv:2305.18466Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples)
-
The Rise of AI Language Pathologists: Exploring Two-level Prompt Learning for Few-shot Weakly-supervised Whole Slide Image Classification 29 May 2023 · 1 repository · arXiv:2305.17891Syntology official (archive's flag): 4 ran · 4 ran (of which 2 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Transformer Language Models Handle Word Frequency in Prediction Head 29 May 2023 · 0 repositories · arXiv:2305.18294
-
Geometric Algebra Transformer 28 May 2023 · 2 repositories · arXiv:2305.18415Syntology community repositories only · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
A New Deep Learning Architecture withInductive Bias Balance for Transformer Oil Temperature Forecasting 28 May 2023 · 1 repository
-
A Quantitative Review on Language Model Efficiency Research 28 May 2023 · 0 repositories · arXiv:2306.01768
-
Bridging the Language Gap: Dynamic Learning Strategies for Improving Multilingual Performance in LLMs 28 May 2023 · 0 repositories · arXiv:2305.17740
-
DPFormer: Learning Differentially Private Transformer on Long-Tailed Data 28 May 2023 · 0 repositories · arXiv:2305.17633
-
Evaluating GPT-3 Generated Explanations for Hateful Content Moderation 28 May 2023 · 1 repository · arXiv:2305.17680
-
Generating EDU Extracts for Plan-Guided Summary Re-Ranking 28 May 2023 · 1 repository · arXiv:2305.17779Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 6 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Knowledge-Augmented Reasoning Distillation for Small Language Models in Knowledge-Intensive Tasks 28 May 2023 · 1 repository · arXiv:2305.18395Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 1 pointer-only (licence)
-
KoSBi: A Dataset for Mitigating Social Bias Risks Towards Safer Large Language Model Application 28 May 2023 · 1 repository · arXiv:2305.17701
-
Large Language Models, scientific knowledge and factuality: A framework to streamline human expert evaluation 28 May 2023 · 1 repository · arXiv:2305.17819
-
Mitigating Label Biases for In-context Learning 28 May 2023 · 1 repository · arXiv:2305.19148Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Reconstructing Sea Surface Temperature Images: A Masked Autoencoder Approach for Cloud Masking and Reconstruction 28 May 2023 · 0 repositories · arXiv:2306.00835
-
Reward Collapse in Aligning Large Language Models 28 May 2023 · 1 repository · arXiv:2305.17608
-
SQuARe: A Large-Scale Dataset of Sensitive Questions and Acceptable Responses Created Through Human-Machine Collaboration 28 May 2023 · 1 repository · arXiv:2305.17696
-
Transfer Learning for Power Outage Detection Task with Limited Training Data 28 May 2023 · 0 repositories · arXiv:2305.17817
-
Analysis over vision-based models for pedestrian action anticipation 27 May 2023 · 0 repositories · arXiv:2305.17451
-
Bridging the Granularity Gap for Acoustic Modeling 27 May 2023 · 1 repository · arXiv:2305.17356
-
Complementary and Integrative Health Lexicon (CIHLex) and Entity Recognition in the Literature 27 May 2023 · 0 repositories · arXiv:2305.17353
-
CrossGET: Cross-Guided Ensemble of Tokens for Accelerating Vision-Language Transformers 27 May 2023 · 1 repository · arXiv:2305.17455Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples)
-
DNA-GPT: Divergent N-Gram Analysis for Training-Free Detection of GPT-Generated Text 27 May 2023 · 1 repository · arXiv:2305.17359Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Graph Inductive Biases in Transformers without Message Passing 27 May 2023 · 2 repositories · arXiv:2305.17589Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified (of 12 harvested samples) · 1 pointer-only (licence)
-
HyperFormer: Learning Expressive Sparse Feature Representations via Hypergraph Transformer 27 May 2023 · 0 repositories · arXiv:2305.17386
-
The Curse of Recursion: Training on Generated Data Makes Models Forget 27 May 2023 · 1 repository · arXiv:2305.17493
-
Scalable Transformer for PDE Surrogate Modeling 27 May 2023 · 1 repository · arXiv:2305.17560
-
SwiftSage: A Generative Agent with Fast and Slow Thinking for Complex Interactive Tasks 27 May 2023 · 2 repositories · arXiv:2305.17390Syntology 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
Towards Explainable Conversational Recommender Systems 27 May 2023 · 1 repository · arXiv:2305.18363
-
What can Large Language Models do in chemistry? A comprehensive benchmark on eight tasks 27 May 2023 · 1 repository · arXiv:2305.18365
-
Zero-TPrune: Zero-Shot Token Pruning through Leveraging of the Attention Graph in Pre-Trained Transformers 27 May 2023 · 0 repositories · arXiv:2305.17328
-
AlignScore: Evaluating Factual Consistency with a Unified Alignment Function 26 May 2023 · 2 repositories · arXiv:2305.16739
-
Backpack Language Models 26 May 2023 · 1 repository · arXiv:2305.16765Syntology 3 ran (of which 3 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; every one of the 3 samples that ran constructed an object rather than computing a result (of 4 harvested samples)
-
Beyond Chain-of-Thought, Effective Graph-of-Thought Reasoning in Language Models 26 May 2023 · 2 repositories · arXiv:2305.16582Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
BiomedGPT: A Generalist Vision-Language Foundation Model for Diverse Biomedical Tasks 26 May 2023 · 1 repository · arXiv:2305.17100Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples) · 7 pointer-only (licence)
-
Chain-of-Thought Hub: A Continuous Effort to Measure Large Language Models' Reasoning Performance 26 May 2023 · 1 repository · arXiv:2305.17306Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples) · 2 pointer-only (licence)
-
ChatGPT: A Study on its Utility for Ubiquitous Software Engineering Tasks 26 May 2023 · 0 repositories · arXiv:2305.16837
-
COMCAT: Towards Efficient Compression and Customization of Attention-Based Vision Models 26 May 2023 · 1 repository · arXiv:2305.17235Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; the one sample that ran constructed an object rather than computing a result (of 2 harvested samples) · 2 pointer-only (licence)
-
Counterfactual reasoning: Testing language models' understanding of hypothetical scenarios 26 May 2023 · 1 repository · arXiv:2305.16572
-
Distinguishing Human Generated Text From ChatGPT Generated Text Using Machine Learning 26 May 2023 · 0 repositories · arXiv:2306.01761
-
Do GPTs Produce Less Literal Translations? 26 May 2023 · 1 repository · arXiv:2305.16806
-
Emergent Agentic Transformer from Chain of Hindsight Experience 26 May 2023 · 0 repositories · arXiv:2305.16554
-
Evaluation of Question Generation Needs More References 26 May 2023 · 0 repositories · arXiv:2305.16626
-
Future-conditioned Unsupervised Pretraining for Decision Transformer 26 May 2023 · 1 repository · arXiv:2305.16683Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples) · 1 pointer-only (licence)
-
Impossible Distillation: from Low-Quality Model to High-Quality Dataset & Model for Summarization and Paraphrasing 26 May 2023 · 0 repositories · arXiv:2305.16635
-
Improving accuracy of GPT-3/4 results on biomedical data using a retrieval-augmented language model 26 May 2023 · 0 repositories · arXiv:2305.17116
-
Improving Position Encoding of Transformers for Multivariate Time Series Classification 26 May 2023 · 1 repository · arXiv:2305.16642
-
Large Language Models as Tool Makers 26 May 2023 · 1 repository · arXiv:2305.17126
-
Learning and Leveraging Verifiers to Improve Planning Capabilities of Pre-trained Language Models 26 May 2023 · 0 repositories · arXiv:2305.17077
-
Learning to Imagine: Visually-Augmented Natural Language Generation 26 May 2023 · 1 repository · arXiv:2305.16944Syntology official: no sample here; runs from other or unrecorded repositories · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
LLMs and the Abstraction and Reasoning Corpus: Successes, Failures, and the Importance of Object-based Representations 26 May 2023 · 1 repository · arXiv:2305.18354
-
NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models 26 May 2023 · 2 repositories · arXiv:2305.16986Syntology official (archive's flag): 2 ran · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 2 pointer-only (licence)
-
Neural Task Synthesis for Visual Programming 26 May 2023 · 1 repository · arXiv:2305.18342
-
On Evaluating Adversarial Robustness of Large Vision-Language Models 26 May 2023 · 1 repository · arXiv:2305.16934Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 2 pointer-only (licence)
-
Playing repeated games with Large Language Models 26 May 2023 · 0 repositories · arXiv:2305.16867
-
TranSFormer: Slow-Fast Transformer for Machine Translation 26 May 2023 · 0 repositories · arXiv:2305.16982
-
A Survey on ChatGPT: AI-Generated Contents, Challenges, and Solutions 25 May 2023 · 0 repositories · arXiv:2305.18339
-
Asking Before Acting: Gather Information in Embodied Decision Making with Language Models 25 May 2023 · 0 repositories · arXiv:2305.15695
-
ChatGPT for PLC/DCS Control Logic Generation 25 May 2023 · 0 repositories · arXiv:2305.15809
-
Concept-Centric Transformers: Enhancing Model Interpretability through Object-Centric Concept Learning within a Shared Global Workspace 25 May 2023 · 2 repositories · arXiv:2305.15775
-
Cross-view Action Recognition Understanding From Exocentric to Egocentric Perspective 25 May 2023 · 0 repositories · arXiv:2305.15699
-
Imitating Task and Motion Planning with Visuomotor Transformers 25 May 2023 · 0 repositories · arXiv:2305.16309
-
Landmark Attention: Random-Access Infinite Context Length for Transformers 25 May 2023 · 2 repositories · arXiv:2305.16300Syntology official (archive's flag): 1 ran · 11 ran (of which 5 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 6 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples)
-
Linguistic Properties of Truthful Response 25 May 2023 · 1 repository · arXiv:2305.15875
-
MERGE: Fast Private Text Generation 25 May 2023 · 1 repository · arXiv:2305.15769
-
Multi-scale Efficient Graph-Transformer for Whole Slide Image Classification 25 May 2023 · 0 repositories · arXiv:2305.15773
-
NexToU: Efficient Topology-Aware U-Net for Medical Image Segmentation 25 May 2023 · 2 repositories · arXiv:2305.15911Syntology official (archive's flag): 2 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Not wacky vs. definitely wacky: A study of scalar adverbs in pretrained language models 25 May 2023 · 0 repositories · arXiv:2305.16426
-
On the Tool Manipulation Capability of Open-source Large Language Models 25 May 2023 · 1 repository · arXiv:2305.16504Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Pre-training Meets Clustering: A Hybrid Extractive Multi-document Summarization Model 25 May 2023 · 1 repository
-
Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation 25 May 2023 · 1 repository · arXiv:2305.15852Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 9 harvested samples)
-
Stecformer: Spatio-temporal Encoding Cascaded Transformer for Multivariate Long-term Time Series Forecasting 25 May 2023 · 0 repositories · arXiv:2305.16370
-
Text-to-Motion Retrieval: Towards Joint Understanding of Human Motion Data and Natural Language 25 May 2023 · 1 repository · arXiv:2305.15842
-
UMat: Uncertainty-Aware Single Image High Resolution Material Capture 25 May 2023 · 0 repositories · arXiv:2305.16312
-
Undetectable Watermarks for Language Models 25 May 2023 · 0 repositories · arXiv:2306.09194
-
UniTRec: A Unified Text-to-Text Transformer and Joint Contrastive Learning Framework for Text-based Recommendation 25 May 2023 · 1 repository · arXiv:2305.15756Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 11 harvested samples) · 6 pointer-only (licence)
-
VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation 25 May 2023 · 0 repositories · arXiv:2305.16107
-
Voyager: An Open-Ended Embodied Agent with Large Language Models 25 May 2023 · 1 repository · arXiv:2305.16291Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
A Joint Time-frequency Domain Transformer for Multivariate Time Series Forecasting 24 May 2023 · 1 repository · arXiv:2305.14649
-
TriMLP: Revenge of a MLP-like Architecture in Sequential Recommendation 24 May 2023 · 1 repository · arXiv:2305.14675
-
A Causal View of Entity Bias in (Large) Language Models 24 May 2023 · 1 repository · arXiv:2305.14695Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
A New Era in Software Security: Towards Self-Healing Software via Large Language Models and Formal Verification 24 May 2023 · 1 repository · arXiv:2305.14752
-
A RelEntLess Benchmark for Modelling Graded Relations between Named Entities 24 May 2023 · 0 repositories · arXiv:2305.15002
-
Adversarial Demonstration Attacks on Large Language Models 24 May 2023 · 0 repositories · arXiv:2305.14950
-
LAraBench: Benchmarking Arabic AI with Large Language Models 24 May 2023 · 0 repositories · arXiv:2305.14982
-
ByteSized32: A Corpus and Challenge Task for Generating Task-Specific World Models Expressed as Text Games 24 May 2023 · 1 repository · arXiv:2305.14879Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples)
-
Chain-of-Questions Training with Latent Answers for Robust Multistep Question Answering 24 May 2023 · 0 repositories · arXiv:2305.14901
-
ChatAgri: Exploring Potentials of ChatGPT on Cross-linguistic Agricultural Text Classification 24 May 2023 · 1 repository · arXiv:2305.15024
-
Clever Hans or Neural Theory of Mind? Stress Testing Social Reasoning in Large Language Models 24 May 2023 · 0 repositories · arXiv:2305.14763
-
A Survey of Diffusion Models in Natural Language Processing 24 May 2023 · 0 repositories · arXiv:2305.14671
-
Don't Take This Out of Context! On the Need for Contextual Models and Evaluations for Stylistic Rewriting 24 May 2023 · 0 repositories · arXiv:2305.14755
-
Don't Trust ChatGPT when Your Question is not in English: A Study of Multilingual Abilities and Types of LLMs 24 May 2023 · 0 repositories · arXiv:2305.16339
-
Editing Common Sense in Transformers 24 May 2023 · 1 repository · arXiv:2305.14956Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples)
-
ExpertPrompting: Instructing Large Language Models to be Distinguished Experts 24 May 2023 · 2 repositories · arXiv:2305.14688Syntology community repositories only · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)