Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 98
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 98 of 190: papers 9,701 to 9,800 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Dont Add, dont Miss: Effective Content Preserving Generation from Pre-Selected Text Spans 13 Oct 2023 · 1 repository · arXiv:2310.09017Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
Vision Transformers increase efficiency of 3D cardiac CT multi-label segmentation 13 Oct 2023 · 1 repository · arXiv:2310.09099
-
From Words and Exercises to Wellness: Farsi Chatbot for Self-Attachment Technique 13 Oct 2023 · 0 repositories · arXiv:2310.09362
-
GLoRE: Evaluating Logical Reasoning of Large Language Models 13 Oct 2023 · 1 repository · arXiv:2310.09107
-
Human-in-the-loop Machine Translation with Large Language Model 13 Oct 2023 · 1 repository · arXiv:2310.08908
-
PaLI-3 Vision Language Models: Smaller, Faster, Stronger 13 Oct 2023 · 1 repository · arXiv:2310.09199Syntology 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 3 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
Retrieval-Generation Alignment for End-to-End Task-Oriented Dialogue System 13 Oct 2023 · 1 repository · arXiv:2310.08877
-
Table-GPT: Table-tuned GPT for Diverse Table Tasks 13 Oct 2023 · 0 repositories · arXiv:2310.09263
-
QUIK: Towards End-to-End 4-Bit Inference on Generative Large Language Models 13 Oct 2023 · 1 repository · arXiv:2310.09259Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 13 harvested samples)
-
Towards Example-Based NMT with Multi-Levenshtein Transformers 13 Oct 2023 · 1 repository · arXiv:2310.08967Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Age Estimation Based on Graph Convolutional Networks and Multi-head Attention Mechanisms 12 Oct 2023 · 0 repositories · arXiv:2310.08064
-
Can GPT models be Financial Analysts? An Evaluation of ChatGPT and GPT-4 on mock CFA Exams 12 Oct 2023 · 0 repositories · arXiv:2310.08678
-
Can Large Language Models Really Improve by Self-critiquing Their Own Plans? 12 Oct 2023 · 0 repositories · arXiv:2310.08118
-
COVID-19 detection using ViT transformer-based approach from Computed Tomography Images 12 Oct 2023 · 1 repository · arXiv:2310.08165
-
Cross-Episodic Curriculum for Transformer Agents 12 Oct 2023 · 0 repositories · arXiv:2310.08549
-
Do pretrained Transformers Learn In-Context by Gradient Descent? 12 Oct 2023 · 0 repositories · arXiv:2310.08540
-
Impact of time and note duration tokenizations on deep learning symbolic music modeling 12 Oct 2023 · 1 repository · arXiv:2310.08497
-
Interpreting Learned Feedback Patterns in Large Language Models 12 Oct 2023 · 1 repository · arXiv:2310.08164Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 1 pointer-only (licence)
-
Investigating the Robustness and Properties of Detection Transformers (DETR) Toward Difficult Images 12 Oct 2023 · 0 repositories · arXiv:2310.08772
-
Jailbreaking Black Box Large Language Models in Twenty Queries 12 Oct 2023 · 1 repository · arXiv:2310.08419Syntology official (archive's flag): 2 ran · 2 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (of 7 harvested samples)
-
Large language models can replicate cross-cultural differences in personality 12 Oct 2023 · 0 repositories · arXiv:2310.10679
-
Multiclass Classification of Policy Documents with Large Language Models 12 Oct 2023 · 0 repositories · arXiv:2310.08167
-
Octopus: Embodied Vision-Language Programmer from Environmental Feedback 12 Oct 2023 · 1 repository · arXiv:2310.08588Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 1 violated, 4 with no contract checked; 6 where Syntology's instrument failed) · 0 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
Prometheus: Inducing Fine-grained Evaluation Capability in Language Models 12 Oct 2023 · 3 repositories · arXiv:2310.08491Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (of 9 harvested samples) · 5 pointer-only (licence)
-
Promptor: A Conversational and Autonomous Prompt Generation Agent for Intelligent Text Entry Techniques 12 Oct 2023 · 0 repositories · arXiv:2310.08101
-
QASiNa: Religious Domain Question Answering using Sirah Nabawiyah 12 Oct 2023 · 1 repository · arXiv:2310.08102
-
Revisiting Data Augmentation for Rotational Invariance in Convolutional Neural Networks 12 Oct 2023 · 0 repositories · arXiv:2310.08429
-
Self-supervised visual learning for analyzing firearms trafficking activities on the Web 12 Oct 2023 · 0 repositories · arXiv:2310.07975
-
Training Generative Question-Answering on Synthetic Data Obtained from an Instruct-tuned Model 12 Oct 2023 · 0 repositories · arXiv:2310.08072
-
Transformer Choice Net: A Transformer Neural Network for Choice Prediction 12 Oct 2023 · 0 repositories · arXiv:2310.08716
-
Transport-Hub-Aware Spatial-Temporal Adaptive Graph Transformer for Traffic Flow Prediction 12 Oct 2023 · 1 repository · arXiv:2310.08328
-
Ziya-Visual: Bilingual Large Vision-Language Model via Multi-Task Instruction Tuning 12 Oct 2023 · 0 repositories · arXiv:2310.08166
-
3D TransUNet: Advancing Medical Image Segmentation through Vision Transformers 11 Oct 2023 · 3 repositories · arXiv:2310.07781
-
Atom-Motif Contrastive Transformer for Molecular Property Prediction 11 Oct 2023 · 0 repositories · arXiv:2310.07351
-
Cognate Transformer for Automated Phonological Reconstruction and Cognate Reflex Prediction 11 Oct 2023 · 1 repository · arXiv:2310.07487
-
Distance Weighted Trans Network for Image Completion 11 Oct 2023 · 0 repositories · arXiv:2310.07440
-
Distilling Efficient Vision Transformers from CNNs for Semantic Segmentation 11 Oct 2023 · 0 repositories · arXiv:2310.07265
-
Diversity of Thought Improves Reasoning Abilities of LLMs 11 Oct 2023 · 0 repositories · arXiv:2310.07088
-
Does resistance to style-transfer equal Global Shape Bias? Measuring network sensitivity to global shape configuration 11 Oct 2023 · 0 repositories · arXiv:2310.07555
-
Ethical Reasoning over Moral Alignment: A Case and Framework for In-Context Ethical Policies in LLMs 11 Oct 2023 · 0 repositories · arXiv:2310.07251
-
Do Large Language Models have Shared Weaknesses in Medical Question Answering? 11 Oct 2023 · 0 repositories · arXiv:2310.07225
-
Found in the Middle: Permutation Self-Consistency Improves Listwise Ranking in Large Language Models 11 Oct 2023 · 1 repository · arXiv:2310.07712Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Generalized Neural Sorting Networks with Error-Free Differentiable Swap Functions 11 Oct 2023 · 1 repository · arXiv:2310.07174Syntology 5 ran (of which 1 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (of 8 harvested samples)
-
Global Minima, Recoverability Thresholds, and Higher-Order Structure in GNNS 11 Oct 2023 · 0 repositories · arXiv:2310.07667
-
Graph Transformer Network for Flood Forecasting with Heterogeneous Covariates 11 Oct 2023 · 0 repositories · arXiv:2310.07631
-
InstructRetro: Instruction Tuning post Retrieval-Augmented Pretraining 11 Oct 2023 · 1 repository · arXiv:2310.07713
-
Large Language Models Are Zero-Shot Time Series Forecasters 11 Oct 2023 · 2 repositories · arXiv:2310.07820Syntology official (archive's flag): 1 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 1 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples)
-
MatFormer: Nested Transformer for Elastic Inference 11 Oct 2023 · 2 repositories · arXiv:2310.07707Syntology 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
NuTime: Numerically Multi-Scaled Embedding for Large-Scale Time-Series Pretraining 11 Oct 2023 · 1 repository · arXiv:2310.07402Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 1 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples) · 1 pointer-only (licence)
-
ProtoHPE: Prototype-guided High-frequency Patch Enhancement for Visible-Infrared Person Re-identification 11 Oct 2023 · 0 repositories · arXiv:2310.07552
-
Relational Prior Knowledge Graphs for Detection and Instance Segmentation 11 Oct 2023 · 1 repository · arXiv:2310.07573
-
Sparse Universal Transformer 11 Oct 2023 · 2 repositories · arXiv:2310.07096Syntology 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples)
-
From Supervised to Generative: A Novel Paradigm for Tabular Deep Learning with Large Language Models 11 Oct 2023 · 0 repositories · arXiv:2310.07338
-
Uncovering Hidden Connections: Iterative Search and Reasoning for Video-grounded Dialog 11 Oct 2023 · 2 repositories · arXiv:2310.07259Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Diffusion Models for Wireless Communications 11 Oct 2023 · 0 repositories · arXiv:2310.07312
-
Supercharging Graph Transformers with Advective Diffusion 10 Oct 2023 · 0 repositories · arXiv:2310.06417
-
Answer Candidate Type Selection: Text-to-Text Language Model for Closed Book Question Answering Meets Knowledge Graphs 10 Oct 2023 · 0 repositories · arXiv:2310.07008
-
Automated clinical coding using off-the-shelf large language models 10 Oct 2023 · 0 repositories · arXiv:2310.06552
-
Efficient Adaptation of Large Vision Transformer via Adapter Re-Composing 10 Oct 2023 · 1 repository · arXiv:2310.06234Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 4 pointer-only (licence)
-
EViT: An Eagle Vision Transformer with Bi-Fovea Self-Attention 10 Oct 2023 · 1 repository · arXiv:2310.06629
-
Generating and Evaluating Tests for K-12 Students with Language Model Simulations: A Case Study on Sentence Reading Efficiency 10 Oct 2023 · 0 repositories · arXiv:2310.06837
-
GeoLLM: Extracting Geospatial Knowledge from Large Language Models 10 Oct 2023 · 1 repository · arXiv:2310.06213Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
GPT-4 as an Agronomist Assistant? Answering Agriculture Exams Using Large Language Models 10 Oct 2023 · 0 repositories · arXiv:2310.06225
-
Humans and language models diverge when predicting repeating text 10 Oct 2023 · 1 repository · arXiv:2310.06408Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
iTransformer: Inverted Transformers Are Effective for Time Series Forecasting 10 Oct 2023 · 11 repositories · arXiv:2310.06625Syntology community repositories only · 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 2 honoured, 4 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples) · 4 pointer-only (licence)
-
Large Language Models for Propaganda Detection 10 Oct 2023 · 2 repositories · arXiv:2310.06422
-
Learning Stackable and Skippable LEGO Bricks for Efficient, Reconfigurable, and Variable-Resolution Diffusion Modeling 10 Oct 2023 · 1 repository · arXiv:2310.06389Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 3 honoured, 0 violated, 5 with no contract checked; 5 where Syntology's instrument failed) · 4 unverified (of 17 harvested samples) · 8 pointer-only (licence)
-
LLMs as Potential Brainstorming Partners for Math and Science Problems 10 Oct 2023 · 0 repositories · arXiv:2310.10677
-
LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression 10 Oct 2023 · 3 repositories · arXiv:2310.06839
-
Multilingual Jailbreak Challenges in Large Language Models 10 Oct 2023 · 1 repository · arXiv:2310.06474
-
NEWTON: Are Large Language Models Capable of Physical Reasoning? 10 Oct 2023 · 0 repositories · arXiv:2310.07018
-
Sparse Fine-tuning for Inference Acceleration of Large Language Models 10 Oct 2023 · 1 repository · arXiv:2310.06927Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
SWE-bench: Can Language Models Resolve Real-World GitHub Issues? 10 Oct 2023 · 8 repositories · arXiv:2310.06770Syntology 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples)
-
WaveNet: Wave-Aware Image Enhancement 10 Oct 2023 · 1 repository
-
What If the TV Was Off? Examining Counterfactual Reasoning Abilities of Multi-modal Language Models 10 Oct 2023 · 1 repository · arXiv:2310.06627
-
Why bother with geometry? On the relevance of linear decompositions of Transformer embeddings 10 Oct 2023 · 1 repository · arXiv:2310.06977
-
A Glance is Enough: Extract Target Sentence By Looking at A keyword 9 Oct 2023 · 0 repositories · arXiv:2310.05352
-
A Meta-Learning Perspective on Transformers for Causal Language Modeling 9 Oct 2023 · 0 repositories · arXiv:2310.05884
-
A Simple and Robust Framework for Cross-Modality Medical Image Segmentation applied to Vision Transformers 9 Oct 2023 · 2 repositories · arXiv:2310.05572
-
Abstractive Summarization of Large Document Collections Using GPT 9 Oct 2023 · 0 repositories · arXiv:2310.05690
-
Automating Customer Service using LangChain: Building custom open-source GPT Chatbot for organizations 9 Oct 2023 · 0 repositories · arXiv:2310.05421
-
Cabbage Sweeter than Cake? Analysing the Potential of Large Language Models for Learning Conceptual Spaces 9 Oct 2023 · 0 repositories · arXiv:2310.05481
-
DyST: Towards Dynamic Neural Scene Representations on Real-World Videos 9 Oct 2023 · 0 repositories · arXiv:2310.06020
-
FireAct: Toward Language Agent Fine-tuning 9 Oct 2023 · 0 repositories · arXiv:2310.05915
-
Foundation Models Meet Visualizations: Challenges and Opportunities 9 Oct 2023 · 0 repositories · arXiv:2310.05771
-
Integrating Graphs with Large Language Models: Methods and Prospects 9 Oct 2023 · 0 repositories · arXiv:2310.05499
-
Integrating Stock Features and Global Information via Large Language Models for Enhanced Stock Return Prediction 9 Oct 2023 · 0 repositories · arXiv:2310.05627
-
Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment Analysis 9 Oct 2023 · 1 repository · arXiv:2310.05804
-
Exploring the Maze of Multilingual Modeling 9 Oct 2023 · 0 repositories · arXiv:2310.05404
-
Memory-Consistent Neural Networks for Imitation Learning 9 Oct 2023 · 0 repositories · arXiv:2310.06171
-
Put Your Money Where Your Mouth Is: Evaluating Strategic Planning and Execution of LLM Agents in an Auction Arena 9 Oct 2023 · 1 repository · arXiv:2310.05746Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 8 harvested samples)
-
Reinforcement Learning in the Era of LLMs: What is Essential? What is needed? An RL Perspective on RLHF, Prompting, and Beyond 9 Oct 2023 · 2 repositories · arXiv:2310.06147Syntology 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
SC-Safety: A Multi-round Open-ended Question Adversarial Safety Benchmark for Large Language Models in Chinese 9 Oct 2023 · 0 repositories · arXiv:2310.05818
-
SimPLR: A Simple and Plain Transformer for Scaling-Efficient Object Detection and Segmentation 9 Oct 2023 · 0 repositories · arXiv:2310.05920
-
Take a Step Back: Evoking Reasoning via Abstraction in Large Language Models 9 Oct 2023 · 0 repositories · arXiv:2310.06117
-
The Importance of Prompt Tuning for Automated Neuron Explanations 9 Oct 2023 · 0 repositories · arXiv:2310.06200
-
The Program Testing Ability of Large Language Models for Code 9 Oct 2023 · 0 repositories · arXiv:2310.05727
-
Transformer Fusion with Optimal Transport 9 Oct 2023 · 1 repository · arXiv:2310.05719Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 7 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Transformers and Large Language Models for Chemistry and Drug Discovery 9 Oct 2023 · 0 repositories · arXiv:2310.06083
-
Are Emily and Greg Still More Employable than Lakisha and Jamal? Investigating Algorithmic Hiring Bias in the Era of ChatGPT 8 Oct 2023 · 0 repositories · arXiv:2310.05135