Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 93
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 93 of 190: papers 9,201 to 9,300 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
LOKE: Linked Open Knowledge Extraction for Automated Knowledge Graph Construction 15 Nov 2023 · 0 repositories · arXiv:2311.09366
-
MELA: Multilingual Evaluation of Linguistic Acceptability 15 Nov 2023 · 1 repository · arXiv:2311.09033
-
Memory Augmented Language Models through Mixture of Word Experts 15 Nov 2023 · 0 repositories · arXiv:2311.10768
-
Progressive Feedback-Enhanced Transformer for Image Forgery Localization 15 Nov 2023 · 1 repository · arXiv:2311.08910
-
Safer-Instruct: Aligning Language Models with Automated Preference Data 15 Nov 2023 · 1 repository · arXiv:2311.08685
-
SparseSpikformer: A Co-Design Framework for Token and Weight Pruning in Spiking Transformer 15 Nov 2023 · 0 repositories · arXiv:2311.08806
-
Token Prediction as Implicit Classification to Identify LLM-Generated Text 15 Nov 2023 · 1 repository · arXiv:2311.08723Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples)
-
ToolTalk: Evaluating Tool-Usage in a Conversational Setting 15 Nov 2023 · 1 repository · arXiv:2311.10775Syntology 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
"We Demand Justice!": Towards Social Context Grounding of Political Texts 15 Nov 2023 · 1 repository · arXiv:2311.09106
-
X-Eval: Generalizable Multi-aspect Text Evaluation via Augmented Instruction Tuning with Auxiliary Evaluation Aspects 15 Nov 2023 · 0 repositories · arXiv:2311.08788
-
Unifying the Perspectives of NLP and Software Engineering: A Survey on Language Models for Code 14 Nov 2023 · 1 repository · arXiv:2311.07989
-
A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily 14 Nov 2023 · 1 repository · arXiv:2311.08268Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 1 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples) · 3 pointer-only (licence)
-
Automated title and abstract screening for scoping reviews using the GPT-4 Large Language Model 14 Nov 2023 · 1 repository · arXiv:2311.07918
-
Comparing Humans, GPT-4, and GPT-4V On Abstraction and Reasoning Tasks 14 Nov 2023 · 0 repositories · arXiv:2311.09247
-
Contrastive Learning for Multi-Object Tracking with Transformers 14 Nov 2023 · 1 repository · arXiv:2311.08043
-
CPopQA: Ranking Cultural Concept Popularity by LLMs 14 Nov 2023 · 0 repositories · arXiv:2311.07897
-
Dual-channel Prototype Network for few-shot Classification of Pathological Images 14 Nov 2023 · 0 repositories · arXiv:2311.07871
-
Evaluating LLMs on Document-Based QA: Exact Answer Selection and Numerical Extraction using Cogtale dataset 14 Nov 2023 · 0 repositories · arXiv:2311.07878
-
Exploring Semi-supervised Hierarchical Stacked Encoder for Legal Judgement Prediction 14 Nov 2023 · 1 repository · arXiv:2311.08103
-
Fair Abstractive Summarization of Diverse Perspectives 14 Nov 2023 · 1 repository · arXiv:2311.07884
-
GMTR: Graph Matching Transformers 14 Nov 2023 · 1 repository · arXiv:2311.08141
-
How good are Large Language Models on African Languages? 14 Nov 2023 · 0 repositories · arXiv:2311.07978
-
Language Models are Better Bug Detector Through Code-Pair Classification 14 Nov 2023 · 1 repository · arXiv:2311.07957
-
Large Language Model-Driven Classroom Flipping: Empowering Student-Centric Peer Questioning with Flipped Interaction 14 Nov 2023 · 0 repositories · arXiv:2311.14708
-
MAgIC: Investigation of Large Language Model Powered Multi-Agent in Cognition, Adaptability, Rationality and Collaboration 14 Nov 2023 · 1 repository · arXiv:2311.08562Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Rotation-Agnostic Image Representation Learning for Digital Pathology 14 Nov 2023 · 0 repositories · arXiv:2311.08359
-
Secure Transformer Inference Protocol 14 Nov 2023 · 1 repository · arXiv:2312.00025Syntology official (archive's flag): 3 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models 14 Nov 2023 · 0 repositories · arXiv:2311.08370
-
Spot: A Natural Language Interface for Geospatial Searches in OSM 14 Nov 2023 · 1 repository · arXiv:2311.08093
-
UT5: Pretraining Non autoregressive T5 with unrolled denoising 14 Nov 2023 · 0 repositories · arXiv:2311.08552
-
A Benchmark to Understand the Role of Knowledge Graphs on Large Language Model's Accuracy for Question Answering on Enterprise SQL Databases 13 Nov 2023 · 0 repositories · arXiv:2311.07509
-
Assessing Logical Puzzle Solving in Large Language Models: Insights from a Minesweeper Case Study 13 Nov 2023 · 1 repository · arXiv:2311.07387Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples)
-
Cross-Axis Transformer with 3D Rotary Positional Embeddings 13 Nov 2023 · 0 repositories · arXiv:2311.07184
-
Do large language models and humans have similar behaviors in causal inference with script knowledge? 13 Nov 2023 · 1 repository · arXiv:2311.07311
-
Fovea Transformer: Efficient Long-Context Modeling with Structured Fine-to-Coarse Attention 13 Nov 2023 · 1 repository · arXiv:2311.07102
-
In-context Learning Generalizes, But Not Always Robustly: The Case of Syntax 13 Nov 2023 · 1 repository · arXiv:2311.07811
-
Interaction is all You Need? A Study of Robots Ability to Understand and Execute 13 Nov 2023 · 1 repository · arXiv:2311.07150
-
It's Not Easy Being Wrong: Large Language Models Struggle with Process of Elimination Reasoning 13 Nov 2023 · 1 repository · arXiv:2311.07532
-
Semantically Grounded QFormer for Efficient Vision Language Understanding 13 Nov 2023 · 0 repositories · arXiv:2311.07449
-
Language Model-In-The-Loop: Data Optimal Approach to Learn-To-Recommend Actions in Text Games 13 Nov 2023 · 0 repositories · arXiv:2311.07687
-
LM-Polygraph: Uncertainty Estimation for Language Models 13 Nov 2023 · 0 repositories · arXiv:2311.07383
-
MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks 13 Nov 2023 · 0 repositories · arXiv:2311.07463
-
On The Truthfulness of 'Surprisingly Likely' Responses of Large Language Models 13 Nov 2023 · 0 repositories · arXiv:2311.07692
-
Speech-based Slot Filling using Large Language Models 13 Nov 2023 · 0 repositories · arXiv:2311.07418
-
STEER: Unified Style Transfer with Expert Reinforcement 13 Nov 2023 · 1 repository · arXiv:2311.07167
-
The Impact of Large Language Models on Scientific Discovery: a Preliminary Study using GPT-4 13 Nov 2023 · 0 repositories · arXiv:2311.07361
-
VerityMath: Advancing Mathematical Reasoning by Self-Verification Through Unit Consistency 13 Nov 2023 · 1 repository · arXiv:2311.07172Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Controllable Topic-Focused Abstractive Summarization 12 Nov 2023 · 0 repositories · arXiv:2311.06724
-
Detecting and Correcting Hate Speech in Multimodal Memes with Large Visual Language Model 12 Nov 2023 · 0 repositories · arXiv:2311.06737
-
Evaluation of GPT-4 for chest X-ray impression generation: A reader study on performance and perception 12 Nov 2023 · 0 repositories · arXiv:2311.06815
-
Flames: Benchmarking Value Alignment of LLMs in Chinese 12 Nov 2023 · 1 repository · arXiv:2311.06899
-
From Complex to Simple: Unraveling the Cognitive Tree for Reasoning with Small Language Models 12 Nov 2023 · 0 repositories · arXiv:2311.06754
-
GIELLM: Japanese General Information Extraction Large Language Model Utilizing Mutual Reinforcement Effect 12 Nov 2023 · 0 repositories · arXiv:2311.06838
-
Large Language Models' Understanding of Math: Source Criticism and Extrapolation 12 Nov 2023 · 0 repositories · arXiv:2311.07618
-
Large Language Models are In-context Teachers for Knowledge Reasoning 12 Nov 2023 · 0 repositories · arXiv:2311.06985
-
TSViT: A Time Series Vision Transformer for Fault Diagnosis 12 Nov 2023 · 0 repositories · arXiv:2311.06916
-
Two Stream Scene Understanding on Graph Embedding 12 Nov 2023 · 0 repositories · arXiv:2311.06746
-
Adversarial Fine-tuning using Generated Respiratory Sound to Address Class Imbalance 11 Nov 2023 · 1 repository · arXiv:2311.06480Syntology official (archive's flag): 3 ran · 3 ran (of which 3 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; every one of the 3 samples that ran constructed an object rather than computing a result (of 3 harvested samples) · 3 pointer-only (licence)
-
CVTHead: One-shot Controllable Head Avatar with Vertex-feature Transformer 11 Nov 2023 · 1 repository · arXiv:2311.06443Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Intentional Biases in LLM Responses 11 Nov 2023 · 0 repositories · arXiv:2311.07611
-
Sparse Attention-Based Neural Networks for Code Classification 11 Nov 2023 · 0 repositories · arXiv:2311.06575
-
Argumentation Element Annotation Modeling using XLNet 10 Nov 2023 · 0 repositories · arXiv:2311.06239
-
Automatic Report Generation for Histopathology images using pre-trained Vision Transformers 10 Nov 2023 · 1 repository · arXiv:2311.06176
-
Data Contamination Quiz: A Tool to Detect and Estimate Contamination in Large Language Models 10 Nov 2023 · 2 repositories · arXiv:2311.06233Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Dual input stream transformer for vertical drift correction in eye-tracking reading data 10 Nov 2023 · 1 repository · arXiv:2311.06095
-
Enhancing Rock Image Segmentation in Digital Rock Physics: A Fusion of Generative AI and State-of-the-Art Neural Networks 10 Nov 2023 · 0 repositories · arXiv:2311.06079
-
Establishing Performance Baselines in Fine-Tuning, Retrieval-Augmented Generation and Soft-Prompting for Non-Specialist LLM Users 10 Nov 2023 · 0 repositories · arXiv:2311.05903
-
Exploring Fine-tuning ChatGPT for News Recommendation 10 Nov 2023 · 0 repositories · arXiv:2311.05850
-
Hiformer: Heterogeneous Feature Interactions Learning with Transformers for Recommender Systems 10 Nov 2023 · 0 repositories · arXiv:2311.05884
-
How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model 10 Nov 2023 · 0 repositories · arXiv:2311.07594
-
Language Models can be Logical Solvers 10 Nov 2023 · 0 repositories · arXiv:2311.06158
-
Making LLMs Worth Every Penny: Resource-Limited Text Classification in Banking 10 Nov 2023 · 0 repositories · arXiv:2311.06102
-
Smart Agent-Based Modeling: On the Use of Large Language Models in Computer Simulations 10 Nov 2023 · 4 repositories · arXiv:2311.06330Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
TransformCode: A Contrastive Learning Framework for Code Embedding via Subtree Transformation 10 Nov 2023 · 1 repository · arXiv:2311.08157
-
Multimodal Foundation Models Exploit Text to Make Medical Image Predictions 9 Nov 2023 · 2 repositories · arXiv:2311.05591
-
BrainNetDiff: Generative AI Empowers Brain Network Generation via Multimodal Diffusion Model 9 Nov 2023 · 0 repositories · arXiv:2311.05199
-
Conic10K: A Challenging Math Problem Understanding and Reasoning Dataset 9 Nov 2023 · 1 repository · arXiv:2311.05113
-
Challenging the Validity of Personality Tests for Large Language Models 9 Nov 2023 · 0 repositories · arXiv:2311.05297
-
Dynamic Association Learning of Self-Attention and Convolution in Image Restoration 9 Nov 2023 · 0 repositories · arXiv:2311.05147
-
GeoFormer: Predicting Human Mobility using Generative Pre-trained Transformer (GPT) 9 Nov 2023 · 0 repositories · arXiv:2311.05092
-
Intelligent Cervical Spine Fracture Detection Using Deep Learning Methods 9 Nov 2023 · 0 repositories · arXiv:2311.05708
-
Large Language Models and Prompt Engineering for Biomedical Query Focused Multi-Document Summarisation 9 Nov 2023 · 0 repositories · arXiv:2311.05169
-
Leveraging Artificial Intelligence Technology for Mapping Research to Sustainable Development Goals: A Case Study 9 Nov 2023 · 0 repositories · arXiv:2311.16162
-
Agent Lumos: Unified and Modular Training for Open-Source Language Agents 9 Nov 2023 · 2 repositories · arXiv:2311.05657Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Protein-ligand binding representation learning from fine-grained interactions 9 Nov 2023 · 0 repositories · arXiv:2311.16160
-
Large Language Models can Strategically Deceive their Users when Put Under Pressure 9 Nov 2023 · 1 repository · arXiv:2311.07590Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Vision Encoder-Decoder Models for AI Coaching 9 Nov 2023 · 2 repositories · arXiv:2311.16161
-
Beyond Size: How Gradients Shape Pruning Decisions in Large Language Models 8 Nov 2023 · 2 repositories · arXiv:2311.04902Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 7 harvested samples) · 1 pointer-only (licence)
-
Data Factors for Better Compositional Generalization 8 Nov 2023 · 1 repository · arXiv:2311.04420Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples)
-
Deep Learning Brasil at ABSAPT 2022: Portuguese Transformer Ensemble Approaches 8 Nov 2023 · 1 repository · arXiv:2311.05051
-
Euclidean, Projective, Conformal: Choosing a Geometric Algebra for Equivariant Transformers 8 Nov 2023 · 1 repository · arXiv:2311.04744Syntology 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
FibroVit—Vision transformer-based framework for detection and classification of pulmonary fibrosis from chest CT images 8 Nov 2023 · 1 repository
-
Hybrid Focal and Full-Range Attention Based Graph Transformers 8 Nov 2023 · 0 repositories · arXiv:2311.04653
-
Loss Masking Is Not Needed in Decoder-only Transformer for Discrete-token-based ASR 8 Nov 2023 · 1 repository · arXiv:2311.04534
-
Massive Editing for Large Language Models via Meta Learning 8 Nov 2023 · 1 repository · arXiv:2311.04661Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
NLQxform: A Language Model-based Question to SPARQL Transformer 8 Nov 2023 · 1 repository · arXiv:2311.07588
-
Rethinking Benchmark and Contamination for Language Models with Rephrased Samples 8 Nov 2023 · 1 repository · arXiv:2311.04850Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
SS-MAE: Spatial-Spectral Masked Auto-Encoder for Multi-Source Remote Sensing Image Classification 8 Nov 2023 · 1 repository · arXiv:2311.04442Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 2 where Syntology's instrument failed) · 6 unverified (of 15 harvested samples) · 15 pointer-only (licence)
-
Towards Few-Annotation Learning in Computer Vision: Application to Image Classification and Object Detection tasks 8 Nov 2023 · 0 repositories · arXiv:2311.04888
-
Vital Sign Forecasting for Sepsis Patients in ICUs 8 Nov 2023 · 0 repositories · arXiv:2311.04770