Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 82
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 82 of 190: papers 8,101 to 8,200 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
ScribFormer: Transformer Makes CNN Work Better for Scribble-based Medical Image Segmentation 3 Feb 2024 · 1 repository · arXiv:2402.02029
-
TCI-Former: Thermal Conduction-Inspired Transformer for Infrared Small Target Detection 3 Feb 2024 · 0 repositories · arXiv:2402.02046
-
Topology-Informed Graph Transformer 3 Feb 2024 · 2 repositories · arXiv:2402.02005Syntology official (archive's flag): 8 ran · 17 ran (of which 0 constructed an object rather than computing a result; 17 with no instrument failure: 0 honoured, 2 violated, 15 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 21 harvested samples) · 21 pointer-only (licence)
-
Hierarchical Structure Enhances the Convergence and Generalizability of Linear Molecular Representation 3 Feb 2024 · 1 repository · arXiv:2402.02164
-
A Data-Driven Analysis of Robust Automatic Piano Transcription 2 Feb 2024 · 0 repositories · arXiv:2402.01424
-
ALERT-Transformer: Bridging Asynchronous and Synchronous Machine Learning for Real-Time Event-based Spatio-Temporal Data 2 Feb 2024 · 0 repositories · arXiv:2402.01393
-
An introduction to graphical tensor notation for mechanistic interpretability 2 Feb 2024 · 0 repositories · arXiv:2402.01790
-
BAT: Learning to Reason about Spatial Sounds with Large Language Models 2 Feb 2024 · 0 repositories · arXiv:2402.01591
-
COMET: Generating Commit Messages using Delta Graph Context Representation 2 Feb 2024 · 0 repositories · arXiv:2402.01841
-
Cross-view Masked Diffusion Transformers for Person Image Synthesis 2 Feb 2024 · 1 repository · arXiv:2402.01516Syntology official (archive's flag): 2 ran · 2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified; every one of the 2 samples that ran constructed an object rather than computing a result (of 5 harvested samples)
-
Can LLMs perform structured graph reasoning? 2 Feb 2024 · 1 repository · arXiv:2402.01805
-
Faster Inference of Integer SWIN Transformer by Removing the GELU Activation 2 Feb 2024 · 0 repositories · arXiv:2402.01169
-
How Can Generative AI Enhance the Well-being of Blind? 2 Feb 2024 · 0 repositories · arXiv:2402.07919
-
Improving Sequential Recommendations with LLMs 2 Feb 2024 · 1 repository · arXiv:2402.01339Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples)
-
Integrating Large Language Models in Causal Discovery: A Statistical Causal Approach 2 Feb 2024 · 2 repositories · arXiv:2402.01454
-
LoTR: Low Tensor Rank Weight Adaptation 2 Feb 2024 · 0 repositories · arXiv:2402.01376
-
Retrieval Augmented End-to-End Spoken Dialog Models 2 Feb 2024 · 0 repositories · arXiv:2402.01828
-
Todyformer: Towards Holistic Dynamic Graph Transformers with Structure-Aware Tokenization 2 Feb 2024 · 0 repositories · arXiv:2402.05944
-
CorpusLM: Towards a Unified Language Model on Corpus for Knowledge-Intensive Tasks 2 Feb 2024 · 0 repositories · arXiv:2402.01176
-
Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape 2 Feb 2024 · 0 repositories · arXiv:2402.01258
-
TravelPlanner: A Benchmark for Real-World Planning with Language Agents 2 Feb 2024 · 2 repositories · arXiv:2402.01622Syntology official (archive's flag): 13 ran · 17 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 1 honoured, 0 violated, 12 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 17 harvested samples) · 4 pointer-only (licence)
-
What Will My Model Forget? Forecasting Forgotten Examples in Language Model Refinement 2 Feb 2024 · 0 repositories · arXiv:2402.01865
-
Dendritic Learning-incorporated Vision Transformer for Image Recognition 1 Feb 2024 · 1 repository
-
FuseFormer: A Transformer for Visual and Thermal Image Fusion 1 Feb 2024 · 0 repositories · arXiv:2402.00971
-
Generation, Distillation and Evaluation of Motivational Interviewing-Style Reflections with a Foundational Language Model 1 Feb 2024 · 0 repositories · arXiv:2402.01051
-
Hierarchical Multi-Label Classification of Online Vaccine Concerns 1 Feb 2024 · 0 repositories · arXiv:2402.01783
-
HiQA: A Hierarchical Contextual Augmentation RAG for Multi-Documents QA 1 Feb 2024 · 0 repositories · arXiv:2402.01767
-
Improving Semantic Control in Discrete Latent Spaces with Transformer Quantized Variational Autoencoders 1 Feb 2024 · 1 repository · arXiv:2402.00723
-
Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing 1 Feb 2024 · 0 repositories · arXiv:2402.00658
-
Merging Multi-Task Models via Weight-Ensembling Mixture of Experts 1 Feb 2024 · 1 repository · arXiv:2402.00433Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Multivariate Probabilistic Time Series Forecasting with Correlated Errors 1 Feb 2024 · 1 repository · arXiv:2402.01000Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Ocassionally Secure: A Comparative Analysis of Code Generation Assistants 1 Feb 2024 · 0 repositories · arXiv:2402.00689
-
On the Psychology of GPT-4: Moderately anxious, slightly masculine, honest, and humble 1 Feb 2024 · 0 repositories · arXiv:2402.01777
-
Investigating Recurrent Transformers with Dynamic Halt 1 Feb 2024 · 1 repository · arXiv:2402.00976
-
Self-Supervised Contrastive Pre-Training for Multivariate Point Processes 1 Feb 2024 · 0 repositories · arXiv:2402.00987
-
SPARQL Generation with Entity Pre-trained GPT for KG Question Answering 1 Feb 2024 · 1 repository · arXiv:2402.00969
-
Tiny Titans: Can Smaller Large Language Models Punch Above Their Weight in the Real World for Meeting Summarization? 1 Feb 2024 · 0 repositories · arXiv:2402.00841
-
Human-mediated Large Language Models for Robotic Intervention in Children with Autism Spectrum Disorders 1 Feb 2024 · 0 repositories · arXiv:2402.00260
-
Understanding the Expressive Power and Mechanisms of Transformer for Sequence Modeling 1 Feb 2024 · 0 repositories · arXiv:2402.00522
-
Code-Aware Prompting: A study of Coverage Guided Test Generation in Regression Setting using LLM 31 Jan 2024 · 0 repositories · arXiv:2402.00097
-
Computation and Parameter Efficient Multi-Modal Fusion Transformer for Cued Speech Recognition 31 Jan 2024 · 0 repositories · arXiv:2401.17604
-
ConSmax: Hardware-Friendly Alternative Softmax with Learnable Parameters 31 Jan 2024 · 1 repository · arXiv:2402.10930
-
Document Structure in Long Document Transformers 31 Jan 2024 · 0 repositories · arXiv:2401.17658
-
EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio Captioning 31 Jan 2024 · 1 repository · arXiv:2401.17690
-
Exploring the limits of decoder-only models trained on public speech recognition corpora 31 Jan 2024 · 0 repositories · arXiv:2402.00235
-
Global-Liar: Factuality of LLMs over Time and Geographic Regions 31 Jan 2024 · 0 repositories · arXiv:2401.17839
-
Graph Transformers without Positional Encodings 31 Jan 2024 · 0 repositories · arXiv:2401.17791
-
Head and Neck Tumor Segmentation from [18F]F-FDG PET/CT Images Based on 3D Diffusion Model 31 Jan 2024 · 0 repositories · arXiv:2401.17593
-
Leveraging Swin Transformer for Local-to-Global Weakly Supervised Semantic Segmentation 31 Jan 2024 · 1 repository · arXiv:2401.17828
-
LLM Voting: Human Choices and AI Collective Decision Making 31 Jan 2024 · 1 repository · arXiv:2402.01766Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Local Feature Matching Using Deep Learning: A Survey 31 Jan 2024 · 1 repository · arXiv:2401.17592
-
Making a Long Story Short in Conversation Modeling 31 Jan 2024 · 0 repositories · arXiv:2402.00143
-
Mitigating the Influence of Distractor Tasks in LMs with Prior-Aware Decoding 31 Jan 2024 · 0 repositories · arXiv:2401.17692
-
Paramanu: A Family of Novel Efficient Generative Foundation Language Models for Indian Languages 31 Jan 2024 · 0 repositories · arXiv:2401.18034
-
Positional Encoding Helps Recurrent Neural Networks Handle a Large Vocabulary 31 Jan 2024 · 1 repository · arXiv:2402.00236
-
RAG-Fusion: a New Take on Retrieval-Augmented Generation 31 Jan 2024 · 0 repositories · arXiv:2402.03367
-
RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval 31 Jan 2024 · 3 repositories · arXiv:2401.18059Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Real Sparks of Artificial Intelligence and the Importance of Inner Interpretability 31 Jan 2024 · 0 repositories · arXiv:2402.00901
-
SCAPE: Searching Conceptual Architecture Prompts using Evolution 31 Jan 2024 · 1 repository · arXiv:2402.00089
-
Scavenging Hyena: Distilling Transformers into Long Convolution Models 31 Jan 2024 · 0 repositories · arXiv:2401.17574
-
Evaluating the Capabilities of LLMs for Supporting Anticipatory Impact Assessment 31 Jan 2024 · 0 repositories · arXiv:2401.18028
-
SWEA: Updating Factual Knowledge in Large Language Models via Subject Word Embedding Altering 31 Jan 2024 · 1 repository · arXiv:2401.17809Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Uncertainty-Aware Explainable Recommendation with Large Language Models 31 Jan 2024 · 0 repositories · arXiv:2402.03366
-
WSC+: Enhancing The Winograd Schema Challenge Using Tree-of-Experts 31 Jan 2024 · 1 repository · arXiv:2401.17703
-
A Cross-Language Investigation into Jailbreak Attacks in Large Language Models 30 Jan 2024 · 0 repositories · arXiv:2401.16765
-
A Preliminary Study on Using Large Language Models in Software Pentesting 30 Jan 2024 · 0 repositories · arXiv:2401.17459
-
Breaking Free Transformer Models: Task-specific Context Attribution Promises Improved Generalizability Without Fine-tuning Pre-trained LLMs 30 Jan 2024 · 1 repository · arXiv:2401.16638
-
CAFCT-Net: A CNN-Transformer Hybrid Network with Contextual and Attentional Feature Fusion for Liver Tumor Segmentation 30 Jan 2024 · 0 repositories · arXiv:2401.16886
-
Conditional and Modal Reasoning in Large Language Models 30 Jan 2024 · 1 repository · arXiv:2401.17169
-
CRUD-RAG: A Comprehensive Chinese Benchmark for Retrieval-Augmented Generation of Large Language Models 30 Jan 2024 · 1 repository · arXiv:2401.17043Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Large Multi-Modal Models (LMMs) as Universal Foundation Models for AI-Native Wireless Systems 30 Jan 2024 · 0 repositories · arXiv:2402.01748
-
LLaMP: Large Language Model Made Powerful for High-fidelity Materials Knowledge Retrieval and Distillation 30 Jan 2024 · 1 repository · arXiv:2401.17244Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models 30 Jan 2024 · 1 repository · arXiv:2401.16745Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples) · 2 pointer-only (licence)
-
OptiState: State Estimation of Legged Robots using Gated Networks with Transformer-based Vision and Kalman Filtering 30 Jan 2024 · 1 repository · arXiv:2401.16719
-
OWSM v3.1: Better and Faster Open Whisper-Style Speech Models based on E-Branchformer 30 Jan 2024 · 1 repository · arXiv:2401.16658
-
Performance Assessment of ChatGPT vs Bard in Detecting Alzheimer's Dementia 30 Jan 2024 · 0 repositories · arXiv:2402.01751
-
Robust Prompt Optimization for Defending Language Models Against Jailbreaking Attacks 30 Jan 2024 · 1 repository · arXiv:2401.17263Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Synthetic Dialogue Dataset Generation using LLM Agents 30 Jan 2024 · 1 repository · arXiv:2401.17461
-
T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & Collectives 30 Jan 2024 · 0 repositories · arXiv:2401.16677
-
Weaver: Foundation Models for Creative Writing 30 Jan 2024 · 0 repositories · arXiv:2401.17268
-
3DG: A Framework for Using Generative AI for Handling Sparse Learner Performance Data From Intelligent Tutoring Systems 29 Jan 2024 · 1 repository · arXiv:2402.01746
-
Enhancing Topological Dependencies in Spatio-Temporal Graphs with Cycle Message Passing Blocks 29 Jan 2024 · 1 repository · arXiv:2401.15894
-
A Survey on Structure-Preserving Graph Transformers 29 Jan 2024 · 0 repositories · arXiv:2401.16176
-
Context-Former: Stitching via Latent Conditioned Sequence Modeling 29 Jan 2024 · 0 repositories · arXiv:2401.16452
-
Development and Testing of a Novel Large Language Model-Based Clinical Decision Support Systems for Medication Safety in 12 Clinical Specialties 29 Jan 2024 · 0 repositories · arXiv:2402.01741
-
Development and Testing of Retrieval Augmented Generation in Large Language Models -- A Case Study Report 29 Jan 2024 · 0 repositories · arXiv:2402.01733
-
Diverse, but Divisive: LLMs Can Exaggerate Gender Differences in Opinion Related to Harms of Misinformation 29 Jan 2024 · 0 repositories · arXiv:2401.16558
-
E-EVAL: A Comprehensive Chinese K-12 Education Evaluation Benchmark for Large Language Models 29 Jan 2024 · 1 repository · arXiv:2401.15927Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Hybrid Transformer and Spatial-Temporal Self-Supervised Learning for Long-term Traffic Prediction 29 Jan 2024 · 0 repositories · arXiv:2401.16453
-
Leveraging Professional Radiologists' Expertise to Enhance LLMs' Evaluation for Radiology Reports 29 Jan 2024 · 0 repositories · arXiv:2401.16578
-
LLaVA-MoLE: Sparse Mixture of LoRA Experts for Mitigating Data Conflicts in Instruction Finetuning MLLMs 29 Jan 2024 · 0 repositories · arXiv:2401.16160
-
LLM4Vuln: A Unified Evaluation Framework for Decoupling and Enhancing LLMs' Vulnerability Reasoning 29 Jan 2024 · 0 repositories · arXiv:2401.16185
-
Prompt4Vis: Prompting Large Language Models with Example Mining and Schema Filtering for Tabular Data Visualization 29 Jan 2024 · 0 repositories · arXiv:2402.07909
-
ReGAL: Refactoring Programs to Discover Generalizable Abstractions 29 Jan 2024 · 1 repository · arXiv:2401.16467Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Response Generation for Cognitive Behavioral Therapy with Large Language Models: Comparative Study with Socratic Questioning 29 Jan 2024 · 0 repositories · arXiv:2401.15966
-
An Insight into Security Code Review with LLMs: Capabilities, Obstacles, and Influential Factors 29 Jan 2024 · 0 repositories · arXiv:2401.16310
-
SHViT: Single-Head Vision Transformer with Memory Efficient Macro Design 29 Jan 2024 · 1 repository · arXiv:2401.16456Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Stolen Subwords: Importance of Vocabularies for Machine Translation Model Stealing 29 Jan 2024 · 1 repository · arXiv:2401.16055
-
ToPro: Token-Level Prompt Decomposition for Cross-Lingual Sequence Labeling Tasks 29 Jan 2024 · 1 repository · arXiv:2401.16589
-
TrackGPT -- A generative pre-trained transformer for cross-domain entity trajectory forecasting 29 Jan 2024 · 0 repositories · arXiv:2402.00066