Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 56
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 56 of 190: papers 5,501 to 5,600 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Sound Tagging in Infant-centric Home Soundscapes 25 Jun 2024 · 0 repositories · arXiv:2406.17190
-
Structured Unrestricted-Rank Matrices for Parameter Efficient Fine-tuning 25 Jun 2024 · 1 repository · arXiv:2406.17740Syntology official (archive's flag): 2 ran · 2 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Task-Agnostic Federated Learning 25 Jun 2024 · 0 repositories · arXiv:2406.17235
-
Temporal-Channel Modeling in Multi-head Self-Attention for Synthetic Speech Detection 25 Jun 2024 · 1 repository · arXiv:2406.17376
-
Towards Optimal Trade-offs in Knowledge Distillation for CNNs and Vision Transformers at the Edge 25 Jun 2024 · 0 repositories · arXiv:2407.12808
-
Univariate Skeleton Prediction in Multivariate Systems Using Transformers 25 Jun 2024 · 1 repository · arXiv:2406.17834
-
Understanding Language Model Circuits through Knowledge Editing 25 Jun 2024 · 0 repositories · arXiv:2406.17241
-
Anomaly Detection of Tabular Data Using LLMs 24 Jun 2024 · 0 repositories · arXiv:2406.16308
-
Attention Instruction: Amplifying Attention in the Middle via Prompting 24 Jun 2024 · 1 repository · arXiv:2406.17095
-
Building on Efficient Foundations: Effectively Training LLMs with Structured Feedforward Layers 24 Jun 2024 · 1 repository · arXiv:2406.16450
-
Classification of Geological Borehole Descriptions Using a Domain Adapted Large Language Model 24 Jun 2024 · 0 repositories · arXiv:2407.10991
-
Diff3Dformer: Leveraging Slice Sequence Diffusion for Enhanced 3D CT Classification with Transformer Networks 24 Jun 2024 · 0 repositories · arXiv:2406.17173
-
DreamBench++: A Human-Aligned Benchmark for Personalized Image Generation 24 Jun 2024 · 1 repository · arXiv:2406.16855Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Evaluation of Instruction-Following Ability for Large Language Models on Story-Ending Generation 24 Jun 2024 · 0 repositories · arXiv:2406.16356
-
Evaluation of Language Models in the Medical Context Under Resource-Constrained Settings 24 Jun 2024 · 1 repository · arXiv:2406.16611
-
Exploring Factual Entailment with NLI: A News Media Study 24 Jun 2024 · 0 repositories · arXiv:2406.16842
-
Feature Fusion for Human Activity Recognition using Parameter-Optimized Multi-Stage Graph Convolutional Network and Transformer Models 24 Jun 2024 · 0 repositories · arXiv:2406.16638
-
Finding Transformer Circuits with Edge Pruning 24 Jun 2024 · 1 repository · arXiv:2406.16778Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 3 honoured, 0 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples)
-
GeoMFormer: A General Architecture for Geometric Molecular Representation Learning 24 Jun 2024 · 1 repository · arXiv:2406.16853
-
GMT: Guided Mask Transformer for Leaf Instance Segmentation 24 Jun 2024 · 1 repository · arXiv:2406.17109
-
Large Language Models in Student Assessment: Comparing ChatGPT and Human Graders 24 Jun 2024 · 0 repositories · arXiv:2406.16510
-
Make Graph Neural Networks Great Again: A Generic Integration Paradigm of Topology-Free Patterns for Traffic Speed Prediction 24 Jun 2024 · 1 repository · arXiv:2406.16992
-
METRIK: Measurement-Efficient Randomized Controlled Trials using Transformers with Input Masking 24 Jun 2024 · 0 repositories · arXiv:2406.16351
-
modeLing: A Novel Dataset for Testing Linguistic Reasoning in Language Models 24 Jun 2024 · 0 repositories · arXiv:2406.17038
-
Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models 24 Jun 2024 · 1 repository · arXiv:2406.17169
-
Multi-Modal Vision Transformers for Crop Mapping from Satellite Image Time Series 24 Jun 2024 · 0 repositories · arXiv:2406.16513
-
On the Role of Long-tail Knowledge in Retrieval Augmented Large Language Models 24 Jun 2024 · 0 repositories · arXiv:2406.16367
-
OTCE: Hybrid SSM and Attention with Cross Domain Mixture of Experts to construct Observer-Thinker-Conceiver-Expresser 24 Jun 2024 · 1 repository · arXiv:2406.16495
-
Panza: Design and Analysis of a Fully-Local Personalized Text Writing Assistant 24 Jun 2024 · 1 repository · arXiv:2407.10994
-
PlagBench: Exploring the Duality of Large Language Models in Plagiarism Generation and Detection 24 Jun 2024 · 0 repositories · arXiv:2406.16288
-
Ragnarök: A Reusable RAG Framework and Baselines for TREC 2024 Retrieval-Augmented Generation Track 24 Jun 2024 · 2 repositories · arXiv:2406.16828Syntology official (archive's flag): 9 ran · 19 ran (of which 0 constructed an object rather than computing a result; 19 with no instrument failure: 0 honoured, 0 violated, 19 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 23 harvested samples)
-
The GPT-WritingPrompts Dataset: A Comparative Analysis of Character Portrayal in Short Stories 24 Jun 2024 · 1 repository · arXiv:2406.16767
-
Towards Better Graph-based Cross-document Relation Extraction via Non-bridge Entity Enhancement and Prediction Debiasing 24 Jun 2024 · 1 repository · arXiv:2406.16529
-
MixTex: Unambiguous Recognition Should Not Rely Solely on Real Data 24 Jun 2024 · 1 repository · arXiv:2406.17148
-
UNO Arena for Evaluating Sequential Decision-Making Capability of Large Language Models 24 Jun 2024 · 0 repositories · arXiv:2406.16382
-
USDC: A Dataset of User Stance and Dogmatism in Long Conversations 24 Jun 2024 · 0 repositories · arXiv:2406.16833
-
Venturing into Uncharted Waters: The Navigation Compass from Transformer to Mamba 24 Jun 2024 · 0 repositories · arXiv:2406.16722
-
Breaking the Frame: Visual Place Recognition by Overlap Prediction 23 Jun 2024 · 1 repository · arXiv:2406.16204Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
EditFollower: Tunable Car Following Models for Customizable Adaptive Cruise Control Systems 23 Jun 2024 · 0 repositories · arXiv:2407.02516
-
Enhancing Commentary Strategies for Imperfect Information Card Games: A Study of Large Language Models in Guandan Commentary 23 Jun 2024 · 1 repository · arXiv:2406.17807
-
Evaluating the Effectiveness of the Foundational Models for Q&A Classification in Mental Health care 23 Jun 2024 · 0 repositories · arXiv:2406.15966
-
GraphEval2000: Benchmarking and Improving Large Language Models on Graph Datasets 23 Jun 2024 · 0 repositories · arXiv:2406.16176
-
Multi-Scale Temporal Difference Transformer for Video-Text Retrieval 23 Jun 2024 · 0 repositories · arXiv:2406.16111
-
Wound Tissue Segmentation in Diabetic Foot Ulcer Images Using Deep Learning: A Pilot Study 23 Jun 2024 · 1 repository · arXiv:2406.16012
-
Beyond Individual Facts: Investigating Categorical Knowledge Locality of Taxonomy and Meronomy Concepts in GPT Models 22 Jun 2024 · 0 repositories · arXiv:2406.15940
-
Can LLMs Generate Visualizations with Dataless Prompts? 22 Jun 2024 · 0 repositories · arXiv:2406.17805
-
Decentralized Transformers with Centralized Aggregation are Sample-Efficient Multi-Agent World Models 22 Jun 2024 · 1 repository · arXiv:2406.15836
-
Enhancing Solar Driver Forecasting with Multivariate Transformers 22 Jun 2024 · 1 repository · arXiv:2406.15847
-
Fast Tree-Field Integrators: From Low Displacement Rank to Topological Transformers 22 Jun 2024 · 1 repository · arXiv:2406.15881Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Ladder: A Model-Agnostic Framework Boosting LLM-based Machine Translation to the Next Level 22 Jun 2024 · 3 repositories · arXiv:2406.15741Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 6 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Soft Masked Mamba Diffusion Model for CT to MRI Conversion 22 Jun 2024 · 1 repository · arXiv:2406.15910
-
SS-GEN: A Social Story Generation Framework with Large Language Models 22 Jun 2024 · 2 repositories · arXiv:2406.15695
-
A GPT-based Code Review System for Programming Language Learning 21 Jun 2024 · 0 repositories · arXiv:2407.04722
-
A SMART Mnemonic Sounds like "Glue Tonic": Mixing LLMs with Student Feedback to Make Mnemonic Learning Stick 21 Jun 2024 · 1 repository · arXiv:2406.15352
-
A Tale of Trust and Accuracy: Base vs. Instruct LLMs in RAG Systems 21 Jun 2024 · 1 repository · arXiv:2406.14972
-
Self-Supervised Adversarial Diffusion Models for Fast MRI Reconstruction 21 Jun 2024 · 0 repositories · arXiv:2406.15656
-
Anime Popularity Prediction Before Huge Investments: a Multimodal Approach Using Deep Learning 21 Jun 2024 · 0 repositories · arXiv:2406.16961
-
Brain-Like Language Processing via a Shallow Untrained Multihead Attention Network 21 Jun 2024 · 1 repository · arXiv:2406.15109
-
Data Efficient Evaluation of Large Language Models and Text-to-Image Models via Adaptive Sampling 21 Jun 2024 · 0 repositories · arXiv:2406.15527
-
Efficient Continual Pre-training by Mitigating the Stability Gap 21 Jun 2024 · 0 repositories · arXiv:2406.14833
-
ESC-Eval: Evaluating Emotion Support Conversations in Large Language Models 21 Jun 2024 · 3 repositories · arXiv:2406.14952Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
How Effective is GPT-4 Turbo in Generating School-Level Questions from Textbooks Based on Bloom's Revised Taxonomy? 21 Jun 2024 · 0 repositories · arXiv:2406.15211
-
Inferring Pluggable Types with Machine Learning 21 Jun 2024 · 0 repositories · arXiv:2406.15676
-
InternLM-Law: An Open Source Chinese Legal Large Language Model 21 Jun 2024 · 1 repository · arXiv:2406.14887
-
LongRAG: Enhancing Retrieval-Augmented Generation with Long-context LLMs 21 Jun 2024 · 0 repositories · arXiv:2406.15319
-
Pistis-RAG: Enhancing Retrieval-Augmented Generation with Human Feedback 21 Jun 2024 · 0 repositories · arXiv:2407.00072
-
Differentiable and Learnable Wireless Simulation with Geometric Transformers 21 Jun 2024 · 0 repositories · arXiv:2406.14995
-
Root Cause Analysis of Anomalies in 5G RAN Using Graph Neural Network and Transformer 21 Jun 2024 · 0 repositories · arXiv:2406.15638
-
SiT: Symmetry-Invariant Transformers for Generalisation in Reinforcement Learning 21 Jun 2024 · 1 repository · arXiv:2406.15025
-
TinyStyler: Efficient Few-Shot Text Style Transfer with Authorship Embeddings 21 Jun 2024 · 1 repository · arXiv:2406.15586
-
Unsupervised Morphological Tree Tokenizer 21 Jun 2024 · 0 repositories · arXiv:2406.15245
-
V-RECS, a Low-Cost LLM4VIS Recommender with Explanations, Captioning and Suggestions 21 Jun 2024 · 1 repository · arXiv:2406.15259
-
What Teaches Robots to Walk, Teaches Them to Trade too -- Regime Adaptive Execution using Informed Data and LLMs 20 Jun 2024 · 0 repositories · arXiv:2406.15508
-
A Large Language Model Outperforms Other Computational Approaches to the High-Throughput Phenotyping of Physician Notes 20 Jun 2024 · 0 repositories · arXiv:2406.14757
-
QPaug: Question and Passage Augmentation for Open-Domain Question Answering of LLMs 20 Jun 2024 · 1 repository · arXiv:2406.14277Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Automatic Labels are as Effective as Manual Labels in Biomedical Images Classification with Deep Learning 20 Jun 2024 · 1 repository · arXiv:2406.14351
-
ChatGPT as Research Scientist: Probing GPT's Capabilities as a Research Librarian, Research Ethicist, Data Generator and Data Predictor 20 Jun 2024 · 0 repositories · arXiv:2406.14765
-
CMTNet: Convolutional Meets Transformer Network for Hyperspectral Images Classification 20 Jun 2024 · 0 repositories · arXiv:2406.14080
-
CodeRAG-Bench: Can Retrieval Augment Code Generation? 20 Jun 2024 · 1 repository · arXiv:2406.14497Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Complexity of Symbolic Representation in Working Memory of Transformer Correlates with the Complexity of a Task 20 Jun 2024 · 0 repositories · arXiv:2406.14213
-
CryptoGPT: a 7B model rivaling GPT-4 in the task of analyzing and classifying real-time financial news 20 Jun 2024 · 0 repositories · arXiv:2406.14039
-
DIRAS: Efficient LLM Annotation of Document Relevance in Retrieval Augmented Generation 20 Jun 2024 · 1 repository · arXiv:2406.14162
-
Enhancing the LLM-Based Robot Manipulation Through Human-Robot Collaboration 20 Jun 2024 · 0 repositories · arXiv:2406.14097
-
Evaluating Implicit Bias in Large Language Models by Attacking From a Psychometric Perspective 20 Jun 2024 · 1 repository · arXiv:2406.14023
-
Evaluating RAG-Fusion with RAGElo: an Automated Elo-based Framework 20 Jun 2024 · 1 repository · arXiv:2406.14783Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Factual Dialogue Summarization via Learning from Large Language Models 20 Jun 2024 · 0 repositories · arXiv:2406.14709
-
Generative AI for Enhancing Active Learning in Education: A Comparative Study of GPT-3.5 and GPT-4 in Crafting Customized Test Questions 20 Jun 2024 · 0 repositories · arXiv:2406.13903
-
How to Compute the Probability of a Word 20 Jun 2024 · 2 repositories · arXiv:2406.14561Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Identifying User Goals from UI Trajectories 20 Jun 2024 · 0 repositories · arXiv:2406.14314
-
Learning to Plan for Retrieval-Augmented Large Language Models from Knowledge Graphs 20 Jun 2024 · 1 repository · arXiv:2406.14282Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
LLaSA: A Multimodal LLM for Human Activity Analysis Through Wearable and Smartphone Sensors 20 Jun 2024 · 1 repository · arXiv:2406.14498Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
MM-GTUNets: Unified Multi-Modal Graph Deep Learning for Brain Disorders Prediction 20 Jun 2024 · 1 repository · arXiv:2406.14455
-
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding 20 Jun 2024 · 1 repository · arXiv:2406.14515Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMs 20 Jun 2024 · 0 repositories · arXiv:2406.13975
-
Persuasiveness of Generated Free-Text Rationales in Subjective Decisions: A Case Study on Pairwise Argument Ranking 20 Jun 2024 · 1 repository · arXiv:2406.13905
-
Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs 20 Jun 2024 · 1 repository · arXiv:2406.14544Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 13 harvested samples)
-
Relation Extraction with Fine-Tuned Large Language Models in Retrieval Augmented Generation Frameworks 20 Jun 2024 · 0 repositories · arXiv:2406.14745
-
RTFormer: Re-parameter TSBN Spiking Transformer 20 Jun 2024 · 0 repositories · arXiv:2406.14180
-
SciDMT: A Large-Scale Corpus for Detecting Scientific Mentions 20 Jun 2024 · 0 repositories · arXiv:2406.14756
-
Seg-LSTM: Performance of xLSTM for Semantic Segmentation of Remotely Sensed Images 20 Jun 2024 · 1 repository · arXiv:2406.14086