Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 63
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 63 of 190: papers 6,201 to 6,300 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Benchmarking Hierarchical Image Pyramid Transformer for the classification of colon biopsies and polyps in histopathology images 24 May 2024 · 0 repositories · arXiv:2405.15127
-
Benchmarking the Performance of Pre-trained LLMs across Urdu NLP Tasks 24 May 2024 · 0 repositories · arXiv:2405.15453
-
Continuously Learning, Adapting, and Improving: A Dual-Process Approach to Autonomous Driving 24 May 2024 · 1 repository · arXiv:2405.15324
-
ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models 24 May 2024 · 1 repository · arXiv:2405.15738Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples) · 1 pointer-only (licence)
-
CulturePark: Boosting Cross-cultural Understanding in Large Language Models 24 May 2024 · 1 repository · arXiv:2405.15145Syntology official: harvested, nothing ran · 0 ran · 7 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Distinguish Any Fake Videos: Unleashing the Power of Large-scale Data and Motion Features 24 May 2024 · 0 repositories · arXiv:2405.15343
-
Evaluating and Safeguarding the Adversarial Robustness of Retrieval-Based In-Context Learning 24 May 2024 · 1 repository · arXiv:2405.15984Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Filtered Corpus Training (FiCT) Shows that Language Models can Generalize from Indirect Evidence 24 May 2024 · 1 repository · arXiv:2405.15750
-
Generalizable and Scalable Multistage Biomedical Concept Normalization Leveraging Large Language Models 24 May 2024 · 1 repository · arXiv:2405.15122
-
GPT is Not an Annotator: The Necessity of Human Annotation in Fairness Benchmark Construction 24 May 2024 · 0 repositories · arXiv:2405.15760
-
Large Language Models Reflect Human Citation Patterns with a Heightened Citation Bias 24 May 2024 · 1 repository · arXiv:2405.15739
-
Learning the Language of Protein Structure 24 May 2024 · 1 repository · arXiv:2405.15840Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Machine Unlearning in Large Language Models 24 May 2024 · 1 repository · arXiv:2405.15152
-
MambaVC: Learned Visual Compression with Selective State Spaces 24 May 2024 · 1 repository · arXiv:2405.15413Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Matchings, Predictions and Counterfactual Harm in Refugee Resettlement Processes 24 May 2024 · 0 repositories · arXiv:2407.13052
-
MeMo: Meaningful, Modular Controllers via Noise Injection 24 May 2024 · 0 repositories · arXiv:2407.01567Syntology 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
MLPs Learn In-Context on Regression and Classification Tasks 24 May 2024 · 2 repositories · arXiv:2405.15618Syntology official (archive's flag): 2 ran · 15 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 16 harvested samples)
-
PoinTramba: A Hybrid Transformer-Mamba Framework for Point Cloud Analysis 24 May 2024 · 1 repository · arXiv:2405.15463Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 1 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples) · 7 pointer-only (licence)
-
Spectraformer: A Unified Random Feature Framework for Transformer 24 May 2024 · 2 repositories · arXiv:2405.15310
-
Steerable Transformers 24 May 2024 · 0 repositories · arXiv:2405.15932
-
Text-guided 3D Human Motion Generation with Keyframe-based Parallel Skip Transformer 24 May 2024 · 0 repositories · arXiv:2405.15439
-
Comet: A Communication-efficient and Performant Approximation for Private Transformer Inference 24 May 2024 · 0 repositories · arXiv:2405.17485
-
The Impact and Opportunities of Generative AI in Fact-Checking 24 May 2024 · 0 repositories · arXiv:2405.15985
-
Towards Better Understanding of In-Context Learning Ability from In-Context Uncertainty Quantification 24 May 2024 · 0 repositories · arXiv:2405.15115
-
The Buffer Mechanism for Multi-Step Information Reasoning in Language Models 24 May 2024 · 0 repositories · arXiv:2405.15302
-
UnitNorm: Rethinking Normalization for Transformers in Time Series 24 May 2024 · 0 repositories · arXiv:2405.15903
-
Zero-Shot Spam Email Classification Using Pre-trained Large Language Models 24 May 2024 · 0 repositories · arXiv:2405.15936
-
3D Learnable Supertoken Transformer for LiDAR Point Cloud Scene Segmentation 23 May 2024 · 0 repositories · arXiv:2405.15826
-
A Declarative System for Optimizing AI Workloads 23 May 2024 · 1 repository · arXiv:2405.14696
-
AGILE: A Novel Reinforcement Learning Framework of LLM Agents 23 May 2024 · 1 repository · arXiv:2405.14751Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
Attending to Topological Spaces: The Cellular Transformer 23 May 2024 · 0 repositories · arXiv:2405.14094
-
AutoCoder: Enhancing Code Large Language Model with AIEV-Instruct 23 May 2024 · 1 repository · arXiv:2405.14906
-
Combining Denoising Autoencoders with Contrastive Learning to fine-tune Transformer Models 23 May 2024 · 1 repository · arXiv:2405.14437
-
DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data 23 May 2024 · 0 repositories · arXiv:2405.14333
-
Dinomaly: The Less Is More Philosophy in Multi-Class Unsupervised Anomaly Detection 23 May 2024 · 2 repositories · arXiv:2405.14325Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 5 where Syntology's instrument failed) · 1 unverified (of 10 harvested samples) · 5 pointer-only (licence)
-
Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion Transformer 23 May 2024 · 1 repository · arXiv:2405.14832Syntology 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples)
-
EditWorld: Simulating World Dynamics for Instruction-Following Image Editing 23 May 2024 · 1 repository · arXiv:2405.14785Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Efficient Medical Question Answering with Knowledge-Augmented Question Generation 23 May 2024 · 1 repository · arXiv:2405.14654
-
Efficient Point Transformer with Dynamic Token Aggregating for LiDAR Point Cloud Processing 23 May 2024 · 0 repositories · arXiv:2405.15827
-
Eliciting Informative Text Evaluations with Large Language Models 23 May 2024 · 1 repository · arXiv:2405.15077Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 14 harvested samples) · 14 pointer-only (licence)
-
Evaluating Large Language Models for Public Health Classification and Extraction Tasks 23 May 2024 · 0 repositories · arXiv:2405.14766
-
Exploring the use of a Large Language Model for data extraction in systematic reviews: a rapid feasibility study 23 May 2024 · 0 repositories · arXiv:2405.14445
-
From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step 23 May 2024 · 1 repository · arXiv:2405.14838Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language Models 23 May 2024 · 2 repositories · arXiv:2405.14831Syntology community repositories only · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Impact of Non-Standard Unicode Characters on Security and Comprehension in Large Language Models 23 May 2024 · 1 repository · arXiv:2405.14490
-
Improving Gloss-free Sign Language Translation by Reducing Representation Density 23 May 2024 · 1 repository · arXiv:2405.14312Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Improving Language Models Trained on Translated Data with Continual Pre-Training and Dictionary Learning Analysis 23 May 2024 · 0 repositories · arXiv:2405.14277
-
JiuZhang3.0: Efficiently Improving Mathematical Reasoning by Training Small Data Synthesis Models 23 May 2024 · 1 repository · arXiv:2405.14365Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
Large Language Models Can Self-Correct with Key Condition Verification 23 May 2024 · 0 repositories · arXiv:2405.14092
-
Leveraging Semantic Segmentation Masks with Embeddings for Fine-Grained Form Classification 23 May 2024 · 0 repositories · arXiv:2405.14162
-
Linking In-context Learning in Transformers to Human Episodic Memory 23 May 2024 · 1 repository · arXiv:2405.14992Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Lorentz-Equivariant Geometric Algebra Transformers for High-Energy Physics 23 May 2024 · 1 repository · arXiv:2405.14806Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Magnetic Resonance Image Processing Transformer for General Accelerated Image Reconstruction 23 May 2024 · 0 repositories · arXiv:2405.15098
-
Not All Language Model Features Are Linear 23 May 2024 · 1 repository · arXiv:2405.14860Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples)
-
Optimizing example selection for retrieval-augmented machine translation with translation memories 23 May 2024 · 0 repositories · arXiv:2405.15070
-
Perception of Knowledge Boundary for Large Language Models through Semi-open-ended Question Answering 23 May 2024 · 0 repositories · arXiv:2405.14383
-
Parameter-free Clipped Gradient Descent Meets Polyak 23 May 2024 · 0 repositories · arXiv:2405.15010
-
PrivCirNet: Efficient Private Inference via Block Circulant Transformation 23 May 2024 · 1 repository · arXiv:2405.14569Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
PuTR: A Pure Transformer for Decoupled and Online Multi-Object Tracking 23 May 2024 · 1 repository · arXiv:2405.14119
-
RaFe: Ranking Feedback Improves Query Rewriting for RAG 23 May 2024 · 0 repositories · arXiv:2405.14431
-
Scalable Visual State Space Model with Fractal Scanning 23 May 2024 · 0 repositories · arXiv:2405.14480
-
ShapeFormer: Shapelet Transformer for Multivariate Time Series Classification 23 May 2024 · 0 repositories · arXiv:2405.14608
-
Sparse-Tuning: Adapting Vision Transformers with Efficient Fine-tuning and Inference 23 May 2024 · 1 repository · arXiv:2405.14700Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 1 honoured, 0 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
Transformers for Image-Goal Navigation 23 May 2024 · 0 repositories · arXiv:2405.14128
-
WISE: Rethinking the Knowledge Memory for Lifelong Model Editing of Large Language Models 23 May 2024 · 1 repository · arXiv:2405.14768Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 9 where Syntology's instrument failed) · 3 unverified (of 16 harvested samples)
-
A Transformer variant for multi-step forecasting of water level and hydrometeorological sensitivity analysis based on explainable artificial intelligence technology 22 May 2024 · 0 repositories · arXiv:2405.13646
-
A General Graph Spectral Wavelet Convolution via Chebyshev Order Decomposition 22 May 2024 · 1 repository · arXiv:2405.13806Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
Affine-based Deformable Attention and Selective Fusion for Semi-dense Matching 22 May 2024 · 0 repositories · arXiv:2405.13874
-
Automated Evaluation of Retrieval-Augmented Language Models with Task-Specific Exam Generation 22 May 2024 · 1 repository · arXiv:2405.13622
-
Automatically Identifying Local and Global Circuits with Linear Computation Graphs 22 May 2024 · 0 repositories · arXiv:2405.13868
-
CViT: Continuous Vision Transformer for Operator Learning 22 May 2024 · 2 repositories · arXiv:2405.13998Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Comparative Analysis of Hyperspectral Image Reconstruction Using Deep Learning for Agricultural and Biological Applications 22 May 2024 · 0 repositories · arXiv:2405.13331
-
Discrete Cosine Transform Based Decorrelated Attention for Vision Transformers 22 May 2024 · 0 repositories · arXiv:2405.13901
-
Evaluating Large Language Models with Human Feedback: Establishing a Swedish Benchmark 22 May 2024 · 1 repository · arXiv:2405.14006
-
FlashRAG: A Modular Toolkit for Efficient Retrieval-Augmented Generation Research 22 May 2024 · 1 repository · arXiv:2405.13576
-
From CNNs to Transformers in Multimodal Human Action Recognition: A Survey 22 May 2024 · 0 repositories · arXiv:2405.15813
-
High Performance P300 Spellers Using GPT2 Word Prediction With Cross-Subject Training 22 May 2024 · 0 repositories · arXiv:2405.13329
-
KU-DMIS at EHRSQL 2024:Generating SQL query via question templatization in EHR 22 May 2024 · 0 repositories · arXiv:2406.00014
-
Leveraging 2D Information for Long-term Time Series Forecasting with Vanilla Transformers 22 May 2024 · 1 repository · arXiv:2405.13810
-
Unlocking the Power of Patch: Patch-Based MLP for Long-Term Time Series Forecasting 22 May 2024 · 0 repositories · arXiv:2405.13575
-
Semantic Equitable Clustering: A Simple and Effective Strategy for Clustering Vision Tokens 22 May 2024 · 0 repositories · arXiv:2405.13337
-
Task-agnostic Decision Transformer for Multi-type Agent Control with Federated Split Training 22 May 2024 · 0 repositories · arXiv:2405.13445
-
TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment 22 May 2024 · 1 repository · arXiv:2405.13911Syntology official (archive's flag): 3 ran · 6 ran (of which 3 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 1 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples) · 3 pointer-only (licence)
-
TrojanRAG: Retrieval-Augmented Generation Can Be Backdoor Driver in Large Language Models 22 May 2024 · 1 repository · arXiv:2405.13401
-
Unsupervised Pre-training with Language-Vision Prompts for Low-Data Instance Segmentation 22 May 2024 · 1 repository · arXiv:2405.13388
-
Why Not Transform Chat Large Language Models to Non-English? 22 May 2024 · 1 repository · arXiv:2405.13923
-
WordGame: Efficient & Effective LLM Jailbreak via Simultaneous Obfuscation in Query and Response 22 May 2024 · 0 repositories · arXiv:2405.14023
-
A Masked Semi-Supervised Learning Approach for Otago Micro Labels Recognition 21 May 2024 · 0 repositories · arXiv:2405.12711
-
BIMM: Brain Inspired Masked Modeling for Video Representation Learning 21 May 2024 · 1 repository · arXiv:2405.12757
-
BiomedParse: a biomedical foundation model for image parsing of everything everywhere all at once 21 May 2024 · 0 repositories · arXiv:2405.12971
-
Enhancing Transformer-based models for Long Sequence Time Series Forecasting via Structured Matrix 21 May 2024 · 1 repository · arXiv:2405.12462
-
Computational Tradeoffs in Image Synthesis: Diffusion, Masked-Token, and Next-Token Prediction 21 May 2024 · 0 repositories · arXiv:2405.13218
-
Generative AI in Cybersecurity: A Comprehensive Review of LLM Applications and Vulnerabilities 21 May 2024 · 0 repositories · arXiv:2405.12750
-
Global-Local Detail Guided Transformer for Sea Ice Recognition in Optical Remote Sensing Images 21 May 2024 · 0 repositories · arXiv:2405.13197
-
GPT-4 Jailbreaks Itself with Near-Perfect Success Using Self-Explanation 21 May 2024 · 0 repositories · arXiv:2405.13077
-
How Reliable AI Chatbots are for Disease Prediction from Patient Complaints? 21 May 2024 · 0 repositories · arXiv:2405.13219
-
Investigating Persuasion Techniques in Arabic: An Empirical Study Leveraging Large Language Models 21 May 2024 · 0 repositories · arXiv:2405.12884
-
Is Dataset Quality Still a Concern in Diagnosis Using Large Foundation Model? 21 May 2024 · 0 repositories · arXiv:2405.12584
-
Mamba in Speech: Towards an Alternative to Self-Attention 21 May 2024 · 1 repository · arXiv:2405.12609
-
Mitigating Overconfidence in Out-of-Distribution Detection by Capturing Extreme Activations 21 May 2024 · 1 repository · arXiv:2405.12658