Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 53
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 53 of 190: papers 5,201 to 5,300 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Beyond Benchmarks: Evaluating Embedding Model Similarity for Retrieval Augmented Generation Systems 11 Jul 2024 · 1 repository · arXiv:2407.08275
-
Brain Tumor Segmentation in MRI Images with 3D U-Net and Contextual Transformer 11 Jul 2024 · 0 repositories · arXiv:2407.08470
-
Converging Paradigms: The Synergy of Symbolic and Connectionist AI in LLM-Empowered Autonomous Agents 11 Jul 2024 · 0 repositories · arXiv:2407.08516
-
Fault Diagnosis in Power Grids with Large Language Model 11 Jul 2024 · 0 repositories · arXiv:2407.08836
-
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision 11 Jul 2024 · 2 repositories · arXiv:2407.08608Syntology official (archive's flag): 16 ran · 16 ran (of which 0 constructed an object rather than computing a result; 15 with no instrument failure: 2 honoured, 0 violated, 13 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 18 harvested samples) · 10 pointer-only (licence)
-
GPT-4 is judged more human than humans in displaced and inverted Turing tests 11 Jul 2024 · 0 repositories · arXiv:2407.08853
-
GraphMamba: An Efficient Graph Structure Learning Vision Mamba for Hyperspectral Image Classification 11 Jul 2024 · 1 repository · arXiv:2407.08255
-
GTA: A Benchmark for General Tool Agents 11 Jul 2024 · 1 repository · arXiv:2407.08713
-
HDT: Hierarchical Document Transformer 11 Jul 2024 · 0 repositories · arXiv:2407.08330
-
Investigating LLMs as Voting Assistants via Contextual Augmentation: A Case Study on the European Parliament Elections 2024 11 Jul 2024 · 0 repositories · arXiv:2407.08495
-
LLMs' morphological analyses of complex FST-generated Finnish words 11 Jul 2024 · 1 repository · arXiv:2407.08269
-
MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine 11 Jul 2024 · 3 repositories · arXiv:2407.08739
-
Multimodal contrastive learning for spatial gene expression prediction using histology images 11 Jul 2024 · 1 repository · arXiv:2407.08216Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 1 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
On the (In)Security of LLM App Stores 11 Jul 2024 · 0 repositories · arXiv:2407.08422
-
Projecting Points to Axes: Oriented Object Detection via Point-Axis Representation 11 Jul 2024 · 1 repository · arXiv:2407.08489
-
Real-Time Anomaly Detection and Reactive Planning with Large Language Models 11 Jul 2024 · 0 repositories · arXiv:2407.08735
-
SALSA: Swift Adaptive Lightweight Self-Attention for Enhanced LiDAR Place Recognition 11 Jul 2024 · 1 repository · arXiv:2407.08260
-
Skywork-Math: Data Scaling Laws for Mathematical Reasoning in Large Language Models -- The Story Goes On 11 Jul 2024 · 0 repositories · arXiv:2407.08348
-
Speculative RAG: Enhancing Retrieval Augmented Generation through Drafting 11 Jul 2024 · 0 repositories · arXiv:2407.08223
-
Spiking Tucker Fusion Transformer for Audio-Visual Zero-Shot Learning 11 Jul 2024 · 0 repositories · arXiv:2407.08130
-
stEnTrans: Transformer-based deep learning for spatial transcriptomics enhancement 11 Jul 2024 · 1 repository · arXiv:2407.08224
-
Synthetic Electroretinogram Signal Generation Using Conditional Generative Adversarial Network for Enhancing Classification of Autism Spectrum Disorder 11 Jul 2024 · 0 repositories · arXiv:2407.08166
-
TractGraphFormer: Anatomically Informed Hybrid Graph CNN-Transformer Network for Classification from Diffusion MRI Tractography 11 Jul 2024 · 0 repositories · arXiv:2407.08883
-
Vox Populi, Vox AI? Using Language Models to Estimate German Public Opinion 11 Jul 2024 · 1 repository · arXiv:2407.08563
-
Arabic Automatic Story Generation with Large Language Models 10 Jul 2024 · 1 repository · arXiv:2407.07551
-
Attribute or Abstain: Large Language Models as Long Document Assistants 10 Jul 2024 · 1 repository · arXiv:2407.07799Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Deep(er) Reconstruction of Imaging Cherenkov Detectors with Swin Transformers and Normalizing Flow Models 10 Jul 2024 · 1 repository · arXiv:2407.07376
-
Evaluating Large Language Models with Grid-Based Game Competitions: An Extensible LLM Benchmark and Leaderboard 10 Jul 2024 · 1 repository · arXiv:2407.07796
-
FACTS About Building Retrieval Augmented Generation-based Chatbots 10 Jul 2024 · 0 repositories · arXiv:2407.07858
-
Real World Federated Learning with a Knowledge Distilled Transformer for Cardiac CT Imaging 10 Jul 2024 · 2 repositories · arXiv:2407.07557
-
FsPONER: Few-shot Prompt Optimization for Named Entity Recognition in Domain-specific Scenarios 10 Jul 2024 · 1 repository · arXiv:2407.08035
-
H-FCBFormer Hierarchical Fully Convolutional Branch Transformer for Occlusal Contact Segmentation with Articulating Paper 10 Jul 2024 · 1 repository · arXiv:2407.07604
-
HAFormer: Unleashing the Power of Hierarchy-Aware Features for Lightweight Semantic Segmentation 10 Jul 2024 · 0 repositories · arXiv:2407.07441
-
KpopMT: Translation Dataset with Terminology for Kpop Fandom 10 Jul 2024 · 1 repository · arXiv:2407.07413
-
LitSearch: A Retrieval Benchmark for Scientific Literature Search 10 Jul 2024 · 1 repository · arXiv:2407.18940Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples)
-
A Guide To Effectively Leveraging LLMs for Low-Resource Text Summarization: Data Augmentation and Semi-supervised Approaches 10 Jul 2024 · 0 repositories · arXiv:2407.07341
-
Multilingual Blending: LLM Safety Alignment Evaluation with Language Mixture 10 Jul 2024 · 0 repositories · arXiv:2407.07342
-
Probability of Differentiation Reveals Brittleness of Homogeneity Bias in GPT-4 10 Jul 2024 · 0 repositories · arXiv:2407.07329
-
Examining Long-Context Large Language Models for Environmental Review Document Comprehension 10 Jul 2024 · 0 repositories · arXiv:2407.07321
-
ROSA: Random Subspace Adaptation for Efficient Fine-Tuning 10 Jul 2024 · 1 repository · arXiv:2407.07802
-
RT-LA-VocE: Real-Time Low-SNR Audio-Visual Speech Enhancement 10 Jul 2024 · 0 repositories · arXiv:2407.07825
-
Swin SMT: Global Sequential Modeling in 3D Medical Image Segmentation 10 Jul 2024 · 1 repository · arXiv:2407.07514
-
Teaching Transformers Causal Reasoning through Axiomatic Training 10 Jul 2024 · 0 repositories · arXiv:2407.07612
-
Toto: Time Series Optimized Transformer for Observability 10 Jul 2024 · 0 repositories · arXiv:2407.07874
-
Video In-context Learning 10 Jul 2024 · 0 repositories · arXiv:2407.07356
-
When to Accept Automated Predictions and When to Defer to Human Judgment? 10 Jul 2024 · 0 repositories · arXiv:2407.07821
-
WorldAPIs: The World Is Worth How Many APIs? A Thought Experiment 10 Jul 2024 · 0 repositories · arXiv:2407.07778
-
A Predictive Model Based on Transformer with Statistical Feature Embedding in Manufacturing Sensor Dataset 9 Jul 2024 · 0 repositories · arXiv:2407.06682
-
A Simple Architecture for Enterprise Large Language Model Applications based on Role based security and Clearance Levels using Retrieval-Augmented Generation or Mixture of Experts 9 Jul 2024 · 0 repositories · arXiv:2407.06718
-
AI AI Bias: Large Language Models Favor Their Own Generated Content 9 Jul 2024 · 1 repository · arXiv:2407.12856
-
Automated Peer Reviewing in Paper SEA: Standardization, Evaluation, and Analysis 9 Jul 2024 · 1 repository · arXiv:2407.12857Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 7 unverified (of 14 harvested samples)
-
CAPformer: Compression-Aware Pre-trained Transformer for Low-Light Image Enhancement 9 Jul 2024 · 0 repositories · arXiv:2407.07056
-
ChatGPT Doesn't Trust Chargers Fans: Guardrail Sensitivity in Context 9 Jul 2024 · 1 repository · arXiv:2407.06866Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
ConvNLP: Image-based AI Text Detection 9 Jul 2024 · 0 repositories · arXiv:2407.07225
-
CorMulT: A Semi-supervised Modality Correlation-aware Multimodal Transformer for Sentiment Analysis 9 Jul 2024 · 0 repositories · arXiv:2407.07046
-
Fine-Tuning Attention Modules Only: Enhancing Weight Disentanglement in Task Arithmetic 9 Jul 2024 · 2 repositories · arXiv:2407.07089Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples)
-
Identification of emotions on Twitter during the 2022 electoral process in Colombia 9 Jul 2024 · 0 repositories · arXiv:2407.07258
-
Measuring Sustainability Intention of ESG Fund Disclosure using Few-Shot Learning 9 Jul 2024 · 0 repositories · arXiv:2407.06893
-
Mixture-of-Modules: Reinventing Transformers as Dynamic Assemblies of Modules 9 Jul 2024 · 0 repositories · arXiv:2407.06677
-
Parameter-Efficient and Memory-Efficient Tuning for Vision Transformer: A Disentangled Approach 9 Jul 2024 · 1 repository · arXiv:2407.06964
-
PEER: Expertizing Domain-Specific Tasks with a Multi-Agent Framework and Tuning Methods 9 Jul 2024 · 1 repository · arXiv:2407.06985
-
Prompting Techniques for Secure Code Generation: A Systematic Investigation 9 Jul 2024 · 0 repositories · arXiv:2407.07064
-
Raply: A profanity-mitigated rap generator 9 Jul 2024 · 0 repositories · arXiv:2407.06941
-
Segment-Based Interactive Machine Translation for Pre-trained Models 9 Jul 2024 · 0 repositories · arXiv:2407.06990
-
Solving General Natural-Language-Description Optimization Problems with Large Language Models 9 Jul 2024 · 0 repositories · arXiv:2407.07924
-
Source Code Summarization in the Era of Large Language Models 9 Jul 2024 · 1 repository · arXiv:2407.07959
-
TrackFormers: In Search of Transformer-Based Particle Tracking for the High-Luminosity LHC Era 9 Jul 2024 · 1 repository · arXiv:2407.07179
-
Using Large Language Models for Generating Smart Contracts for Health Insurance from Textual Policies 9 Jul 2024 · 0 repositories · arXiv:2407.07019
-
Using Pretrained Large Language Model with Prompt Engineering to Answer Biomedical Questions 9 Jul 2024 · 0 repositories · arXiv:2407.06779
-
3D Vision and Language Pretraining with Large-Scale Synthetic Data 8 Jul 2024 · 1 repository · arXiv:2407.06084Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples) · 1 pointer-only (licence)
-
SOLO: A Single Transformer for Scalable Vision-Language Modeling 8 Jul 2024 · 1 repository · arXiv:2407.06438Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
CharSS: Character-Level Transformer Model for Sanskrit Word Segmentation 8 Jul 2024 · 0 repositories · arXiv:2407.06331
-
CodeUpdateArena: Benchmarking Knowledge Editing on API Updates 8 Jul 2024 · 1 repository · arXiv:2407.06249
-
Cross-domain Few-shot In-context Learning for Enhancing Traffic Sign Recognition 8 Jul 2024 · 0 repositories · arXiv:2407.05814
-
Deep Learning-based Anomaly Detection and Log Analysis for Computer Networks 8 Jul 2024 · 0 repositories · arXiv:2407.05639
-
Fast On-device LLM Inference with NPUs 8 Jul 2024 · 1 repository · arXiv:2407.05858Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples)
-
Generative Debunking of Climate Misinformation 8 Jul 2024 · 0 repositories · arXiv:2407.05599
-
InverseCoder: Self-improving Instruction-Tuned Code LLMs with Inverse-Instruct 8 Jul 2024 · 1 repository · arXiv:2407.05700Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
Large Language Models Understand Layout 8 Jul 2024 · 1 repository · arXiv:2407.05750
-
Learning Lane Graphs from Aerial Imagery Using Transformers 8 Jul 2024 · 0 repositories · arXiv:2407.05687
-
Controllable and Reliable Knowledge-Intensive Task-Oriented Conversational Agents with Declarative Genie Worksheets 8 Jul 2024 · 1 repository · arXiv:2407.05674
-
Meme Analysis using LLM-based Contextual Information and U-net Encapsulated Transformer 8 Jul 2024 · 1 repository
-
MSTF: Multiscale Transformer for Incomplete Trajectory Prediction 8 Jul 2024 · 0 repositories · arXiv:2407.05671
-
Multi-Label Plant Species Classification with Self-Supervised Vision Transformers 8 Jul 2024 · 1 repository · arXiv:2407.06298
-
On the Power of Convolution Augmented Transformer 8 Jul 2024 · 0 repositories · arXiv:2407.05591
-
Potential of Multimodal Large Language Models for Data Mining of Medical Images and Free-text Reports 8 Jul 2024 · 0 repositories · arXiv:2407.05758
-
Pruning Large Language Models to Intra-module Low-rank Architecture with Transitional Activations 8 Jul 2024 · 1 repository · arXiv:2407.05690
-
STMR: Spiral Transformer for Hand Mesh Reconstruction 8 Jul 2024 · 0 repositories · arXiv:2407.05967
-
Surprising gender biases in GPT 8 Jul 2024 · 0 repositories · arXiv:2407.06003
-
T2VSafetyBench: Evaluating the Safety of Text-to-Video Generative Models 8 Jul 2024 · 0 repositories · arXiv:2407.05965
-
Tailor3D: Customized 3D Assets Editing and Generation with Dual-Side Images 8 Jul 2024 · 0 repositories · arXiv:2407.06191
-
Towards Optimizing and Evaluating a Retrieval Augmented QA Chatbot using LLMs with Human in the Loop 8 Jul 2024 · 0 repositories · arXiv:2407.05925
-
TransMA: an explainable multi-modal deep learning model for predicting properties of ionizable lipid nanoparticles in mRNA delivery 8 Jul 2024 · 1 repository · arXiv:2407.05736
-
Vision-Braille: An End-to-End Tool for Chinese Braille Image-to-Text Translation 8 Jul 2024 · 0 repositories · arXiv:2407.06048
-
WSI-VQA: Interpreting Whole Slide Images by Generative Visual Question Answering 8 Jul 2024 · 1 repository · arXiv:2407.05603Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Enhancing Computer Programming Education with LLMs: A Study on Effective Prompt Engineering for Python Code Generation 7 Jul 2024 · 0 repositories · arXiv:2407.05437
-
Just read twice: closing the recall gap for recurrent language models 7 Jul 2024 · 1 repository · arXiv:2407.05483Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Large Language Model as an Assignment Evaluator: Insights, Feedback, and Challenges in a 1000+ Student Course 7 Jul 2024 · 0 repositories · arXiv:2407.05216
-
Learning Motion Blur Robust Vision Transformers with Dynamic Early Exit for Real-Time UAV Tracking 7 Jul 2024 · 0 repositories · arXiv:2407.05383
-
Mamba Hawkes Process 7 Jul 2024 · 0 repositories · arXiv:2407.05302