Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 24
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 24 of 190: papers 2,301 to 2,400 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Generative Artificial Intelligence-Supported Pentesting: A Comparison between Claude Opus, GPT-4, and Copilot 12 Jan 2025 · 0 repositories · arXiv:2501.06963
-
MFConvTr: Multi-Frequency Convolutional Transformer for Fetal Arrhythmia Detection in Non-Invasive fECG 12 Jan 2025 · 0 repositories · arXiv:2501.16649
-
MiniRAG: Towards Extremely Simple Retrieval-Augmented Generation 12 Jan 2025 · 1 repository · arXiv:2501.06713Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Transforming Vision Transformer: Towards Efficient Multi-Task Asynchronous Learning 12 Jan 2025 · 1 repository · arXiv:2501.06884Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
ZNO-Eval: Benchmarking reasoning capabilities of large language models in Ukrainian 12 Jan 2025 · 1 repository · arXiv:2501.06715
-
A Comparative Performance Analysis of Classification and Segmentation Models on Bangladeshi Pothole Dataset 11 Jan 2025 · 0 repositories · arXiv:2501.06602
-
Assessing instructor-AI cooperation for grading essay-type questions in an introductory sociology course 11 Jan 2025 · 1 repository · arXiv:2501.06461
-
CeViT: Copula-Enhanced Vision Transformer in multi-task learning and bi-group image covariates with an application to myopia screening 11 Jan 2025 · 1 repository · arXiv:2501.06540
-
First Token Probability Guided RAG for Telecom Question Answering 11 Jan 2025 · 0 repositories · arXiv:2501.06468
-
Flash Window Attention: speedup the attention computation for Swin Transformer 11 Jan 2025 · 2 repositories · arXiv:2501.06480
-
FocusDD: Real-World Scene Infusion for Robust Dataset Distillation 11 Jan 2025 · 0 repositories · arXiv:2501.06405
-
Ladder-residual: parallelism-aware architecture for accelerating large model inference with communication overlapping 11 Jan 2025 · 1 repository · arXiv:2501.06589Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 5 where Syntology's instrument failed) · 1 unverified (of 10 harvested samples) · 1 pointer-only (licence)
-
Tensor Product Attention Is All You Need 11 Jan 2025 · 1 repository · arXiv:2501.06425Syntology official (archive's flag): 5 ran · 8 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 5 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples) · 2 pointer-only (licence)
-
Ultra Memory-Efficient On-FPGA Training of Transformers via Tensor-Compressed Optimization 11 Jan 2025 · 0 repositories · arXiv:2501.06663
-
A Holistically Point-guided Text Framework for Weakly-Supervised Camouflaged Object Detection 10 Jan 2025 · 0 repositories · arXiv:2501.06038
-
An Attention-Guided Deep Learning Approach for Classifying 39 Skin Lesion Types 10 Jan 2025 · 1 repository · arXiv:2501.05991
-
Analyzing Spatio-Temporal Dynamics of Dissolved Oxygen for the River Thames using Superstatistical Methods and Machine Learning 10 Jan 2025 · 0 repositories · arXiv:2501.07599
-
Binary Event-Driven Spiking Transformer 10 Jan 2025 · 0 repositories · arXiv:2501.05904
-
Bridging Dialects: Translating Standard Bangla to Regional Variants Using Neural Models 10 Jan 2025 · 0 repositories · arXiv:2501.05749
-
Iconicity in Large Language Models 10 Jan 2025 · 0 repositories · arXiv:2501.05643
-
Merging Feed-Forward Sublayers for Compressed Transformers 10 Jan 2025 · 1 repository · arXiv:2501.06126
-
Mix-QViT: Mixed-Precision Vision Transformer Quantization Driven by Layer Importance and Quantization Sensitivity 10 Jan 2025 · 0 repositories · arXiv:2501.06357
-
Model Inversion in Split Learning for Personalized LLMs: New Insights from Information Bottleneck Theory 10 Jan 2025 · 0 repositories · arXiv:2501.05965
-
MSCViT: A Small-size ViT architecture with Multi-Scale Self-Attention Mechanism for Tiny Datasets 10 Jan 2025 · 0 repositories · arXiv:2501.06040
-
Multi-subject Open-set Personalization in Video Generation 10 Jan 2025 · 0 repositories · arXiv:2501.06187
-
Swin-X2S: Reconstructing 3D Shape from 2D Biplanar X-ray with Swin Transformers 10 Jan 2025 · 0 repositories · arXiv:2501.05961
-
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer 10 Jan 2025 · 0 repositories · arXiv:2501.06320
-
VideoRAG: Retrieval-Augmented Generation over Video Corpus 10 Jan 2025 · 1 repository · arXiv:2501.05874Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Weakly Supervised Segmentation of Hyper-Reflective Foci with Compact Convolutional Transformers and SAM2 10 Jan 2025 · 0 repositories · arXiv:2501.05933
-
A General Retrieval-Augmented Generation Framework for Multimodal Case-Based Reasoning Applications 9 Jan 2025 · 0 repositories · arXiv:2501.05030
-
Biomedical Relation Extraction via Adaptive Document-Relation Cross-Mapping and Concept Unique Identifier 9 Jan 2025 · 0 repositories · arXiv:2501.05155
-
Large language models streamline automated systematic review: A preliminary study 9 Jan 2025 · 0 repositories · arXiv:2502.15702
-
LLMQuoter: Enhancing RAG Capabilities Through Efficient Quote Extraction From Large Contexts 9 Jan 2025 · 1 repository · arXiv:2501.05554
-
LongViTU: Instruction Tuning for Long-Form Video Understanding 9 Jan 2025 · 0 repositories · arXiv:2501.05037
-
OpenAI ChatGPT interprets Radiological Images: GPT-4 as a Medical Doctor for a Fast Check-Up 9 Jan 2025 · 0 repositories · arXiv:2501.06269
-
Optimizing Multitask Industrial Processes with Predictive Action Guidance 9 Jan 2025 · 0 repositories · arXiv:2501.05108
-
RAG-WM: An Efficient Black-Box Watermarking Approach for Retrieval-Augmented Generation of Large Language Models 9 Jan 2025 · 0 repositories · arXiv:2501.05249
-
SpecTf: Transformers Enable Data-Driven Imaging Spectroscopy Cloud Detection 9 Jan 2025 · 1 repository · arXiv:2501.04916
-
The dynamics of meaning through time: Assessment of Large Language Models 9 Jan 2025 · 0 repositories · arXiv:2501.05552
-
UAV-VLA: Vision-Language-Action System for Large Scale Aerial Mission Generation 9 Jan 2025 · 1 repository · arXiv:2501.05014
-
A partition cover approach to tokenization 8 Jan 2025 · 1 repository · arXiv:2501.06246
-
Advancing Retrieval-Augmented Generation for Persian: Development of Language Models, Comprehensive Benchmarks, and Best Practices for Optimization 8 Jan 2025 · 0 repositories · arXiv:2501.04858
-
Circuit Complexity Bounds for Visual Autoregressive Model 8 Jan 2025 · 0 repositories · arXiv:2501.04299
-
Integrating LLMs with ITS: Recent Advances, Potentials, Challenges, and Future Directions 8 Jan 2025 · 0 repositories · arXiv:2501.04437
-
Knowledge Retrieval Based on Generative AI 8 Jan 2025 · 0 repositories · arXiv:2501.04635
-
MB-TaylorFormer V2: Improved Multi-branch Linear Transformer Expanded by Taylor Formula for Image Restoration 8 Jan 2025 · 2 repositories · arXiv:2501.04486
-
Multi-task retriever fine-tuning for domain-specific and efficient RAG 8 Jan 2025 · 0 repositories · arXiv:2501.04652
-
Re-ranking the Context for Multimodal Retrieval Augmented Generation 8 Jan 2025 · 0 repositories · arXiv:2501.04695
-
Scaling Large Language Model Training on Frontier with Low-Bandwidth Partitioning 8 Jan 2025 · 0 repositories · arXiv:2501.04266
-
AuxDepthNet: Real-Time Monocular 3D Object Detection with Depth-Sensitive Features 7 Jan 2025 · 0 repositories · arXiv:2501.03700
-
CFFormer: Cross CNN-Transformer Channel Attention and Spatial Feature Fusion for Improved Segmentation of Low Quality Medical Images 7 Jan 2025 · 0 repositories · arXiv:2501.03629
-
Efficient and Accurate Tuberculosis Diagnosis: Attention Residual U-Net and Vision Transformer Based Detection Framework 7 Jan 2025 · 0 repositories · arXiv:2501.03538
-
Finding A Voice: Evaluating African American Dialect Generation for Chatbot Technology 7 Jan 2025 · 1 repository · arXiv:2501.03441
-
How to Select Pre-Trained Code Models for Reuse? A Learning Perspective 7 Jan 2025 · 1 repository · arXiv:2501.03783
-
Language and Planning in Robotic Navigation: A Multilingual Evaluation of State-of-the-Art Models 7 Jan 2025 · 0 repositories · arXiv:2501.05478
-
LM-Net: A Light-weight and Multi-scale Network for Medical Image Segmentation 7 Jan 2025 · 1 repository · arXiv:2501.03838
-
MTRAG: A Multi-Turn Conversational Benchmark for Evaluating Retrieval-Augmented Generation Systems 7 Jan 2025 · 1 repository · arXiv:2501.03468
-
Practical Design and Benchmarking of Generative AI Applications for Surgical Billing and Coding 7 Jan 2025 · 0 repositories · arXiv:2501.05479
-
RAG-Check: Evaluating Multimodal Retrieval Augmented Generation Performance 7 Jan 2025 · 0 repositories · arXiv:2501.03995
-
Reading with Intent -- Neutralizing Intent 7 Jan 2025 · 0 repositories · arXiv:2501.03475
-
SNR-EQ-JSCC: Joint Source-Channel Coding with SNR-Based Embedding and Query 7 Jan 2025 · 0 repositories · arXiv:2501.04732
-
Three-dimensional attention Transformer for state evaluation in real-time strategy games 7 Jan 2025 · 0 repositories · arXiv:2501.03832
-
A Novel Vision Transformer for Camera-LiDAR Fusion based Traffic Object Segmentation 6 Jan 2025 · 0 repositories · arXiv:2501.02858
-
Adaptive Pruning of Pretrained Transformer via Differential Inclusions 6 Jan 2025 · 0 repositories · arXiv:2501.03289
-
CHAT: Beyond Contrastive Graph Transformer for Link Prediction in Heterogeneous Networks 6 Jan 2025 · 0 repositories · arXiv:2501.02760
-
Developing an Artificial Intelligence Tool for Personalized Breast Cancer Treatment Plans based on the NCCN Guidelines 6 Jan 2025 · 0 repositories · arXiv:2502.15698
-
Intelligent logistics management robot path planning algorithm integrating transformer and GCN network 6 Jan 2025 · 0 repositories · arXiv:2501.02749
-
FlippedRAG: Black-Box Opinion Manipulation Adversarial Attacks to Retrieval-Augmented Generation Models 6 Jan 2025 · 0 repositories · arXiv:2501.02968
-
GLoG-CSUnet: Enhancing Vision Transformers with Adaptable Radiomic Features for Medical Image Segmentation 6 Jan 2025 · 1 repository · arXiv:2501.02788
-
Integrating Language-Image Prior into EEG Decoding for Cross-Task Zero-Calibration RSVP-BCI 6 Jan 2025 · 0 repositories · arXiv:2501.02841
-
Mixture-of-Experts Graph Transformers for Interpretable Particle Collision Detection 6 Jan 2025 · 1 repository · arXiv:2501.03432
-
Political Events using RAG with LLMs 6 Jan 2025 · 0 repositories · arXiv:2502.15701
-
QuIM-RAG: Advancing Retrieval-Augmented Generation with Inverted Question Matching for Enhanced QA Performance 6 Jan 2025 · 0 repositories · arXiv:2501.02702
-
SALT: Sales Autocompletion Linked Business Tables Dataset 6 Jan 2025 · 1 repository · arXiv:2501.03413
-
Sensorformer: Cross-patch attention with global-patch compression is effective for high-dimensional multivariate time series forecasting 6 Jan 2025 · 0 repositories · arXiv:2501.03284
-
Sequence Complementor: Complementing Transformers For Time Series Forecasting with Learnable Sequences 6 Jan 2025 · 0 repositories · arXiv:2501.02735
-
Tree-based RAG-Agent Recommendation System: A Case Study in Medical Test Data 6 Jan 2025 · 0 repositories · arXiv:2501.02727
-
VicSim: Enhancing Victim Simulation with Emotional and Linguistic Fidelity 6 Jan 2025 · 0 repositories · arXiv:2501.03139
-
Decoding fMRI Data into Captions using Prefix Language Modeling 5 Jan 2025 · 1 repository · arXiv:2501.02570
-
DeTrack: In-model Latent Denoising Learning for Visual Object Tracking 5 Jan 2025 · 0 repositories · arXiv:2501.02467
-
Empowering Bengali Education with AI: Solving Bengali Math Word Problems through Transformer Models 5 Jan 2025 · 0 repositories · arXiv:2501.02599
-
Evaluating Large Language Models Against Human Annotators in Latent Content Analysis: Sentiment, Political Leaning, Emotional Intensity, and Sarcasm 5 Jan 2025 · 0 repositories · arXiv:2501.02532
-
GS-DiT: Advancing Video Generation with Pseudo 4D Gaussian Fields through Efficient Dense 3D Point Tracking 5 Jan 2025 · 0 repositories · arXiv:2501.02690
-
HonkaiChat: Companions from Anime that feel alive! 5 Jan 2025 · 0 repositories · arXiv:2501.03277
-
LWFNet: Coherent Doppler Wind Lidar-Based Network for Wind Field Retrieval 5 Jan 2025 · 0 repositories · arXiv:2501.02613
-
Towards New Benchmark for AI Alignment & Sentiment Analysis in Socially Important Issues: A Comparative Study of Human and LLMs in the Context of AGI 5 Jan 2025 · 0 repositories · arXiv:2501.02531
-
Examining the Robustness of Homogeneity Bias to Hyperparameter Adjustments in GPT-4 4 Jan 2025 · 0 repositories · arXiv:2501.02211
-
Exploring the Capabilities and Limitations of Large Language Models for Radiation Oncology Decision Support 4 Jan 2025 · 0 repositories · arXiv:2501.02346
-
Graph-Aware Isomorphic Attention for Adaptive Dynamics in Transformers 4 Jan 2025 · 1 repository · arXiv:2501.02393
-
Knowledge Graph Retrieval-Augmented Generation for LLM-based Recommendation 4 Jan 2025 · 0 repositories · arXiv:2501.02226
-
The Application of Large Language Models in Recommendation Systems 4 Jan 2025 · 0 repositories · arXiv:2501.02178
-
A Separable Self-attention Inspired by the State Space Model for Computer Vision 3 Jan 2025 · 1 repository · arXiv:2501.02040
-
A Survey on Large Language Models with some Insights on their Capabilities and Limitations 3 Jan 2025 · 0 repositories · arXiv:2501.04040
-
AgentRefine: Enhancing Agent Generalization through Refinement Tuning 3 Jan 2025 · 0 repositories · arXiv:2501.01702
-
BARTPredict: Empowering IoT Security with LLM-Driven Cyber Threat Prediction 3 Jan 2025 · 0 repositories · arXiv:2501.01664
-
Classifier-Guided Captioning Across Modalities 3 Jan 2025 · 0 repositories · arXiv:2501.03183
-
End-to-End Long Document Summarization using Gradient Caching 3 Jan 2025 · 0 repositories · arXiv:2501.01805
-
LLMs & Legal Aid: Understanding Legal Needs Exhibited Through User Queries 3 Jan 2025 · 0 repositories · arXiv:2501.01711
-
MIRAGE: Exploring How Large Language Models Perform in Complex Social Interactive Environments 3 Jan 2025 · 1 repository · arXiv:2501.01652
-
PersonaAI: Leveraging Retrieval-Augmented Generation and Personalized Context for AI-Driven Digital Avatars 3 Jan 2025 · 0 repositories · arXiv:2503.15489