Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 48
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 48 of 190: papers 4,701 to 4,800 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Generative AI for automatic topic labelling 13 Aug 2024 · 0 repositories · arXiv:2408.07003
-
Harnessing Earnings Reports for Stock Predictions: A QLoRA-Enhanced LLM Approach 13 Aug 2024 · 0 repositories · arXiv:2408.06634
-
Leveraging Language Models for Emotion and Behavior Analysis in Education 13 Aug 2024 · 0 repositories · arXiv:2408.06874
-
Optimal Preprocessing for Joint Detection and Classification of Wireless Communication Signals in Congested Spectrum Using Computer Vision Methods 13 Aug 2024 · 0 repositories · arXiv:2408.06545
-
Pragmatic inference of scalar implicature by LLMs 13 Aug 2024 · 0 repositories · arXiv:2408.06673
-
Spectrum Prediction With Deep 3D Pyramid Vision Transformer Learning 13 Aug 2024 · 1 repository · arXiv:2408.06870
-
Sumotosima: A Framework and Dataset for Classifying and Summarizing Otoscopic Images 13 Aug 2024 · 1 repository · arXiv:2408.06755
-
Unlocking Efficiency: Adaptive Masking for Gene Transformer Models 13 Aug 2024 · 1 repository · arXiv:2408.07180
-
Using Advanced LLMs to Enhance Smaller LLMs: An Interpretable Knowledge Distillation Approach 13 Aug 2024 · 0 repositories · arXiv:2408.07238
-
VulCatch: Enhancing Binary Vulnerability Detection through CodeT5 Decompilation and KAN Advanced Feature Extraction 13 Aug 2024 · 0 repositories · arXiv:2408.07181
-
Advanced Vision Transformers and Open-Set Learning for Robust Mosquito Classification: A Novel Approach to Entomological Studies 12 Aug 2024 · 0 repositories · arXiv:2408.06457
-
Bayesian inference to improve quality of Retrieval Augmented Generation 12 Aug 2024 · 0 repositories · arXiv:2408.08901
-
Body Transformer: Leveraging Robot Embodiment for Policy Learning 12 Aug 2024 · 0 repositories · arXiv:2408.06316
-
Cross-Lingual Conversational Speech Summarization with Large Language Models 12 Aug 2024 · 0 repositories · arXiv:2408.06484
-
DPDETR: Decoupled Position Detection Transformer for Infrared-Visible Object Detection 12 Aug 2024 · 0 repositories · arXiv:2408.06123
-
Enhancing 3D Transformer Segmentation Model for Medical Image with Token-level Representation Learning 12 Aug 2024 · 1 repository · arXiv:2408.05889
-
HAT: History-Augmented Anchor Transformer for Online Temporal Action Localization 12 Aug 2024 · 1 repository · arXiv:2408.06437Syntology official (archive's flag): 13 ran · 13 ran (of which 3 constructed an object rather than computing a result; 12 with no instrument failure: 1 honoured, 0 violated, 11 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 14 harvested samples) · 2 pointer-only (licence)
-
Med42-v2: A Suite of Clinical LLMs 12 Aug 2024 · 0 repositories · arXiv:2408.06142
-
Optimizing RAG Techniques for Automotive Industry PDF Chatbots: A Case Study with Locally Deployed Ollama Models 12 Aug 2024 · 0 repositories · arXiv:2408.05933
-
PAFormer: Part Aware Transformer for Person Re-identification 12 Aug 2024 · 0 repositories · arXiv:2408.05918
-
PhaGO: Protein function annotation for bacteriophages by integrating the genomic context 12 Aug 2024 · 0 repositories · arXiv:2408.06402
-
Spacetime E(n)-Transformer: Equivariant Attention for Spatio-temporal Graphs 12 Aug 2024 · 1 repository · arXiv:2408.06039
-
The Language of Trauma: Modeling Traumatic Event Descriptions Across Domains with Explainable AI 12 Aug 2024 · 0 repositories · arXiv:2408.05977
-
Utilize Transformers for translating Wikipedia category names 12 Aug 2024 · 0 repositories · arXiv:2408.06124
-
GPT-4 Emulates Average-Human Emotional Cognition from a Third-Person Perspective 11 Aug 2024 · 0 repositories · arXiv:2408.13718
-
HySparK: Hybrid Sparse Masking for Large Scale Medical Image Pre-Training 11 Aug 2024 · 1 repository · arXiv:2408.05815
-
Kov: Transferable and Naturalistic Black-Box LLM Attacks using Markov Decision Processes and Tree Search 11 Aug 2024 · 1 repository · arXiv:2408.08899
-
Sampling Foundational Transformer: A Theoretical Perspective 11 Aug 2024 · 0 repositories · arXiv:2408.05822
-
U-DECN: End-to-End Underwater Object Detection ConvNet with Improved DeNoising Training 11 Aug 2024 · 1 repository · arXiv:2408.05780
-
PhishLang: A Real-Time, Fully Client-Side Phishing Detection Framework Using MobileBERT 11 Aug 2024 · 2 repositories · arXiv:2408.05667
-
BeyondCT: A deep learning model for predicting pulmonary function from chest CT scans 10 Aug 2024 · 0 repositories · arXiv:2408.05645
-
Chain of Condition: Construct, Verify and Solve Conditions for Conditional Question Answering 10 Aug 2024 · 0 repositories · arXiv:2408.05442
-
Improving Whisper's Recognition Performance for Under-Represented Language Kazakh Leveraging Unpaired Speech and Text 10 Aug 2024 · 0 repositories · arXiv:2408.05554
-
Modeling Multi-Step Scientific Processes with Graph Transformer Networks 10 Aug 2024 · 0 repositories · arXiv:2408.05425
-
PersonViT: Large-scale Self-supervised Vision Transformer for Person Re-Identification 10 Aug 2024 · 1 repository · arXiv:2408.05398
-
PointMT: Efficient Point Cloud Analysis with Hybrid MLP-Transformer Architecture 10 Aug 2024 · 0 repositories · arXiv:2408.05508
-
SWIFT:A Scalable lightWeight Infrastructure for Fine-Tuning 10 Aug 2024 · 3 repositories · arXiv:2408.05517Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
A Hybrid RAG System with Comprehensive Enhancement on Complex Reasoning 9 Aug 2024 · 0 repositories · arXiv:2408.05141
-
ChatGPT Meets Iris Biometrics 9 Aug 2024 · 0 repositories · arXiv:2408.04868
-
ConfusedPilot: Confused Deputy Risks in RAG-based LLMs 9 Aug 2024 · 0 repositories · arXiv:2408.04870
-
DeepInteraction++: Multi-Modality Interaction for Autonomous Driving 9 Aug 2024 · 1 repository · arXiv:2408.05075Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
COAST: Enhancing the Code Debugging Ability of LLMs through Communicative Agent Based Data Synthesis 9 Aug 2024 · 1 repository · arXiv:2408.05006
-
Evaluating the capability of large language models to personalize science texts for diverse middle-school-age learners 9 Aug 2024 · 0 repositories · arXiv:2408.05204
-
Examining the Behavior of LLM Architectures Within the Framework of Standardized National Exams in Brazil 9 Aug 2024 · 0 repositories · arXiv:2408.05035
-
From Text to Insight: Leveraging Large Language Models for Performance Evaluation in Management 9 Aug 2024 · 0 repositories · arXiv:2408.05328
-
HybridRAG: Integrating Knowledge Graphs and Vector Retrieval Augmented Generation for Efficient Information Extraction 9 Aug 2024 · 0 repositories · arXiv:2408.04948
-
Large Language Models and Thematic Analysis: Human-AI Synergy in Researching Hate Speech on Social Media 9 Aug 2024 · 0 repositories · arXiv:2408.05126
-
LLMJudge: LLMs for Relevance Judgments 9 Aug 2024 · 1 repository · arXiv:2408.08896
-
MIDI-to-Tab: Guitar Tablature Inference via Masked Language Modeling 9 Aug 2024 · 0 repositories · arXiv:2408.05024
-
Rag and Roll: An End-to-End Evaluation of Indirect Prompt Manipulations in LLM-based Application Frameworks 9 Aug 2024 · 0 repositories · arXiv:2408.05025
-
Retrieval-augmented code completion for local projects using large language models 9 Aug 2024 · 0 repositories · arXiv:2408.05026
-
Attention Mechanism and Context Modeling System for Text Mining Machine Translation 8 Aug 2024 · 0 repositories · arXiv:2408.04216
-
Can GPT-4 Models Detect Misleading Visualizations? 8 Aug 2024 · 0 repositories · arXiv:2408.12617
-
EfficientRAG: Efficient Retriever for Multi-Hop Question Answering 8 Aug 2024 · 1 repository · arXiv:2408.04259Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Hybrid Student-Teacher Large Language Model Refinement for Cancer Toxicity Symptom Extraction 8 Aug 2024 · 0 repositories · arXiv:2408.04775
-
M2EF-NNs: Multimodal Multi-instance Evidence Fusion Neural Networks for Cancer Survival Prediction 8 Aug 2024 · 0 repositories · arXiv:2408.04170
-
Medical Graph RAG: Towards Safe Medical Large Language Model via Graph Retrieval-Augmented Generation 8 Aug 2024 · 1 repository · arXiv:2408.04187Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Multi-Turn Context Jailbreak Attack on Large Language Models From First Principles 8 Aug 2024 · 0 repositories · arXiv:2408.04686
-
Scalable Transformer for High Dimensional Multivariate Time Series Forecasting 8 Aug 2024 · 1 repository · arXiv:2408.04245
-
SCENE: Evaluating Explainable AI Techniques Using Soft Counterfactuals 8 Aug 2024 · 0 repositories · arXiv:2408.04575
-
Survey: Transformer-based Models in Data Modality Conversion 8 Aug 2024 · 0 repositories · arXiv:2408.04723
-
Towards Explainable Network Intrusion Detection using Large Language Models 8 Aug 2024 · 0 repositories · arXiv:2408.04342
-
Towards Resilient and Efficient LLMs: A Comparative Study of Efficiency, Performance, and Adversarial Robustness 8 Aug 2024 · 0 repositories · arXiv:2408.04585
-
Transformer Explainer: Interactive Learning of Text-Generative Models 8 Aug 2024 · 1 repository · arXiv:2408.04619
-
UHNet: An Ultra-Lightweight and High-Speed Edge Detection Network 8 Aug 2024 · 0 repositories · arXiv:2408.04258
-
A Comparison of LLM Finetuning Methods & Evaluation Metrics with Travel Chatbot Use Case 7 Aug 2024 · 0 repositories · arXiv:2408.03562
-
Bi-Level Spatial and Channel-aware Transformer for Learned Image Compression 7 Aug 2024 · 0 repositories · arXiv:2408.03842
-
Can Rule-Based Insights Enhance LLMs for Radiology Report Classification? Introducing the RadPrompt Methodology 7 Aug 2024 · 0 repositories · arXiv:2408.04121
-
Could ChatGPT get an Engineering Degree? Evaluating Higher Education Vulnerability to AI Assistants 7 Aug 2024 · 0 repositories · arXiv:2408.11841
-
Early Prediction of Causes (not Effects) in Healthcare by Long-Term Clinical Time Series Forecasting 7 Aug 2024 · 1 repository · arXiv:2408.03816
-
VulScribeR: Exploring RAG-based Vulnerability Augmentation with LLMs 7 Aug 2024 · 1 repository · arXiv:2408.04125
-
FMiFood: Multi-modal Contrastive Learning for Food Image Classification 7 Aug 2024 · 0 repositories · arXiv:2408.03922
-
No-Reference Image Quality Assessment with Global-Local Progressive Integration and Semantic-Aligned Quality Transfer 7 Aug 2024 · 1 repository · arXiv:2408.03885
-
Image-to-LaTeX Converter for Mathematical Formulas and Text 7 Aug 2024 · 1 repository · arXiv:2408.04015
-
Inter-Series Transformer: Attending to Products in Time Series Forecasting 7 Aug 2024 · 0 repositories · arXiv:2408.03872
-
Is Child-Directed Speech Effective Training Data for Language Models? 7 Aug 2024 · 1 repository · arXiv:2408.03617Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 1 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples) · 2 pointer-only (licence)
-
JARViS: Detecting Actions in Video Using Unified Actor-Scene Context Relation Modeling 7 Aug 2024 · 0 repositories · arXiv:2408.03612
-
Query3D: LLM-Powered Open-Vocabulary Scene Segmentation with Language Embedded 3D Gaussian 7 Aug 2024 · 1 repository · arXiv:2408.03516
-
MaxMind: A Memory Loop Network to Enhance Software Productivity based on Large Language Models 7 Aug 2024 · 0 repositories · arXiv:2408.03841
-
PackMamba: Efficient Processing of Variable-Length Sequences in Mamba training 7 Aug 2024 · 0 repositories · arXiv:2408.03865
-
PaveCap: The First Multimodal Framework for Comprehensive Pavement Condition Assessment with Dense Captioning and PCI Estimation 7 Aug 2024 · 1 repository · arXiv:2408.04110
-
RailTrack-DaViT: A Vision Transformer-Based Approach for Automated Railway Track Defect Detection 7 Aug 2024 · 1 repository
-
SocFedGPT: Federated GPT-based Adaptive Content Filtering System Leveraging User Interactions in Social Networks 7 Aug 2024 · 0 repositories · arXiv:2408.05243
-
Surgformer: Surgical Transformer with Hierarchical Temporal Attention for Surgical Phase Recognition 7 Aug 2024 · 1 repository · arXiv:2408.03867
-
SwinShadow: Shifted Window for Ambiguous Adjacent Shadow Detection 7 Aug 2024 · 1 repository · arXiv:2408.03521
-
Intermediate direct preference optimization 6 Aug 2024 · 0 repositories · arXiv:2408.02923
-
Data Poisoning in LLMs: Jailbreak-Tuning and Scaling Laws 6 Aug 2024 · 2 repositories · arXiv:2408.02946Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Empathy Level Alignment via Reinforcement Learning for Empathetic Response Generation 6 Aug 2024 · 1 repository · arXiv:2408.02976
-
Analysis of Argument Structure Constructions in a Deep Recurrent Language Model 6 Aug 2024 · 0 repositories · arXiv:2408.03062
-
Evaluating the Translation Performance of Large Language Models Based on Euas-20 6 Aug 2024 · 0 repositories · arXiv:2408.03119
-
Leveraging Entity Information for Cross-Modality Correlation Learning: The Entity-Guided Multimodal Summarization 6 Aug 2024 · 1 repository · arXiv:2408.03149
-
ReSyncer: Rewiring Style-based Generator for Unified Audio-Visually Synced Facial Performer 6 Aug 2024 · 0 repositories · arXiv:2408.03284
-
MDT-A2G: Exploring Masked Diffusion Transformers for Co-Speech Gesture Generation 6 Aug 2024 · 0 repositories · arXiv:2408.03312
-
Advancing EEG-Based Gaze Prediction Using Depthwise Separable Convolution and Enhanced Pre-Processing 6 Aug 2024 · 1 repository · arXiv:2408.03480
-
Can LLMs Serve As Time Series Anomaly Detectors? 6 Aug 2024 · 0 repositories · arXiv:2408.03475
-
FLASH: Federated Learning-Based LLMs for Advanced Query Processing in Social Networks through RAG 6 Aug 2024 · 0 repositories · arXiv:2408.05242
-
LLM-Aided Compilation for Tensor Accelerators 6 Aug 2024 · 0 repositories · arXiv:2408.03408
-
LLM-based MOFs Synthesis Condition Extraction using Few-Shot Demonstrations 6 Aug 2024 · 0 repositories · arXiv:2408.04665
-
Set2Seq Transformer: Learning Permutation Aware Set Representations of Artistic Sequences 6 Aug 2024 · 0 repositories · arXiv:2408.03404
-
TF-Locoformer: Transformer with Local Modeling by Convolution for Speech Separation and Enhancement 6 Aug 2024 · 1 repository · arXiv:2408.03440