Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 49
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 49 of 190: papers 4,801 to 4,900 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
The Use of Large Language Models (LLM) for Cyber Threat Intelligence (CTI) in Cybercrime Forums 6 Aug 2024 · 0 repositories · arXiv:2408.03354
-
TrafficGPT: An LLM Approach for Open-Set Encrypted Traffic Classification 6 Aug 2024 · 1 repository
-
AssemAI: Interpretable Image-Based Anomaly Detection for Manufacturing Pipelines 5 Aug 2024 · 1 repository · arXiv:2408.02181
-
Is Large Language Model Good at Database Knob Tuning? A Comprehensive Experimental Evaluation 5 Aug 2024 · 0 repositories · arXiv:2408.02213
-
Cross-modulated Attention Transformer for RGBT Tracking 5 Aug 2024 · 0 repositories · arXiv:2408.02222
-
Do Large Language Models Speak All Languages Equally? A Comparative Study in Low-Resource Settings 5 Aug 2024 · 0 repositories · arXiv:2408.02237
-
DRFormer: Multi-Scale Transformer Utilizing Diverse Receptive Fields for Long Time-Series Forecasting 5 Aug 2024 · 1 repository · arXiv:2408.02279
-
The NPU-ASLP System Description for Visual Speech Recognition in CNVSRC 2024 5 Aug 2024 · 1 repository · arXiv:2408.02369
-
Why Are My Prompts Leaked? Unraveling Prompt Extraction Threats in Customized Large Language Models 5 Aug 2024 · 1 repository · arXiv:2408.02416
-
RAG Foundry: A Framework for Enhancing LLMs for Retrieval Augmented Generation 5 Aug 2024 · 2 repositories · arXiv:2408.02545
-
On Using Quasirandom Sequences in Machine Learning for Model Weight Initialization 5 Aug 2024 · 1 repository · arXiv:2408.02654
-
Wiping out the limitations of Large Language Models -- A Taxonomy for Retrieval Augmented Generation 5 Aug 2024 · 0 repositories · arXiv:2408.02854
-
AppAgent v2: Advanced Agent for Flexible Mobile Interactions 5 Aug 2024 · 0 repositories · arXiv:2408.11824
-
LLM Agents Improve Semantic Code Search 5 Aug 2024 · 0 repositories · arXiv:2408.11058
-
SEAS: Self-Evolving Adversarial Safety Optimization for Large Language Models 5 Aug 2024 · 1 repository · arXiv:2408.02632
-
Self-Taught Evaluators 5 Aug 2024 · 0 repositories · arXiv:2408.02666
-
Toward Attention-based TinyML: A Heterogeneous Accelerated Architecture and Automated Deployment Flow 5 Aug 2024 · 1 repository · arXiv:2408.02473
-
XMainframe: A Large Language Model for Mainframe Modernization 5 Aug 2024 · 1 repository · arXiv:2408.04660
-
DiReCT: Diagnostic Reasoning for Clinical Notes via Large Language Models 4 Aug 2024 · 1 repository · arXiv:2408.01933Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 1 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
ML-EAT: A Multilevel Embedding Association Test for Interpretable and Transparent Social Science 4 Aug 2024 · 1 repository · arXiv:2408.01966
-
AdaCBM: An Adaptive Concept Bottleneck Model for Explainable and Accurate Diagnosis 4 Aug 2024 · 1 repository · arXiv:2408.02001
-
MedSyn: LLM-based Synthetic Medical Text Generation Framework 4 Aug 2024 · 1 repository · arXiv:2408.02056
-
KAN-RCBEVDepth: A multi-modal fusion algorithm in object detection for autonomous driving 4 Aug 2024 · 0 repositories · arXiv:2408.02088
-
Effective Demonstration Annotation for In-Context Learning via Language Model-Based Determinantal Point Process 4 Aug 2024 · 0 repositories · arXiv:2408.02103
-
Leveraging Large Language Models with Chain-of-Thought and Prompt Engineering for Traffic Crash Severity Analysis and Inference 4 Aug 2024 · 0 repositories · arXiv:2408.04652
-
Advancing Mental Health Pre-Screening: A New Custom GPT for Psychological Distress Assessment 3 Aug 2024 · 0 repositories · arXiv:2408.01614
-
Self-Emotion Blended Dialogue Generation in Social Simulation Agents 3 Aug 2024 · 0 repositories · arXiv:2408.01633
-
Stimulating Imagination: Towards General-purpose Object Rearrangement 3 Aug 2024 · 0 repositories · arXiv:2408.01655
-
A Novel Evaluation Framework for Image2Text Generation 3 Aug 2024 · 0 repositories · arXiv:2408.01723
-
LAM3D: Leveraging Attention for Monocular 3D Object Detection 3 Aug 2024 · 0 repositories · arXiv:2408.01739
-
GLDiTalker: Speech-Driven 3D Facial Animation with Graph Latent Diffusion Transformer 3 Aug 2024 · 0 repositories · arXiv:2408.01826
-
Efficient Solutions For An Intriguing Failure of LLMs: Long Context Window Does Not Mean LLMs Can Analyze Long Sequences Flawlessly 3 Aug 2024 · 0 repositories · arXiv:2408.01866
-
MALADE: Orchestration of LLM-powered Agents with Retrieval Augmented Generation for Pharmacovigilance 3 Aug 2024 · 1 repository · arXiv:2408.01869
-
Building Trust in Mental Health Chatbots: Safety Metrics and LLM-Based Evaluation Tools 3 Aug 2024 · 0 repositories · arXiv:2408.04650
-
Distinguishing Chatbot from Human 3 Aug 2024 · 0 repositories · arXiv:2408.04647
-
JambaTalk: Speech-Driven 3D Talking Head Generation Based on Hybrid Transformer-Mamba Language Model 3 Aug 2024 · 0 repositories · arXiv:2408.01627
-
POA: Pre-training Once for Models of All Sizes 2 Aug 2024 · 1 repository · arXiv:2408.01031
-
MambaST: A Plug-and-Play Cross-Spectral Spatial-Temporal Fuser for Efficient Pedestrian Detection 2 Aug 2024 · 1 repository · arXiv:2408.01037
-
LLM as Runtime Error Handler: A Promising Pathway to Adaptive Self-Healing of Software Systems 2 Aug 2024 · 0 repositories · arXiv:2408.01055
-
Leveraging Encoder-only Large Language Models for Mobile App Review Feature Extraction 2 Aug 2024 · 1 repository · arXiv:2408.01063
-
BioRAG: A RAG-LLM Framework for Biological Question Reasoning 2 Aug 2024 · 0 repositories · arXiv:2408.01107
-
An Efficient and Effective Transformer Decoder-Based Framework for Multi-Task Visual Grounding 2 Aug 2024 · 1 repository · arXiv:2408.01120Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 15 harvested samples) · 15 pointer-only (licence)
-
A Survey of Mamba 2 Aug 2024 · 0 repositories · arXiv:2408.01129
-
Rethinking Pre-Trained Feature Extractor Selection in Multiple Instance Learning for Whole Slide Image Classification 2 Aug 2024 · 1 repository · arXiv:2408.01167
-
Nested Music Transformer: Sequentially Decoding Compound Tokens in Symbolic Music and Audio Generation 2 Aug 2024 · 1 repository · arXiv:2408.01180
-
High-Throughput Phenotyping of Clinical Text Using Large Language Models 2 Aug 2024 · 0 repositories · arXiv:2408.01214
-
Multi-head Spatial-Spectral Mamba for Hyperspectral Image Classification 2 Aug 2024 · 1 repository · arXiv:2408.01224
-
HeteroMorpheus: Universal Control Based on Morphological Heterogeneity Modeling 2 Aug 2024 · 1 repository · arXiv:2408.01230
-
WaveMamba: Spatial-Spectral Wavelet Mamba for Hyperspectral Image Classification 2 Aug 2024 · 0 repositories · arXiv:2408.01231
-
RAGEval: Scenario Specific RAG Evaluation Dataset Generation Framework 2 Aug 2024 · 1 repository · arXiv:2408.01262
-
Underwater Object Detection Enhancement via Channel Stabilization 2 Aug 2024 · 1 repository · arXiv:2408.01293
-
Spatial and Spatial-Spectral Morphological Mamba for Hyperspectral Image Classification 2 Aug 2024 · 2 repositories · arXiv:2408.01372
-
NOLO: Navigate Only Look Once 2 Aug 2024 · 0 repositories · arXiv:2408.01384
-
Pre-trained Language Models Improve the Few-shot Prompt Ability of Decision Transformer 2 Aug 2024 · 0 repositories · arXiv:2408.01402
-
THOR2: Topological Analysis for 3D Shape and Color-Based Human-Inspired Object Recognition in Unseen Environments 2 Aug 2024 · 1 repository · arXiv:2408.01579
-
Evaluating the Impact of Advanced LLM Techniques on AI-Lecture Tutors for a Robotics Course 2 Aug 2024 · 0 repositories · arXiv:2408.04645
-
Enhanced Structured State Space Models via Grouped FIR Filtering and Attention Sink Mechanisms 1 Aug 2024 · 0 repositories · arXiv:2408.00244
-
OTAD: An Optimal Transport-Induced Robust Model for Agnostic Adversarial Attack 1 Aug 2024 · 0 repositories · arXiv:2408.00329
-
Improving Retrieval-Augmented Generation in Medicine with Iterative Follow-up Questions 1 Aug 2024 · 1 repository · arXiv:2408.00727
-
Leaf Angle Estimation using Mask R-CNN and LETR Vision Transformer 1 Aug 2024 · 0 repositories · arXiv:2408.00749
-
AgentGen: Enhancing Planning Abilities for Large Language Model based Agent via Environment and Task Generation 1 Aug 2024 · 1 repository · arXiv:2408.00764Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 9 harvested samples)
-
Granting GPT-4 License and Opportunity: Enhancing Accuracy and Confidence Estimation for Few-Shot Event Detection 1 Aug 2024 · 0 repositories · arXiv:2408.00914
-
Automatic Pull Request Description Generation Using LLMs: A T5 Model Approach 1 Aug 2024 · 0 repositories · arXiv:2408.00921
-
Advancing Medical Image Segmentation: Morphology-Driven Learning with Diffusion Transformer 1 Aug 2024 · 1 repository · arXiv:2408.00347
-
DNTextSpotter: Arbitrary-Shaped Scene Text Spotting via Improved Denoising Training 1 Aug 2024 · 1 repository · arXiv:2408.00355
-
Cross-Scan Mamba with Masked Training for Robust Spectral Imaging 1 Aug 2024 · 0 repositories · arXiv:2408.00629
-
Hybrid Querying Over Relational Databases and Large Language Models 1 Aug 2024 · 0 repositories · arXiv:2408.00884
-
What comes after transformers? -- A selective survey connecting ideas in deep learning 1 Aug 2024 · 0 repositories · arXiv:2408.00386
-
Multi-Level Querying using A Knowledge Pyramid 31 Jul 2024 · 0 repositories · arXiv:2407.21276
-
SAKR: Enhancing Retrieval-Augmented Generation via Streaming Algorithm and K-Means Clustering 31 Jul 2024 · 0 repositories · arXiv:2407.21300
-
MetaOpenFOAM: an LLM-based multi-agent framework for CFD 31 Jul 2024 · 1 repository · arXiv:2407.21320Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Performance of Recent Large Language Models for a Low-Resourced Language 31 Jul 2024 · 0 repositories · arXiv:2407.21330
-
Improving Faithfulness of Large Language Models in Summarization via Sliding Generation and Self-Consistency 31 Jul 2024 · 0 repositories · arXiv:2407.21443
-
Generative Expressive Conversational Speech Synthesis 31 Jul 2024 · 1 repository · arXiv:2407.21491
-
FSSC: Federated Learning of Transformer Neural Networks for Semantic Image Communication 31 Jul 2024 · 0 repositories · arXiv:2407.21507
-
Interpreting and learning voice commands with a Large Language Model for a robot system 31 Jul 2024 · 0 repositories · arXiv:2407.21512
-
Can LLMs "Reason" in Music? An Evaluation of LLMs' Capability of Music Understanding and Generation 31 Jul 2024 · 0 repositories · arXiv:2407.21531
-
PMoE: Progressive Mixture of Experts with Asymmetric Transformer for Continual Learning 31 Jul 2024 · 0 repositories · arXiv:2407.21571
-
Dynamic Object Queries for Transformer-based Incremental Object Detection 31 Jul 2024 · 0 repositories · arXiv:2407.21687
-
Adaptive Retrieval-Augmented Generation for Conversational Systems 31 Jul 2024 · 0 repositories · arXiv:2407.21712
-
Gemma 2: Improving Open Language Models at a Practical Size 31 Jul 2024 · 0 repositories · arXiv:2408.00118
-
Automated Software Vulnerability Static Code Analysis Using Generative Pre-Trained Transformer Models 31 Jul 2024 · 0 repositories · arXiv:2408.00197
-
Recording First-person Experiences to Build a New Type of Foundation Model 31 Jul 2024 · 0 repositories · arXiv:2408.02680
-
A New Type of Foundation Model Based on Recordings of People's Emotions and Physiology 31 Jul 2024 · 0 repositories · arXiv:2408.00030
-
A Simple Low-bit Quantization Framework for Video Snapshot Compressive Imaging 31 Jul 2024 · 1 repository · arXiv:2407.21517
-
CC-SAM: SAM with Cross-feature Attention and Context for Ultrasound Image Segmentation 31 Jul 2024 · 0 repositories · arXiv:2408.00181
-
MART: MultiscAle Relational Transformer Networks for Multi-agent Trajectory Prediction 31 Jul 2024 · 1 repository · arXiv:2407.21635
-
On-the-fly Point Feature Representation for Point Clouds Analysis 31 Jul 2024 · 0 repositories · arXiv:2407.21335
-
RoadFormer+: Delivering RGB-X Scene Parsing through Scale-Aware Information Decoupling and Advanced Heterogeneous Feature Fusion 31 Jul 2024 · 0 repositories · arXiv:2407.21631
-
Semantic Successive Refinement: A Generative AI-aided Semantic Communication Framework 31 Jul 2024 · 0 repositories · arXiv:2408.05112
-
The Llama 3 Herd of Models 31 Jul 2024 · 5 repositories · arXiv:2407.21783Syntology 9 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 1 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples)
-
Tora: Trajectory-oriented Diffusion Transformer for Video Generation 31 Jul 2024 · 1 repository · arXiv:2407.21705Syntology official (archive's flag): 9 ran · 9 ran (of which 4 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 1 violated, 6 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 12 harvested samples)
-
Zero Shot Health Trajectory Prediction Using Transformer 30 Jul 2024 · 1 repository · arXiv:2407.21124Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples)
-
Decomposed Prompting to Answer Questions on a Course Discussion Board 30 Jul 2024 · 1 repository · arXiv:2407.21170
-
GenRec: Generative Sequential Recommendation with Large Language Models 30 Jul 2024 · 1 repository · arXiv:2407.21191
-
Be aware of overfitting by hyperparameter optimization! 30 Jul 2024 · 0 repositories · arXiv:2407.20786
-
BERT and LLMs-Based avGFP Brightness Prediction and Mutation Design 30 Jul 2024 · 0 repositories · arXiv:2407.20534
-
Breaking Agents: Compromising Autonomous LLM Agents Through Malfunction Amplification 30 Jul 2024 · 0 repositories · arXiv:2407.20859
-
Comparison of Large Language Models for Generating Contextually Relevant Questions 30 Jul 2024 · 1 repository · arXiv:2407.20578
-
Effectively Leveraging CLIP for Generating Situational Summaries of Images and Videos 30 Jul 2024 · 1 repository · arXiv:2407.20642