Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 44
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 44 of 190: papers 4,301 to 4,400 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Multimodal Emotion Recognition with Vision-language Prompting and Modality Dropout 11 Sep 2024 · 0 repositories · arXiv:2409.07078
-
SimulBench: Evaluating Language Models with Creative Simulation Tasks 11 Sep 2024 · 0 repositories · arXiv:2409.07641
-
SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis 11 Sep 2024 · 1 repository · arXiv:2409.07556Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 11 harvested samples)
-
Swin-LiteMedSAM: A Lightweight Box-Based Segment Anything Model for Large-Scale Medical Image Datasets 11 Sep 2024 · 1 repository · arXiv:2409.07172Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Token Turing Machines are Efficient Vision Models 11 Sep 2024 · 1 repository · arXiv:2409.07613Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 1 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 14 harvested samples) · 1 pointer-only (licence)
-
Towards Fairer Health Recommendations: finding informative unbiased samples via Word Sense Disambiguation 11 Sep 2024 · 0 repositories · arXiv:2409.07424
-
VMAS: Video-to-Music Generation via Semantic Alignment in Web Music Videos 11 Sep 2024 · 0 repositories · arXiv:2409.07450
-
Weather-Informed Probabilistic Forecasting and Scenario Generation in Power Systems 11 Sep 2024 · 0 repositories · arXiv:2409.07637
-
Mapping Biomedical Ontology Terms to IDs: Effect of Domain Prevalence on Prediction Accuracy 11 Sep 2024 · 0 repositories · arXiv:2409.13746
-
KAG: Boosting LLMs in Professional Domains via Knowledge Augmented Generation 10 Sep 2024 · 1 repository · arXiv:2409.13731Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
A Dataset for Evaluating LLM-based Evaluation Functions for Research Question Extraction Task 10 Sep 2024 · 0 repositories · arXiv:2409.06883
-
A Practical Gated Recurrent Transformer Network Incorporating Multiple Fusions for Video Denoising 10 Sep 2024 · 0 repositories · arXiv:2409.06603
-
Accelerating Large Language Model Pretraining via LFR Pedagogy: Learn, Focus, and Review 10 Sep 2024 · 0 repositories · arXiv:2409.06131
-
Adaptive Transformer Modelling of Density Function for Nonparametric Survival Analysis 10 Sep 2024 · 1 repository · arXiv:2409.06209Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 0 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
AgileIR: Memory-Efficient Group Shifted Windows Attention for Agile Image Restoration 10 Sep 2024 · 0 repositories · arXiv:2409.06206
-
Can Large Language Models Unlock Novel Scientific Research Ideas? 10 Sep 2024 · 1 repository · arXiv:2409.06185Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 9 unverified (of 11 harvested samples)
-
DA-MoE: Towards Dynamic Expert Allocation for Mixture-of-Experts Models 10 Sep 2024 · 0 repositories · arXiv:2409.06669
-
Generative AI for Requirements Engineering: A Systematic Literature Review 10 Sep 2024 · 0 repositories · arXiv:2409.06741
-
GroUSE: A Benchmark to Evaluate Evaluators in Grounded Question Answering 10 Sep 2024 · 1 repository · arXiv:2409.06595
-
Knowledge Distillation via Query Selection for Detection Transformer 10 Sep 2024 · 0 repositories · arXiv:2409.06443
-
Lightweight single-image super-resolution network based on dual paths 10 Sep 2024 · 0 repositories · arXiv:2409.06590
-
Static for Dynamic: Towards a Deeper Understanding of Dynamic Facial Expressions Using Static Expression Data 10 Sep 2024 · 1 repository · arXiv:2409.06154
-
What is the Role of Small Models in the LLM Era: A Survey 10 Sep 2024 · 1 repository · arXiv:2409.06857
-
Classification performance and reproducibility of GPT-4 omni for information extraction from veterinary electronic health records 9 Sep 2024 · 1 repository · arXiv:2409.13727
-
Rule Extrapolation in Language Models: A Study of Compositional Generalization on OOD Prompts 9 Sep 2024 · 1 repository · arXiv:2409.13728
-
AbGPT: De Novo Antibody Design via Generative Language Modeling 9 Sep 2024 · 1 repository · arXiv:2409.06090
-
Assessing SPARQL capabilities of Large Language Models 9 Sep 2024 · 2 repositories · arXiv:2409.05925
-
Deep Generative Model for Mechanical System Configuration Design 9 Sep 2024 · 0 repositories · arXiv:2409.06016
-
DriveScape: Towards High-Resolution Controllable Multi-View Driving Video Generation 9 Sep 2024 · 0 repositories · arXiv:2409.05463
-
DSDFormer: An Innovative Transformer-Mamba Framework for Robust High-Precision Driver Distraction Identification 9 Sep 2024 · 0 repositories · arXiv:2409.05587
-
Elsevier Arena: Human Evaluation of Chemistry/Biology/Health Foundational Large Language Models 9 Sep 2024 · 0 repositories · arXiv:2409.05486
-
Exploring Rich Subjective Quality Information for Image Quality Assessment in the Wild 9 Sep 2024 · 0 repositories · arXiv:2409.05540
-
FairHome: A Fair Housing and Fair Lending Dataset 9 Sep 2024 · 0 repositories · arXiv:2409.05990
-
Harmonic Reasoning in Large Language Models 9 Sep 2024 · 0 repositories · arXiv:2409.05521
-
Identifying the sources of ideological bias in GPT models through linguistic variation in output 9 Sep 2024 · 0 repositories · arXiv:2409.06043
-
MemoRAG: Moving towards Next-Gen RAG Via Memory-Inspired Knowledge Discovery 9 Sep 2024 · 1 repository · arXiv:2409.05591Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples) · 6 pointer-only (licence)
-
Regression with Large Language Models for Materials and Molecular Property Prediction 9 Sep 2024 · 0 repositories · arXiv:2409.06080
-
Retrofitting Temporal Graph Neural Networks with Transformer 9 Sep 2024 · 1 repository · arXiv:2409.05477
-
Revisiting the Solution of Meta KDD Cup 2024: CRAG 9 Sep 2024 · 1 repository · arXiv:2409.15337
-
RotCAtt-TransUNet++: Novel Deep Neural Network for Sophisticated Cardiac Segmentation 9 Sep 2024 · 1 repository · arXiv:2409.05280
-
Towards Building a Robust Knowledge Intensive Question Answering Model with Large Language Models 9 Sep 2024 · 0 repositories · arXiv:2409.05385
-
Zero-shot Outlier Detection via Prior-data Fitted Networks: Model Selection Bygone! 9 Sep 2024 · 0 repositories · arXiv:2409.05672
-
Audio-Guided Fusion Techniques for Multimodal Emotion Analysis 8 Sep 2024 · 0 repositories · arXiv:2409.05007
-
Lung-DETR: Deformable Detection Transformer for Sparse Lung Nodule Anomaly Detection 8 Sep 2024 · 0 repositories · arXiv:2409.05200
-
OneGen: Efficient One-Pass Unified Generation and Retrieval for LLMs 8 Sep 2024 · 1 repository · arXiv:2409.05152
-
Vision-fused Attack: Advancing Aggressive and Stealthy Adversarial Text against Neural Machine Translation 8 Sep 2024 · 1 repository · arXiv:2409.05021Syntology official: harvested, nothing ran · 0 ran · 4 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Activation Function Optimization Scheme for Image Classification 7 Sep 2024 · 1 repository · arXiv:2409.04915
-
Towards Weather-Robust 3D Human Body Reconstruction: Millimeter-Wave Radar-Based Dataset, Benchmark, and Multi-Modal Fusion 7 Sep 2024 · 0 repositories · arXiv:2409.04851
-
Cross-attention Inspired Selective State Space Models for Target Sound Extraction 7 Sep 2024 · 1 repository · arXiv:2409.04803Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Efficient Training of Transformers for Molecule Property Prediction on Small-scale Datasets 7 Sep 2024 · 0 repositories · arXiv:2409.04909
-
MuAP: Multi-step Adaptive Prompt Learning for Vision-Language Model with Missing Modality 7 Sep 2024 · 0 repositories · arXiv:2409.04693
-
NapTune: Efficient Model Tuning for Mood Classification using Previous Night's Sleep Measures along with Wearable Time-series 7 Sep 2024 · 0 repositories · arXiv:2409.04723
-
Swin Transformer for Robust Differentiation of Real and Synthetic Images: Intra- and Inter-Dataset Analysis 7 Sep 2024 · 0 repositories · arXiv:2409.04734
-
VidLPRO: A Video-Language Pre-training Framework for Robotic and Laparoscopic Surgery 7 Sep 2024 · 0 repositories · arXiv:2409.04732
-
ActionFlow: Equivariant, Accurate, and Efficient Policies with Spatially Symmetric Flow Matching 6 Sep 2024 · 0 repositories · arXiv:2409.04576
-
Advancing SEM Based Nano-Scale Defect Analysis in Semiconductor Manufacturing for Advanced IC Nodes 6 Sep 2024 · 0 repositories · arXiv:2409.04310
-
AnyMatch -- Efficient Zero-Shot Entity Matching with a Small Language Model 6 Sep 2024 · 1 repository · arXiv:2409.04073
-
BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer Training 6 Sep 2024 · 1 repository · arXiv:2409.04599
-
Column Vocabulary Association (CVA): semantic interpretation of dataless tables 6 Sep 2024 · 0 repositories · arXiv:2409.13709
-
Combining LLMs and Knowledge Graphs to Reduce Hallucinations in Question Answering 6 Sep 2024 · 0 repositories · arXiv:2409.04181
-
GALLa: Graph Aligned Large Language Models for Improved Source Code Understanding 6 Sep 2024 · 0 repositories · arXiv:2409.04183
-
Qihoo-T2X: An Efficient Proxy-Tokenized Diffusion Transformer for Text-to-Any-Task 6 Sep 2024 · 1 repository · arXiv:2409.04005
-
Retrieval Augmented Generation-Based Incident Resolution Recommendation System for IT Support 6 Sep 2024 · 0 repositories · arXiv:2409.13707
-
Towards Safer Online Spaces: Simulating and Assessing Intervention Strategies for Eating Disorder Discussions 6 Sep 2024 · 0 repositories · arXiv:2409.04043
-
UI-JEPA: Towards Active Perception of User Intent through Onscreen User Activity 6 Sep 2024 · 0 repositories · arXiv:2409.04081
-
Bypassing DARCY Defense: Indistinguishable Universal Adversarial Triggers 5 Sep 2024 · 0 repositories · arXiv:2409.03183
-
CACER: Clinical Concept Annotations for Cancer Events and Relations 5 Sep 2024 · 1 repository · arXiv:2409.03905
-
Causal Temporal Representation Learning with Nonstationary Sparse Transition 5 Sep 2024 · 1 repository · arXiv:2409.03142Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Characterizing Massive Activations of Attention Mechanism in Graph Neural Networks 5 Sep 2024 · 1 repository · arXiv:2409.03463
-
Evaluating Open-Source Sparse Autoencoders on Disentangling Factual Knowledge in GPT-2 Small 5 Sep 2024 · 1 repository · arXiv:2409.04478Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
LMLT: Low-to-high Multi-Level Vision Transformer for Image Super-Resolution 5 Sep 2024 · 1 repository · arXiv:2409.03516
-
MARAGS: A Multi-Adapter System for Multi-Task Retrieval Augmented Generation Question Answering 5 Sep 2024 · 0 repositories · arXiv:2409.03171
-
MaterialBENCH: Evaluating College-Level Materials Science Problem-Solving Abilities of Large Language Models 5 Sep 2024 · 0 repositories · arXiv:2409.03161
-
MVTN: A Multiscale Video Transformer Network for Hand Gesture Recognition 5 Sep 2024 · 1 repository · arXiv:2409.03890
-
Onboard Satellite Image Classification for Earth Observation: A Comparative Study of ViT Models 5 Sep 2024 · 1 repository · arXiv:2409.03901
-
RAG based Question-Answering for Contextual Response Prediction System 5 Sep 2024 · 0 repositories · arXiv:2409.03708
-
Revolutionizing Database Q&A with Large Language Models: Comprehensive Benchmark and Evaluation 5 Sep 2024 · 1 repository · arXiv:2409.04475
-
Sketch: A Toolkit for Streamlining LLM Operations 5 Sep 2024 · 0 repositories · arXiv:2409.03346
-
Why mamba is effective? Exploit Linear Transformer-Mamba Network for Multi-Modality Image Fusion 5 Sep 2024 · 0 repositories · arXiv:2409.03223
-
xLAM: A Family of Large Action Models to Empower AI Agent Systems 5 Sep 2024 · 1 repository · arXiv:2409.03215
-
A Comparative Study on Large Language Models for Log Parsing 4 Sep 2024 · 0 repositories · arXiv:2409.02474
-
GenDFIR: Advancing Cyber Incident Timeline Analysis Through Retrieval Augmented Generation and Large Language Models 4 Sep 2024 · 0 repositories · arXiv:2409.02572
-
Causality-Aware Transformer Networks for Robotic Navigation 4 Sep 2024 · 0 repositories · arXiv:2409.02669
-
Detecting Calls to Action in Multimodal Content: Analysis of the 2021 German Federal Election Campaign on Instagram 4 Sep 2024 · 0 repositories · arXiv:2409.02690
-
Diversify-verify-adapt: Efficient and Robust Retrieval-Augmented Ambiguous Question Answering 4 Sep 2024 · 0 repositories · arXiv:2409.02361
-
Historical German Text Normalization Using Type- and Token-Based Language Modeling 4 Sep 2024 · 0 repositories · arXiv:2409.02841
-
How DREAMS are made: Emulating Satellite Galaxy and Subhalo Populations with Diffusion Models and Point Clouds 4 Sep 2024 · 1 repository · arXiv:2409.02980Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 11 harvested samples)
-
How Privacy-Savvy Are Large Language Models? A Case Study on Compliance and Privacy Technical Review 4 Sep 2024 · 0 repositories · arXiv:2409.02375
-
Hypothesizing Missing Causal Variables with LLMs 4 Sep 2024 · 1 repository · arXiv:2409.02604
-
iConFormer: Dynamic Parameter-Efficient Tuning with Input-Conditioned Adaptation 4 Sep 2024 · 0 repositories · arXiv:2409.02838
-
Incorporating Like-Minded Peers to Overcome Friend Data Sparsity in Session-Based Social Recommendations 4 Sep 2024 · 0 repositories · arXiv:2409.02702
-
Irrelevant Alternatives Bias Large Language Model Hiring Decisions 4 Sep 2024 · 0 repositories · arXiv:2409.15299
-
Large Language Models as Efficient Reward Function Searchers for Custom-Environment Multi-Objective Reinforcement Learning 4 Sep 2024 · 0 repositories · arXiv:2409.02428
-
Leveraging Interpretability in the Transformer to Automate the Proactive Scaling of Cloud Resources 4 Sep 2024 · 0 repositories · arXiv:2409.03103
-
LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via a Hybrid Architecture 4 Sep 2024 · 1 repository · arXiv:2409.02889Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
MoA is All You Need: Building LLM Research Team using Mixture of Agents 4 Sep 2024 · 0 repositories · arXiv:2409.07487
-
MobileUNETR: A Lightweight End-To-End Hybrid Vision Transformer For Efficient Medical Image Segmentation 4 Sep 2024 · 1 repository · arXiv:2409.03062
-
More is More: Addition Bias in Large Language Models 4 Sep 2024 · 1 repository · arXiv:2409.02569
-
Robust Text-to-Cypher Using Combination of BERT, GraphSAGE, and Transformer (CoBGT) Model 4 Sep 2024 · 0 repositories
-
Towards Data-Centric Face Anti-Spoofing: Improving Cross-domain Generalization via Physics-based Data Synthesis 4 Sep 2024 · 0 repositories · arXiv:2409.03501