Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 84
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 84 of 190: papers 8,301 to 8,400 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
AttentionLego: An Open-Source Building Block For Spatially-Scalable Large Language Model Accelerator With Processing-In-Memory Technology 21 Jan 2024 · 0 repositories · arXiv:2401.11459
-
CheX-GPT: Harnessing Large Language Models for Enhanced Chest X-ray Report Labeling 21 Jan 2024 · 2 repositories · arXiv:2401.11505
-
Enhancing Recommendation Diversity by Re-ranking with Large Language Models 21 Jan 2024 · 0 repositories · arXiv:2401.11506
-
Epilepsy Seizure Detection and Prediction using an Approximate Spiking Convolutional Transformer 21 Jan 2024 · 0 repositories · arXiv:2402.09424
-
Finding a Needle in the Adversarial Haystack: A Targeted Paraphrasing Approach For Uncovering Edge Cases with Minimal Distribution Distortion 21 Jan 2024 · 1 repository · arXiv:2401.11373
-
Freely Long-Thinking Transformer (FraiLT) 21 Jan 2024 · 0 repositories · arXiv:2401.11626
-
Language Models as Hierarchy Encoders 21 Jan 2024 · 1 repository · arXiv:2401.11374Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples)
-
LLMRA: Multi-modal Large Language Model based Restoration Assistant 21 Jan 2024 · 0 repositories · arXiv:2401.11401
-
ProLex: A Benchmark for Language Proficiency-oriented Lexical Substitution 21 Jan 2024 · 1 repository · arXiv:2401.11356
-
Scalable High-Resolution Pixel-Space Image Synthesis with Hourglass Diffusion Transformers 21 Jan 2024 · 1 repository · arXiv:2401.11605
-
Training microrobots to swim by a large language model 21 Jan 2024 · 0 repositories · arXiv:2402.00044
-
BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models 20 Jan 2024 · 1 repository · arXiv:2401.12242Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 15 harvested samples)
-
DengueNet: Dengue Prediction using Spatiotemporal Satellite Imagery for Resource-Limited Countries 20 Jan 2024 · 1 repository · arXiv:2401.11114
-
Drop your Decoder: Pre-training with Bag-of-Word Prediction for Dense Passage Retrieval 20 Jan 2024 · 3 repositories · arXiv:2401.11248Syntology official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Enhancing Large Language Models for Clinical Decision Support by Incorporating Clinical Practice Guidelines 20 Jan 2024 · 0 repositories · arXiv:2401.11120
-
Evaluating and Enhancing Large Language Models Performance in Domain-specific Medicine: Osteoarthritis Management with DocOA 20 Jan 2024 · 0 repositories · arXiv:2401.12998
-
Density Adaptive Attention is All You Need: Robust Parameter-Efficient Fine-Tuning Across Multiple Modalities 20 Jan 2024 · 2 repositories · arXiv:2401.11143Syntology community repositories only · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images 20 Jan 2024 · 1 repository · arXiv:2401.11170Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
LMUFormer: Low Complexity Yet Powerful Spiking Model With Legendre Memory Units 20 Jan 2024 · 1 repository · arXiv:2402.04882Syntology official (archive's flag): 9 ran · 9 ran (of which 1 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 12 harvested samples)
-
Prompt-RAG: Pioneering Vector Embedding-Free Retrieval-Augmented Generation in Niche Domains, Exemplified by Korean Medicine 20 Jan 2024 · 0 repositories · arXiv:2401.11246
-
Uncertainty-aware Bridge based Mobile-Former Network for Event-based Pattern Recognition 20 Jan 2024 · 1 repository · arXiv:2401.11123
-
Unfair TOS: An Automated Approach using Customized BERT 20 Jan 2024 · 0 repositories · arXiv:2401.11207
-
AAT: Adapting Audio Transformer for Various Acoustics Recognition Tasks 19 Jan 2024 · 1 repository · arXiv:2401.10544
-
Attentive Fusion: A Transformer-based Approach to Multimodal Hate Speech Detection 19 Jan 2024 · 2 repositories · arXiv:2401.10653
-
DeepRLI: A Multi-objective Framework for Universal Protein--Ligand Interaction Prediction 19 Jan 2024 · 1 repository · arXiv:2401.10806
-
FinLLMs: A Framework for Financial Reasoning Dataset Generation with Large Language Models 19 Jan 2024 · 0 repositories · arXiv:2401.10744
-
LangBridge: Multilingual Reasoning Without Multilingual Supervision 19 Jan 2024 · 1 repository · arXiv:2401.10695Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 7 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
M2ORT: Many-To-One Regression Transformer for Spatial Transcriptomics Prediction from Histopathology Images 19 Jan 2024 · 1 repository · arXiv:2401.10608Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
MDGNN: Multi-Relational Dynamic Graph Neural Network for Comprehensive and Dynamic Stock Investment Prediction 19 Jan 2024 · 0 repositories · arXiv:2402.06633
-
Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences 19 Jan 2024 · 1 repository · arXiv:2401.10529
-
Mining experimental data from Materials Science literature with Large Language Models: an evaluation study 19 Jan 2024 · 1 repository · arXiv:2401.11052
-
Reinforcement learning for question answering in programming domain using public community scoring as a human feedback 19 Jan 2024 · 0 repositories · arXiv:2401.10882
-
Speech Swin-Transformer: Exploring a Hierarchical Transformer with Shifted Windows for Speech Emotion Recognition 19 Jan 2024 · 0 repositories · arXiv:2401.10536
-
Understanding Video Transformers via Universal Concept Discovery 19 Jan 2024 · 0 repositories · arXiv:2401.10831
-
An Empirical Study on the Impact of Positional Encoding in Transformer-based Monaural Speech Enhancement 18 Jan 2024 · 0 repositories · arXiv:2401.09686
-
Beyond Traditional Benchmarks: Analyzing Behaviors of Open LLMs on Data-to-Text Generation 18 Jan 2024 · 0 repositories · arXiv:2401.10186
-
BlenDA: Domain Adaptive Object Detection through diffusion-based blending 18 Jan 2024 · 1 repository · arXiv:2401.09921
-
ChatQA: Surpassing GPT-4 on Conversational QA and RAG 18 Jan 2024 · 0 repositories · arXiv:2401.10225
-
Code Prompting Elicits Conditional Reasoning Abilities in Text+Code LLMs 18 Jan 2024 · 1 repository · arXiv:2401.10065Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
Controllable Decontextualization of Yes/No Question and Answers into Factual Statements 18 Jan 2024 · 0 repositories · arXiv:2401.09775
-
Curriculum Recommendations Using Transformer Base Model with InfoNCE Loss And Language Switching Method 18 Jan 2024 · 0 repositories · arXiv:2401.09699
-
Exploring General Intelligence via Gated Graph Transformer in Functional Connectivity Studies 18 Jan 2024 · 0 repositories · arXiv:2401.10348
-
Gender Bias in Machine Translation and The Era of Large Language Models 18 Jan 2024 · 0 repositories · arXiv:2401.10016
-
Image Translation as Diffusion Visual Programmers 18 Jan 2024 · 0 repositories · arXiv:2401.09742
-
Leveraging Biases in Large Language Models: "bias-kNN'' for Effective Few-Shot Learning 18 Jan 2024 · 0 repositories · arXiv:2401.09783
-
R-Judge: Benchmarking Safety Risk Awareness for LLM Agents 18 Jan 2024 · 1 repository · arXiv:2401.10019Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Reconstructing the Invisible: Video Frame Restoration through Siamese Masked Conditional Variational Autoencoder 18 Jan 2024 · 0 repositories · arXiv:2401.10402
-
Self-Rewarding Language Models 18 Jan 2024 · 3 repositories · arXiv:2401.10020Syntology 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 2 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (of 11 harvested samples)
-
Towards Principled Graph Transformers 18 Jan 2024 · 1 repository · arXiv:2401.10119Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
When Neural Code Completion Models Size up the Situation: Attaining Cheaper and Faster Completion through Dynamic Model Inference 18 Jan 2024 · 1 repository · arXiv:2401.09964
-
AttackEval: How to Evaluate the Effectiveness of Jailbreak Attacking on Large Language Models 17 Jan 2024 · 0 repositories · arXiv:2401.09002
-
Augmenting Math Word Problems via Iterative Question Composing 17 Jan 2024 · 1 repository · arXiv:2401.09003Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
BENO: Boundary-embedded Neural Operators for Elliptic PDEs 17 Jan 2024 · 1 repository · arXiv:2401.09323Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 9 harvested samples)
-
Bridging Research and Readers: A Multi-Modal Automated Academic Papers Interpretation System 17 Jan 2024 · 1 repository · arXiv:2401.09150
-
COCO is "ALL'' You Need for Visual Instruction Fine-tuning 17 Jan 2024 · 0 repositories · arXiv:2401.08968
-
Deciphering Textual Authenticity: A Generalized Strategy through the Lens of Large Language Semantics for Detecting Human vs. Machine-Generated Text 17 Jan 2024 · 1 repository · arXiv:2401.09407
-
Dynamic Relation Transformer for Contextual Text Block Detection 17 Jan 2024 · 0 repositories · arXiv:2401.09232
-
Efficient generative adversarial networks using linear additive-attention Transformers 17 Jan 2024 · 2 repositories · arXiv:2401.09596
-
From User Surveys to Telemetry-Driven AI Agents: Exploring the Potential of Personalized Productivity Solutions 17 Jan 2024 · 0 repositories · arXiv:2401.08960
-
Impact of Large Language Model Assistance on Patients Reading Clinical Notes: A Mixed-Methods Study 17 Jan 2024 · 0 repositories · arXiv:2401.09637
-
Improving Classification Performance With Human Feedback: Label a few, we label the rest 17 Jan 2024 · 0 repositories · arXiv:2401.09555
-
Learning from Implicit User Feedback, Emotions and Demographic Information in Task-Oriented and Document-Grounded Dialogues 17 Jan 2024 · 1 repository · arXiv:2401.09248
-
MADA: Meta-Adaptive Optimizers through hyper-gradient Descent 17 Jan 2024 · 0 repositories · arXiv:2401.08893
-
MSHyper: Multi-Scale Hypergraph Transformer for Long-Range Time Series Forecasting 17 Jan 2024 · 1 repository · arXiv:2401.09261Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Siamese Meets Diffusion Network: SMDNet for Enhanced Change Detection in High-Resolution RS Imagery 17 Jan 2024 · 0 repositories · arXiv:2401.09325
-
Evaluating LLMs' Mathematical and Coding Competency through Ontology-guided Interventions 17 Jan 2024 · 1 repository · arXiv:2401.09395Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Learning from Imperfect Demonstrations with Self-Supervision for Robotic Manipulation 17 Jan 2024 · 0 repositories · arXiv:2401.08957
-
SymTC: A Symbiotic Transformer-CNN Net for Instance Segmentation of Lumbar Spine MRI 17 Jan 2024 · 1 repository · arXiv:2401.09627
-
Trapped in texture bias? A large scale comparison of deep instance segmentation 17 Jan 2024 · 1 repository · arXiv:2401.09109
-
Application of LLM Agents in Recruitment: A Novel Framework for Resume Screening 16 Jan 2024 · 0 repositories · arXiv:2401.08315
-
B-Cos Aligned Transformers Learn Human-Interpretable Features 16 Jan 2024 · 0 repositories · arXiv:2401.08868
-
Code Generation with AlphaCodium: From Prompt Engineering to Flow Engineering 16 Jan 2024 · 2 repositories · arXiv:2401.08500Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation 16 Jan 2024 · 1 repository · arXiv:2401.08417Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples)
-
EmoLLMs: A Series of Emotional Large Language Models and Annotation Tools for Comprehensive Affective Analysis 16 Jan 2024 · 1 repository · arXiv:2401.08508Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples)
-
Enhancing Robustness of LLM-Synthetic Text Detectors for Academic Writing: A Comprehensive Analysis 16 Jan 2024 · 0 repositories · arXiv:2401.08046
-
Exploiting Inter-Layer Expert Affinity for Accelerating Mixture-of-Experts Model Inference 16 Jan 2024 · 1 repository · arXiv:2401.08383
-
Forging Vision Foundation Models for Autonomous Driving: Challenges, Methodologies, and Opportunities 16 Jan 2024 · 1 repository · arXiv:2401.08045
-
Hidden flaws behind expert-level accuracy of multimodal GPT-4 vision in medicine 16 Jan 2024 · 0 repositories · arXiv:2401.08396
-
Importance-Aware Image Segmentation-based Semantic Communication for Autonomous Driving 16 Jan 2024 · 0 repositories · arXiv:2401.10153
-
MARIO: MAth Reasoning with code Interpreter Output -- A Reproducible Pipeline 16 Jan 2024 · 2 repositories · arXiv:2401.08190Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
MMToM-QA: Multimodal Theory of Mind Question Answering 16 Jan 2024 · 1 repository · arXiv:2401.08743
-
RAG vs Fine-tuning: Pipelines, Tradeoffs, and a Case Study on Agriculture 16 Jan 2024 · 0 repositories · arXiv:2401.08406
-
RoTBench: A Multi-Level Benchmark for Evaluating the Robustness of Large Language Models in Tool Learning 16 Jan 2024 · 1 repository · arXiv:2401.08326Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 1 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 15 harvested samples)
-
Small Object Detection by DETR via Information Augmentation and Adaptive Feature Fusion 16 Jan 2024 · 0 repositories · arXiv:2401.08017
-
Solving Continual Offline Reinforcement Learning with Decision Transformer 16 Jan 2024 · 0 repositories · arXiv:2401.08478
-
Statistical Test for Attention Map in Vision Transformer 16 Jan 2024 · 1 repository · arXiv:2401.08169
-
Transcending the Limit of Local Window: Advanced Super-Resolution Transformer with Adaptive Token Dictionary 16 Jan 2024 · 1 repository · arXiv:2401.08209Syntology official (archive's flag): 11 ran · 11 ran (of which 9 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 15 harvested samples)
-
Tuning Language Models by Proxy 16 Jan 2024 · 2 repositories · arXiv:2401.08565Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Video Quality Assessment Based on Swin TransformerV2 and Coarse to Fine Strategy 16 Jan 2024 · 0 repositories · arXiv:2401.08522
-
A Novel Approach for Automatic Program Repair using Round-Trip Translation with Large Language Models 15 Jan 2024 · 1 repository · arXiv:2401.07994
-
Combining Image- and Geometric-based Deep Learning for Shape Regression: A Comparison to Pixel-level Methods for Segmentation in Chest X-Ray 15 Jan 2024 · 0 repositories · arXiv:2401.07542
-
Consolidating Trees of Robotic Plans Generated Using Large Language Models to Improve Reliability 15 Jan 2024 · 0 repositories · arXiv:2401.07868
-
Exploiting GPT-4 Vision for Zero-shot Point Cloud Understanding 15 Jan 2024 · 0 repositories · arXiv:2401.07572
-
Graph database while computationally efficient filters out quickly the ESG integrated equities in investment management 15 Jan 2024 · 0 repositories · arXiv:2401.07483
-
Graph Transformer GANs with Graph Masked Modeling for Architectural Layout Generation 15 Jan 2024 · 0 repositories · arXiv:2401.07721
-
Milestones in Bengali Sentiment Analysis leveraging Transformer-models: Fundamentals, Challenges and Future Directions 15 Jan 2024 · 0 repositories · arXiv:2401.07847
-
One for All: Toward Unified Foundation Models for Earth Vision 15 Jan 2024 · 0 repositories · arXiv:2401.07527
-
The Chronicles of RAG: The Retriever, the Chunk and the Generator 15 Jan 2024 · 0 repositories · arXiv:2401.07883
-
3D Landmark Detection on Human Point Clouds: A Benchmark and A Dual Cascade Point Transformer Framework 14 Jan 2024 · 0 repositories · arXiv:2401.07251
-
Harnessing Large Language Models Over Transformer Models for Detecting Bengali Depressive Social Media Text: A Comprehensive Study 14 Jan 2024 · 1 repository · arXiv:2401.07310