Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 70
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 70 of 190: papers 6,901 to 7,000 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Xiwu: A Basis Flexible and Learnable LLM for High Energy Physics 8 Apr 2024 · 1 repository · arXiv:2404.08001
-
A Multi-Level Framework for Accelerating Training Transformer Models 7 Apr 2024 · 1 repository · arXiv:2404.07999Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
MM-MATH: Advancing Multimodal Math Evaluation with Process Evaluation and Fine-grained Classification 7 Apr 2024 · 1 repository · arXiv:2404.05091
-
Contextual Chart Generation for Cyber Deception 7 Apr 2024 · 0 repositories · arXiv:2404.04854
-
CSA-Trans: Code Structure Aware Transformer for AST 7 Apr 2024 · 1 repository · arXiv:2404.05767Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Data Bias According to Bipol: Men are Naturally Right and It is the Role of Women to Follow Their Lead 7 Apr 2024 · 1 repository · arXiv:2404.04838
-
Dual-Scale Transformer for Large-Scale Single-Pixel Imaging 7 Apr 2024 · 1 repository · arXiv:2404.05001Syntology official (archive's flag): 6 ran · 6 ran (of which 3 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
GvT: A Graph-based Vision Transformer with Talking-Heads Utilizing Sparsity, Trained from Scratch on Small Datasets 7 Apr 2024 · 0 repositories · arXiv:2404.04924
-
Initial Exploration of Zero-Shot Privacy Utility Tradeoffs in Tabular Data Using GPT-4 7 Apr 2024 · 0 repositories · arXiv:2404.05047
-
Joint Reconstruction of 3D Human and Object via Contact-Based Refinement Transformer 7 Apr 2024 · 1 repository · arXiv:2404.04819Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
LHU-Net: A Light Hybrid U-Net for Cost-Efficient, High-Performance Volumetric Medical Image Segmentation 7 Apr 2024 · 1 repository · arXiv:2404.05102Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
PagPassGPT: Pattern Guided Password Guessing via Generative Pretrained Transformer 7 Apr 2024 · 1 repository · arXiv:2404.04886
-
Rethinking Diffusion Model for Multi-Contrast MRI Super-Resolution 7 Apr 2024 · 1 repository · arXiv:2404.04785Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 7 where Syntology's instrument failed) · 0 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
VMambaMorph: a Multi-Modality Deformable Image Registration Framework based on Visual State Space Model with Cross-Scan Module 7 Apr 2024 · 1 repository · arXiv:2404.05105
-
Convolutional Neural Network Transformer (CNNT) for Fluorescence Microscopy image Denoising with Improved Generalization and Fast Adaptation 6 Apr 2024 · 0 repositories · arXiv:2404.04726
-
Empowering Image Recovery_ A Multi-Attention Approach 6 Apr 2024 · 0 repositories · arXiv:2404.04617
-
Decomposition-based Unsupervised Domain Adaptation for Remote Sensing Image Semantic Segmentation 6 Apr 2024 · 1 repository · arXiv:2404.04531Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples)
-
IITK at SemEval-2024 Task 2: Exploring the Capabilities of LLMs for Safe Biomedical Natural Language Inference for Clinical Trials 6 Apr 2024 · 1 repository · arXiv:2404.04510
-
Large Language Model (LLM) AI text generation detection based on transformer deep learning algorithm 6 Apr 2024 · 0 repositories · arXiv:2405.06652
-
MACM: Utilizing a Multi-Agent System for Condition Mining in Solving Complex Mathematical Problems 6 Apr 2024 · 1 repository · arXiv:2404.04735Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Mixed-Query Transformer: A Unified Image Segmentation Architecture 6 Apr 2024 · 0 repositories · arXiv:2404.04469
-
RecGPT: Generative Personalized Prompts for Sequential Recommendation via ChatGPT Training Paradigm 6 Apr 2024 · 0 repositories · arXiv:2404.08675
-
Bayesian Additive Regression Networks 5 Apr 2024 · 1 repository · arXiv:2404.04425
-
Cleared for Takeoff? Compositional & Conditional Reasoning may be the Achilles Heel to (Flight-Booking) Language Agents 5 Apr 2024 · 0 repositories · arXiv:2404.04237
-
Effects of Different Prompts on the Quality of GPT-4 Responses to Dementia Care Questions 5 Apr 2024 · 0 repositories · arXiv:2404.08674
-
Context-Aware Aerial Object Detection: Leveraging Inter-Object and Background Relationships 5 Apr 2024 · 0 repositories · arXiv:2404.04140
-
PhysPT: Physics-aware Pretrained Transformer for Estimating Human Dynamics from Monocular Videos 5 Apr 2024 · 0 repositories · arXiv:2404.04430
-
player2vec: A Language Modeling Approach to Understand Player Behavior in Games 5 Apr 2024 · 0 repositories · arXiv:2404.04234
-
Re-pseudonymization Strategies for Smart Meter Data Are Not Robust to Deep Learning Profiling Attacks 5 Apr 2024 · 0 repositories · arXiv:2404.03948
-
Scope Ambiguities in Large Language Models 5 Apr 2024 · 1 repository · arXiv:2404.04332
-
The NES Video-Music Database: A Dataset of Symbolic Video Game Music Paired with Gameplay Videos 5 Apr 2024 · 1 repository · arXiv:2404.04420
-
Uformer: A UNet-Transformer fused robust end-to-end deep learning framework for real-time denoising of lung sounds 5 Apr 2024 · 0 repositories · arXiv:2404.04365
-
A Comprehensive Survey on Self-Supervised Learning for Recommendation 4 Apr 2024 · 1 repository · arXiv:2404.03354
-
A Directional Diffusion Graph Transformer for Recommendation 4 Apr 2024 · 0 repositories · arXiv:2404.03326
-
AutoWebGLM: A Large Language Model-based Web Navigating Agent 4 Apr 2024 · 1 repository · arXiv:2404.03648Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
Capabilities of Large Language Models in Control Engineering: A Benchmark Study on GPT-4, Claude 3 Opus, and Gemini 1.0 Ultra 4 Apr 2024 · 0 repositories · arXiv:2404.03647
-
CBR-RAG: Case-Based Reasoning for Retrieval Augmented Generation in LLMs for Legal Question Answering 4 Apr 2024 · 1 repository · arXiv:2404.04302
-
CONFLARE: CONFormal LArge language model REtrieval 4 Apr 2024 · 1 repository · arXiv:2404.04287
-
Conversational Disease Diagnosis via External Planner-Controlled Large Language Models 4 Apr 2024 · 1 repository · arXiv:2404.04292
-
Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences 4 Apr 2024 · 0 repositories · arXiv:2404.03715
-
Do Large Language Models Rank Fairly? An Empirical Study on the Fairness of LLMs as Rankers 4 Apr 2024 · 0 repositories · arXiv:2404.03192
-
Evaluating LLMs at Detecting Errors in LLM Responses 4 Apr 2024 · 1 repository · arXiv:2404.03602Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
NLP at UC Santa Cruz at SemEval-2024 Task 5: Legal Answer Validation using Few-Shot Multi-Choice QA 4 Apr 2024 · 1 repository · arXiv:2404.03150
-
On the Surprising Efficacy of Distillation as an Alternative to Pre-Training Small Models 4 Apr 2024 · 1 repository · arXiv:2404.03263Syntology official (archive's flag): 4 ran · 4 ran (of which 3 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
On the Theoretical Expressive Power and the Design Space of Higher-Order Graph Transformers 4 Apr 2024 · 1 repository · arXiv:2404.03380Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples) · 5 pointer-only (licence)
-
Probing Large Language Models for Scalar Adjective Lexical Semantics and Scalar Diversity Pragmatics 4 Apr 2024 · 1 repository · arXiv:2404.03301
-
RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis 4 Apr 2024 · 0 repositories · arXiv:2404.03204
-
Reason from Fallacy: Enhancing Large Language Models' Logical Reasoning through Logical Fallacy Understanding 4 Apr 2024 · 0 repositories · arXiv:2404.04293
-
Sailor: Open Language Models for South-East Asia 4 Apr 2024 · 3 repositories · arXiv:2404.03608Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 1 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 12 harvested samples) · 11 pointer-only (licence)
-
Towards Automated Movie Trailer Generation 4 Apr 2024 · 0 repositories · arXiv:2404.03477
-
Adaptive Cross-lingual Text Classification through In-Context One-Shot Demonstrations 3 Apr 2024 · 1 repository · arXiv:2404.02452Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
AI-Tutoring in Software Engineering Education 3 Apr 2024 · 0 repositories · arXiv:2404.02548
-
An Incomplete Loop: Deductive, Inductive, and Abductive Learning in Large Language Models 3 Apr 2024 · 0 repositories · arXiv:2404.03028
-
Attributions toward Artificial Agents in a modified Moral Turing Test 3 Apr 2024 · 0 repositories · arXiv:2406.11854
-
BCAmirs at SemEval-2024 Task 4: Beyond Words: A Multimodal and Multilingual Exploration of Persuasion in Memes 3 Apr 2024 · 1 repository · arXiv:2404.03022
-
Benchmarking Large Language Models for Persian: A Preliminary Study Focusing on ChatGPT 3 Apr 2024 · 1 repository · arXiv:2404.02403
-
Conifer: Improving Complex Constrained Instruction-Following Ability of Large Language Models 3 Apr 2024 · 1 repository · arXiv:2404.02823Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Cross-Architecture Transfer Learning for Linear-Cost Inference Transformers 3 Apr 2024 · 0 repositories · arXiv:2404.02684
-
Decision Transformer as a Foundation Model for Partially Observable Continuous Control 3 Apr 2024 · 0 repositories · arXiv:2404.02407
-
DeiT-LT Distillation Strikes Back for Vision Transformer Training on Long-Tailed Datasets 3 Apr 2024 · 2 repositories · arXiv:2404.02900
-
DPFT: Dual Perspective Fusion Transformer for Camera-Radar-based Object Detection 3 Apr 2024 · 1 repository · arXiv:2404.03015Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples)
-
FlightScope: An Experimental Comparative Review of Aircraft Detection Algorithms in Satellite Imagery 3 Apr 2024 · 1 repository · arXiv:2404.02877
-
Foundation Models for Structural Health Monitoring 3 Apr 2024 · 1 repository · arXiv:2404.02944
-
GPT-DETOX: An In-Context Learning-Based Paraphraser for Text Detoxification 3 Apr 2024 · 0 repositories · arXiv:2404.03052
-
Jump Self-attention: Capturing High-order Statistics in Transformers 3 Apr 2024 · 0 repositories
-
On Linearizing Structured Data in Encoder-Decoder Language Models: Insights from Text-to-SQL 3 Apr 2024 · 0 repositories · arXiv:2404.02389
-
On the Scalability of Diffusion-based Text-to-Image Generation 3 Apr 2024 · 0 repositories · arXiv:2404.02883
-
Optimizing the Deployment of Tiny Transformers on Low-Power MCUs 3 Apr 2024 · 1 repository · arXiv:2404.02945
-
RS-Mamba for Large Remote Sensing Image Dense Prediction 3 Apr 2024 · 1 repository · arXiv:2404.02668Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Task Agnostic Architecture for Algorithm Induction via Implicit Composition 3 Apr 2024 · 0 repositories · arXiv:2404.02450
-
Transformer-based Stagewise Decomposition for Large-Scale Multistage Stochastic Optimization 3 Apr 2024 · 0 repositories · arXiv:2404.02583
-
uTeBC-NLP at SemEval-2024 Task 9: Can LLMs be Lateral Thinkers? 3 Apr 2024 · 1 repository · arXiv:2404.02474Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Accelerating Transformer Pre-training with 2:4 Sparsity 2 Apr 2024 · 2 repositories · arXiv:2404.01847Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples)
-
Advancing LLM Reasoning Generalists with Preference Trees 2 Apr 2024 · 1 repository · arXiv:2404.02078Syntology official (archive's flag): 18 ran · 18 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 1 violated, 12 with no contract checked; 5 where Syntology's instrument failed) · 2 unverified (of 20 harvested samples) · 2 pointer-only (licence)
-
ASTRA: An Action Spotting TRAnsformer for Soccer Videos 2 Apr 2024 · 1 repository · arXiv:2404.01891
-
Stereotype Detection in LLMs: A Multiclass, Explainable, and Benchmark-Driven Approach 2 Apr 2024 · 0 repositories · arXiv:2404.01768
-
Automated User Story Generation with Test Case Specification Using Large Language Model 2 Apr 2024 · 0 repositories · arXiv:2404.01558
-
Bidirectional Multi-Scale Implicit Neural Representations for Image Deraining 2 Apr 2024 · 1 repository · arXiv:2404.01547Syntology official (archive's flag): 19 ran · 19 ran (of which 11 constructed an object rather than computing a result; 16 with no instrument failure: 0 honoured, 0 violated, 16 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 21 harvested samples) · 21 pointer-only (licence)
-
CLAPNQ: Cohesive Long-form Answers from Passages in Natural Questions for RAG systems 2 Apr 2024 · 1 repository · arXiv:2404.02103
-
CMAT: A Multi-Agent Collaboration Tuning Framework for Enhancing Small Language Models 2 Apr 2024 · 1 repository · arXiv:2404.01663
-
Collapse of Self-trained Language Models 2 Apr 2024 · 1 repository · arXiv:2404.02305Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples)
-
Comparative Study of Domain Driven Terms Extraction Using Large Language Models 2 Apr 2024 · 0 repositories · arXiv:2404.02330
-
CSST Strong Lensing Preparation: a Framework for Detecting Strong Lenses in the Multi-color Imaging Survey by the China Survey Space Telescope (CSST) 2 Apr 2024 · 0 repositories · arXiv:2404.01780
-
Deconstructing In-Context Learning: Understanding Prompts via Corruption 2 Apr 2024 · 1 repository · arXiv:2404.02054
-
EGTR: Extracting Graph from Transformer for Scene Graph Generation 2 Apr 2024 · 1 repository · arXiv:2404.02072Syntology official (archive's flag): 16 ran · 16 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 1 honoured, 0 violated, 11 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 17 harvested samples) · 4 pointer-only (licence)
-
Enhancing Inference Efficiency of Large Language Models: Investigating Optimization Strategies and Architectural Innovations 2 Apr 2024 · 0 repositories · arXiv:2404.05741
-
Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack 2 Apr 2024 · 0 repositories · arXiv:2404.01833
-
IndoCulture: Exploring Geographically-Influenced Cultural Commonsense Reasoning Across Eleven Indonesian Provinces 2 Apr 2024 · 0 repositories · arXiv:2404.01854
-
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks 2 Apr 2024 · 1 repository · arXiv:2404.02151Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples)
-
METAL: Towards Multilingual Meta-Evaluation 2 Apr 2024 · 0 repositories · arXiv:2404.01667
-
Minimize Quantization Output Error with Bias Compensation 2 Apr 2024 · 1 repository · arXiv:2404.01892
-
Octopus: On-device language model for function calling of software APIs 2 Apr 2024 · 0 repositories · arXiv:2404.01549
-
Octopus v2: On-device language model for super agent 2 Apr 2024 · 0 repositories · arXiv:2404.01744
-
PATCH! Psychometrics-Assis{T}ed Ben{CH}marking of Large Language Models against Human Populations: A Case Study of Proficiency in 8th Grade Mathematics 2 Apr 2024 · 1 repository · arXiv:2404.01799
-
Symbolic Prompt Program Search: A Structure-Aware Approach to Efficient Compile-Time Prompt Optimization 2 Apr 2024 · 1 repository · arXiv:2404.02319
-
RAT: Retrieval-Augmented Transformer for Click-Through Rate Prediction 2 Apr 2024 · 1 repository · arXiv:2404.02249
-
Release of Pre-Trained Models for the Japanese Language 2 Apr 2024 · 0 repositories · arXiv:2404.01657
-
Samba: Semantic Segmentation of Remotely Sensed Images with State Space Model 2 Apr 2024 · 1 repository · arXiv:2404.01705Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples) · 7 pointer-only (licence)
-
Scene Adaptive Sparse Transformer for Event-based Object Detection 2 Apr 2024 · 1 repository · arXiv:2404.01882Syntology official (archive's flag): 19 ran · 19 ran (of which 0 constructed an object rather than computing a result; 19 with no instrument failure: 0 honoured, 0 violated, 19 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 23 harvested samples)
-
SGSH: Stimulate Large Language Models with Skeleton Heuristics for Knowledge Base Question Generation 2 Apr 2024 · 1 repository · arXiv:2404.01923