Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 54
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 54 of 190: papers 5,301 to 5,400 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
MINDECHO: Role-Playing Language Agents for Key Opinion Leaders 7 Jul 2024 · 0 repositories · arXiv:2407.05305
-
Multimodal Prompt Learning with Missing Modalities for Sentiment Analysis and Emotion Recognition 7 Jul 2024 · 1 repository · arXiv:2407.05374Syntology official (archive's flag): 2 ran · 2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified; every one of the 2 samples that ran constructed an object rather than computing a result (of 6 harvested samples)
-
PTaRL: Prototype-based Tabular Representation Learning via Space Calibration 7 Jul 2024 · 1 repository · arXiv:2407.05364Syntology 4 ran (of which 3 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples) · 1 pointer-only (licence)
-
CLIPVQA:Video Quality Assessment via CLIP 6 Jul 2024 · 1 repository · arXiv:2407.04928
-
EVA-Score: Evaluating Abstractive Long-form Summarization on Informativeness through Extraction and Validation 6 Jul 2024 · 0 repositories · arXiv:2407.04969
-
How do you know that? Teaching Generative Language Models to Reference Answers to Biomedical Questions 6 Jul 2024 · 1 repository · arXiv:2407.05015
-
Integer-only Quantized Transformers for Embedded FPGA-based Time-series Forecasting in AIoT 6 Jul 2024 · 0 repositories · arXiv:2407.11041
-
PRANCE: Joint Token-Optimization and Structural Channel-Pruning for Adaptive ViT Inference 6 Jul 2024 · 1 repository · arXiv:2407.05010
-
RULE: Reliable Multimodal RAG for Factuality in Medical Vision Language Models 6 Jul 2024 · 1 repository · arXiv:2407.05131Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
SHINE: Saliency-aware HIerarchical NEgative Ranking for Compositional Temporal Grounding 6 Jul 2024 · 1 repository · arXiv:2407.05118Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 1 violated, 8 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
Solving for X and Beyond: Can Large Language Models Solve Complex Math Problems with More-Than-Two Unknowns? 6 Jul 2024 · 1 repository · arXiv:2407.05134
-
The Solution for the AIGC Inference Performance Optimization Competition 6 Jul 2024 · 0 repositories · arXiv:2407.04991
-
Are LLMs Correctly Integrated into Software Systems? 6 Jul 2024 · 0 repositories · arXiv:2407.05138
-
Aligning Model Evaluations with Human Preferences: Mitigating Token Count Bias in Language Model Assessments 5 Jul 2024 · 0 repositories · arXiv:2407.12847
-
ANAH-v2: Scaling Analytical Hallucination Annotation of Large Language Models 5 Jul 2024 · 1 repository · arXiv:2407.04693Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 4 harvested samples)
-
Are Large Language Models Strategic Decision Makers? A Study of Performance and Bias in Two-Player Non-Zero-Sum Games 5 Jul 2024 · 0 repositories · arXiv:2407.04467
-
Associative Recurrent Memory Transformer 5 Jul 2024 · 1 repository · arXiv:2407.04841Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 10 unverified (of 14 harvested samples)
-
GPT vs RETRO: Exploring the Intersection of Retrieval and Parameter-Efficient Fine-Tuning 5 Jul 2024 · 0 repositories · arXiv:2407.04528
-
HCS-TNAS: Hybrid Constraint-driven Semi-supervised Transformer-NAS for Ultrasound Image Segmentation 5 Jul 2024 · 0 repositories · arXiv:2407.04203
-
Improving ensemble extreme precipitation forecasts using generative artificial intelligence 5 Jul 2024 · 0 repositories · arXiv:2407.04882
-
Learning to (Learn at Test Time): RNNs with Expressive Hidden States 5 Jul 2024 · 3 repositories · arXiv:2407.04620Syntology official (archive's flag): 2 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
Robust Decision Transformer: Tackling Data Corruption in Offline RL via Sequence Modeling 5 Jul 2024 · 0 repositories · arXiv:2407.04285
-
Strengthening Structural Inductive Biases by Pre-training to Perform Syntactic Transformations 5 Jul 2024 · 1 repository · arXiv:2407.04543
-
Using LLMs to label medical papers according to the CIViC evidence model 5 Jul 2024 · 1 repository · arXiv:2407.04466
-
A Computer Vision Approach to Estimate the Localized Sea State 4 Jul 2024 · 0 repositories · arXiv:2407.03755
-
ADAPT: Multimodal Learning for Detecting Physiological Changes under Missing Modalities 4 Jul 2024 · 1 repository · arXiv:2407.03836
-
Adaptive Step-size Perception Unfolding Network with Non-local Hybrid Attention for Hyperspectral Image Reconstruction 4 Jul 2024 · 0 repositories · arXiv:2407.04024
-
Diverse and Fine-Grained Instruction-Following Ability Exploration with Synthetic Data 4 Jul 2024 · 0 repositories · arXiv:2407.03942
-
DSLR: Document Refinement with Sentence-Level Re-ranking and Reconstruction to Enhance Retrieval-Augmented Generation 4 Jul 2024 · 0 repositories · arXiv:2407.03627
-
Evaluating Language Model Context Windows: A "Working Memory" Test and Inference-time Correction 4 Jul 2024 · 1 repository · arXiv:2407.03651
-
From Data to Commonsense Reasoning: The Use of Large Language Models for Explainable AI 4 Jul 2024 · 0 repositories · arXiv:2407.03778
-
Generalizing Graph Transformers Across Diverse Graphs and Tasks via Pre-Training on Industrial-Scale Data 4 Jul 2024 · 0 repositories · arXiv:2407.03953
-
GPT-4 vs. Human Translators: A Comprehensive Evaluation of Translation Quality Across Languages, Domains, and Expertise Levels 4 Jul 2024 · 0 repositories · arXiv:2407.03658
-
QET: Enhancing Quantized LLM Parameters and KV cache Compression through Element Substitution and Residual Clustering 4 Jul 2024 · 0 repositories · arXiv:2407.03637
-
NutriBench: A Dataset for Evaluating Large Language Models on Nutrition Estimation from Meal Descriptions 4 Jul 2024 · 0 repositories · arXiv:2407.12843
-
On the Benchmarking of LLMs for Open-Domain Dialogue Evaluation 4 Jul 2024 · 0 repositories · arXiv:2407.03841
-
On the Effectiveness of Acoustic BPE in Decoder-Only TTS 4 Jul 2024 · 0 repositories · arXiv:2407.03892
-
Controllable Conversations: Planning-Based Dialogue Agent with Large Language Models 4 Jul 2024 · 1 repository · arXiv:2407.03884
-
Query-Guided Self-Supervised Summarization of Nursing Notes 4 Jul 2024 · 0 repositories · arXiv:2407.04125
-
Question-Analysis Prompting Improves LLM Performance in Reasoning Tasks 4 Jul 2024 · 0 repositories · arXiv:2407.03624
-
Slice-100K: A Multimodal Dataset for Extrusion-based 3D Printing 4 Jul 2024 · 0 repositories · arXiv:2407.04180
-
Solving Zebra Puzzles Using Constraint-Guided Multi-Agent Systems 4 Jul 2024 · 0 repositories · arXiv:2407.03956
-
Towards Automating Text Annotation: A Case Study on Semantic Proximity Annotation using GPT-4 4 Jul 2024 · 0 repositories · arXiv:2407.04130
-
A Comparative Study of DSL Code Generation: Fine-Tuning vs. Optimized Retrieval Augmentation 3 Jul 2024 · 0 repositories · arXiv:2407.02742
-
A Unified Framework for 3D Scene Understanding 3 Jul 2024 · 1 repository · arXiv:2407.03263Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples)
-
AgentInstruct: Toward Generative Teaching with Agentic Flows 3 Jul 2024 · 0 repositories · arXiv:2407.03502
-
Gradient descent with generalized Newton's method 3 Jul 2024 · 1 repository · arXiv:2407.02772Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Fine-Grained Scene Image Classification with Modality-Agnostic Adapter 3 Jul 2024 · 1 repository · arXiv:2407.02769
-
Fisher-aware Quantization for DETR Detectors with Critical-category Objectives 3 Jul 2024 · 0 repositories · arXiv:2407.03442
-
Graph and Skipped Transformer: Exploiting Spatial and Temporal Modeling Capacities for Efficient 3D Human Pose Estimation 3 Jul 2024 · 0 repositories · arXiv:2407.02990
-
Human-like Linguistic Biases in Neural Speech Models: Phonetic Categorization and Phonotactic Constraints in Wav2Vec2.0 3 Jul 2024 · 1 repository · arXiv:2407.03005
-
Improving LLM Abilities in Idiomatic Translation 3 Jul 2024 · 0 repositories · arXiv:2407.03518
-
ISWSST: Index-space-wave State Superposition Transformers for Multispectral Remotely Sensed Imagery Semantic Segmentation 3 Jul 2024 · 0 repositories · arXiv:2407.03033
-
LANE: Logic Alignment of Non-tuning Large Language Models and Online Recommendation Systems for Explainable Reason Generation 3 Jul 2024 · 0 repositories · arXiv:2407.02833
-
Large Language Models as Evaluators for Scientific Synthesis 3 Jul 2024 · 0 repositories · arXiv:2407.02977
-
Learning to Reduce: Towards Improving Performance of Large Language Models on Structured Data 3 Jul 2024 · 0 repositories · arXiv:2407.02750
-
MVGT: A Multi-view Graph Transformer Based on Spatial Relations for EEG Emotion Recognition 3 Jul 2024 · 0 repositories · arXiv:2407.03131
-
ObfuscaTune: Obfuscated Offsite Fine-tuning and Inference of Proprietary LLMs on Private Datasets 3 Jul 2024 · 0 repositories · arXiv:2407.02960
-
On Large Language Models in National Security Applications 3 Jul 2024 · 1 repository · arXiv:2407.03453
-
OSPC: Artificial VLM Features for Hateful Meme Detection 3 Jul 2024 · 0 repositories · arXiv:2407.12836
-
RDBE: Reasoning Distillation-Based Evaluation Enhances Automatic Essay Scoring 3 Jul 2024 · 0 repositories · arXiv:2407.13781
-
Regurgitative Training: The Value of Real Data in Training Large Language Models 3 Jul 2024 · 0 repositories · arXiv:2407.12835
-
Self-supervised Vision Transformer are Scalable Generative Models for Domain Generalization 3 Jul 2024 · 1 repository · arXiv:2407.02900
-
SemioLLM: Assessing Large Language Models for Semiological Analysis in Epilepsy Research 3 Jul 2024 · 0 repositories · arXiv:2407.03004
-
Generative AI Enables EEG Super-Resolution via Spatio-Temporal Adaptive Diffusion Learning 3 Jul 2024 · 0 repositories · arXiv:2407.03089
-
TheoremLlama: Transforming General-Purpose LLMs into Lean4 Experts 3 Jul 2024 · 1 repository · arXiv:2407.03203Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; the one sample that ran constructed an object rather than computing a result (of 2 harvested samples) · 2 pointer-only (licence)
-
A Depression Detection Method Based on Multi-Modal Feature Fusion Using Cross-Attention 2 Jul 2024 · 0 repositories · arXiv:2407.12825
-
Assessing the Code Clone Detection Capability of Large Language Models 2 Jul 2024 · 0 repositories · arXiv:2407.02402
-
Beyond Numeric Awards: In-Context Dueling Bandits with LLM Agents 2 Jul 2024 · 0 repositories · arXiv:2407.01887
-
Deep Learning Based Apparent Diffusion Coefficient Map Generation from Multi-parametric MR Images for Patients with Diffuse Gliomas 2 Jul 2024 · 0 repositories · arXiv:2407.02616
-
Fake News Detection and Manipulation Reasoning via Large Vision-Language Models 2 Jul 2024 · 0 repositories · arXiv:2407.02042
-
GPTCast: a weather language model for precipitation nowcasting 2 Jul 2024 · 1 repository · arXiv:2407.02089
-
GRASP: A Grid-Based Benchmark for Evaluating Commonsense Spatial Reasoning 2 Jul 2024 · 0 repositories · arXiv:2407.01892
-
Improving Visual Storytelling with Multimodal Large Language Models 2 Jul 2024 · 0 repositories · arXiv:2407.02586
-
Integrate the Essence and Eliminate the Dross: Fine-Grained Self-Consistency for Free-Form Language Generation 2 Jul 2024 · 1 repository · arXiv:2407.02056Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
LLM-Select: Feature Selection with Large Language Models 2 Jul 2024 · 0 repositories · arXiv:2407.02694
-
MeMemo: On-device Retrieval Augmentation for Private and Personalized Text Generation 2 Jul 2024 · 1 repository · arXiv:2407.01972
-
Language Model Alignment in Multilingual Trolley Problems 2 Jul 2024 · 2 repositories · arXiv:2407.02273Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Open foundation models for Azerbaijani language 2 Jul 2024 · 0 repositories · arXiv:2407.02337
-
OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation 2 Jul 2024 · 0 repositories · arXiv:2407.02371
-
RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs 2 Jul 2024 · 0 repositories · arXiv:2407.02485
-
SeqAR: Jailbreak LLMs with Sequential Auto-Generated Characters 2 Jul 2024 · 1 repository · arXiv:2407.01902Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
The Art of Saying No: Contextual Noncompliance in Language Models 2 Jul 2024 · 1 repository · arXiv:2407.12043Syntology 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
Embedded Prompt Tuning: Towards Enhanced Calibration of Pretrained Models for Medical Images 1 Jul 2024 · 1 repository · arXiv:2407.01003
-
BERGEN: A Benchmarking Library for Retrieval-Augmented Generation 1 Jul 2024 · 1 repository · arXiv:2407.01102Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Deciphering the Factors Influencing the Efficacy of Chain-of-Thought: Probability, Memorization, and Noisy Reasoning 1 Jul 2024 · 1 repository · arXiv:2407.01687Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Domain Influence in MRI Medical Image Segmentation: spatial versus k-space inputs 1 Jul 2024 · 1 repository · arXiv:2407.01367
-
Face4RAG: Factual Consistency Evaluation for Retrieval Augmented Generation in Chinese 1 Jul 2024 · 0 repositories · arXiv:2407.01080
-
Ground Every Sentence: Improving Retrieval-Augmented LLMs with Interleaved Reference-Claim Generation 1 Jul 2024 · 0 repositories · arXiv:2407.01796
-
How Does Overparameterization Affect Features? 1 Jul 2024 · 0 repositories · arXiv:2407.00968
-
Hybrid RAG-empowered Multi-modal LLM for Secure Data Management in Internet of Medical Things: A Diffusion-based Contract Approach 1 Jul 2024 · 0 repositories · arXiv:2407.00978
-
Hypformer: Exploring Efficient Hyperbolic Transformer Fully in Hyperbolic Space 1 Jul 2024 · 1 repository · arXiv:2407.01290Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples) · 1 pointer-only (licence)
-
Image-to-Text Logic Jailbreak: Your Imagination can Help You Do Anything 1 Jul 2024 · 0 repositories · arXiv:2407.02534
-
Increasing Model Capacity for Free: A Simple Strategy for Parameter Efficient Fine-tuning 1 Jul 2024 · 1 repository · arXiv:2407.01320Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples) · 7 pointer-only (licence)
-
Investigating the potential of Sparse Mixtures-of-Experts for multi-domain neural machine translation 1 Jul 2024 · 0 repositories · arXiv:2407.01126
-
Large Language Model Enhanced Knowledge Representation Learning: A Survey 1 Jul 2024 · 0 repositories · arXiv:2407.00936
-
Multi-branch CNN and grouping cascade attention for medical image classification 1 Jul 2024 · 0 repositories
-
Multi-State-Action Tokenisation in Decision Transformers for Multi-Discrete Action Spaces 1 Jul 2024 · 0 repositories · arXiv:2407.01310
-
Papez: Resource-Efficient Speech Separation with Auditory Working Memory 1 Jul 2024 · 1 repository · arXiv:2407.00888
-
Pictures Of MIDI: Controlled Music Generation via Graphical Prompts for Image-Based Diffusion Inpainting 1 Jul 2024 · 0 repositories · arXiv:2407.01499