Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 23
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 23 of 190: papers 2,201 to 2,300 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
SRMT: Shared Memory for Multi-agent Lifelong Pathfinding 22 Jan 2025 · 1 repository · arXiv:2501.13200
-
T-Graphormer: Using Transformers for Spatiotemporal Forecasting 22 Jan 2025 · 1 repository · arXiv:2501.13274
-
A Survey of Graph Retrieval-Augmented Generation for Customized Large Language Models 21 Jan 2025 · 1 repository · arXiv:2501.13958
-
Advancing the Understanding and Evaluation of AR-Generated Scenes: When Vision-Language Models Shine and Stumble 21 Jan 2025 · 1 repository · arXiv:2501.13964
-
ALoFTRAG: Automatic Local Fine Tuning for Retrieval Augmented Generation 21 Jan 2025 · 1 repository · arXiv:2501.11929
-
Automatic Labelling with Open-source LLMs using Dynamic Label Schema Integration 21 Jan 2025 · 0 repositories · arXiv:2501.12332
-
Continuous 3D Perception Model with Persistent State 21 Jan 2025 · 0 repositories · arXiv:2501.12387
-
Divide-Then-Aggregate: An Efficient Tool Learning Method via Parallel Tool Invocation 21 Jan 2025 · 0 repositories · arXiv:2501.12432
-
DLEN: Dual Branch of Transformer for Low-Light Image Enhancement in Dual Domains 21 Jan 2025 · 0 repositories · arXiv:2501.12235
-
Episodic Memories Generation and Evaluation Benchmark for Large Language Models 21 Jan 2025 · 1 repository · arXiv:2501.13121Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 14 harvested samples) · 14 pointer-only (licence)
-
FOCUS: First Order Concentrated Updating Scheme 21 Jan 2025 · 0 repositories · arXiv:2501.12243
-
Harnessing Generative Pre-Trained Transformer for Datacenter Packet Trace Generation 21 Jan 2025 · 0 repositories · arXiv:2501.12033
-
Med-R²: Crafting Trustworthy LLM Physicians via Retrieval and Reasoning of Evidence-Based Medicine 21 Jan 2025 · 1 repository · arXiv:2501.11885
-
Network-informed Prompt Engineering against Organized Astroturf Campaigns under Extreme Class Imbalance 21 Jan 2025 · 1 repository · arXiv:2501.11849
-
Panoramic Interests: Stylistic-Content Aware Personalized Headline Generation 21 Jan 2025 · 1 repository · arXiv:2501.11900
-
Towards Accurate Unified Anomaly Segmentation 21 Jan 2025 · 1 repository · arXiv:2501.12295
-
Vision-Language Models for Automated Chest X-ray Interpretation: Leveraging ViT and GPT-2 21 Jan 2025 · 0 repositories · arXiv:2501.12356
-
Adaptive parameters identification for nonlinear dynamics using deep permutation invariant networks 20 Jan 2025 · 0 repositories · arXiv:2501.11350
-
DLinear-based Prediction of Remaining Useful Life of Lithium-Ion Batteries: Feature Engineering through Explainable Artificial Intelligence 20 Jan 2025 · 0 repositories · arXiv:2501.11542
-
Explainable Lane Change Prediction for Near-Crash Scenarios Using Knowledge Graph Embeddings and Retrieval Augmented Generation 20 Jan 2025 · 0 repositories · arXiv:2501.11560
-
Generative AI-enabled Blockage Prediction for Robust Dual-Band mmWave Communication 20 Jan 2025 · 0 repositories · arXiv:2501.11763
-
Glinthawk: A Two-Tiered Architecture for Offline LLM Inference 20 Jan 2025 · 1 repository · arXiv:2501.11779
-
KEIR @ ECIR 2025: The Second Workshop on Knowledge-Enhanced Information Retrieval 20 Jan 2025 · 0 repositories · arXiv:2501.11499
-
PIKE-RAG: sPecIalized KnowledgE and Rationale Augmented Generation 20 Jan 2025 · 1 repository · arXiv:2501.11551Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Synthetic Data Can Mislead Evaluations: Membership Inference as Machine Text Detection 20 Jan 2025 · 0 repositories · arXiv:2501.11786
-
Trustformer: A Trusted Federated Transformer 20 Jan 2025 · 0 repositories · arXiv:2501.11706
-
TutorLLM: Customizing Learning Recommendations with Knowledge Tracing and Retrieval-Augmented Generation 20 Jan 2025 · 0 repositories · arXiv:2502.15709
-
Chain-of-Reasoning: Towards Unified Mathematical Reasoning in Large Language Models via a Multi-Paradigm Perspective 19 Jan 2025 · 0 repositories · arXiv:2501.11110
-
From Arabic Text to Puzzles: LLM-Driven Development of Arabic Educational Crosswords 19 Jan 2025 · 0 repositories · arXiv:2501.11035
-
A CNN-Transformer for Classification of Longitudinal 3D MRI Images -- A Case Study on Hepatocellular Carcinoma Prediction 18 Jan 2025 · 1 repository · arXiv:2501.10733
-
Dynamic Trend Fusion Module for Traffic Flow Prediction 18 Jan 2025 · 1 repository · arXiv:2501.10796
-
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models 18 Jan 2025 · 0 repositories · arXiv:2501.10714
-
GEC-RAG: Improving Generative Error Correction via Retrieval-Augmented Generation for Automatic Speech Recognition Systems 18 Jan 2025 · 0 repositories · arXiv:2501.10734
-
LD-DETR: Loop Decoder DEtection TRansformer for Video Moment Retrieval and Highlight Detection 18 Jan 2025 · 1 repository · arXiv:2501.10787
-
Neural Algorithmic Reasoning for Hypergraphs with Looped Transformers 18 Jan 2025 · 0 repositories · arXiv:2501.10688
-
Visual RAG: Expanding MLLM visual knowledge without fine-tuning 18 Jan 2025 · 0 repositories · arXiv:2501.10834
-
4bit-Quantization in Vector-Embedding for RAG 17 Jan 2025 · 1 repository · arXiv:2501.10534
-
AirRAG: Activating Intrinsic Reasoning for Retrieval Augmented Generation via Tree-based Search 17 Jan 2025 · 0 repositories · arXiv:2501.10053
-
Bias in Decision-Making for AI's Ethical Dilemmas: A Comparative Study of ChatGPT and Claude 17 Jan 2025 · 1 repository · arXiv:2501.10484
-
Enhancing the Reliability in Machine Learning for Gravitational Wave Parameter Estimation with Attention-Based Models 17 Jan 2025 · 0 repositories · arXiv:2501.10486
-
PaSa: An LLM Agent for Comprehensive Academic Paper Search 17 Jan 2025 · 1 repository · arXiv:2501.10120Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Passage Segmentation of Documents for Extractive Question Answering 17 Jan 2025 · 0 repositories · arXiv:2501.09940
-
Self-Clustering Graph Transformer Approach to Model Resting-State Functional Brain Activity 17 Jan 2025 · 0 repositories · arXiv:2501.16345
-
A Simple Aerial Detection Baseline of Multimodal Language Models 16 Jan 2025 · 1 repository · arXiv:2501.09720
-
Confidence Estimation for Error Detection in Text-to-SQL Systems 16 Jan 2025 · 1 repository · arXiv:2501.09527
-
Generalized Single-Image-Based Morphing Attack Detection Using Deep Representations from Vision Transformer 16 Jan 2025 · 0 repositories · arXiv:2501.09817
-
HSPFormer: Hierarchical Spatial Perception Transformer for Semantic Segmentation 16 Jan 2025 · 1 repository
-
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation 16 Jan 2025 · 0 repositories · arXiv:2501.09755
-
Perspective Transition of Large Language Models for Solving Subjective Tasks 16 Jan 2025 · 0 repositories · arXiv:2501.09265
-
Practical Continual Forgetting for Pre-trained Vision Models 16 Jan 2025 · 1 repository · arXiv:2501.09705
-
Towards Robust and Realistic Human Pose Estimation via WiFi Signals 16 Jan 2025 · 1 repository · arXiv:2501.09411
-
Unified Face Matching and Physical-Digital Spoofing Attack Detection 16 Jan 2025 · 0 repositories · arXiv:2501.09635
-
Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG 15 Jan 2025 · 1 repository · arXiv:2501.09136
-
Attention is All You Need Until You Need Retention 15 Jan 2025 · 0 repositories · arXiv:2501.09166
-
BRIGHT-VO: Brightness-Guided Hybrid Transformer for Visual Odometry with Multi-modality Refinement Module 15 Jan 2025 · 1 repository · arXiv:2501.08659
-
CT-PatchTST: Channel-Time Patch Time-Series Transformer for Long-Term Renewable Energy Forecasting 15 Jan 2025 · 0 repositories · arXiv:2501.08620
-
Enhanced Large Language Models for Effective Screening of Depression and Anxiety 15 Jan 2025 · 0 repositories · arXiv:2501.08769
-
Generative AI Takes a Statistics Exam: A Comparison of Performance between ChatGPT3.5, ChatGPT4, and ChatGPT4o-mini 15 Jan 2025 · 0 repositories · arXiv:2501.09171
-
MIAFEx: An Attention-based Feature Extraction Method for Medical Image Classification 15 Jan 2025 · 0 repositories · arXiv:2501.08562
-
Multi-View Transformers for Airway-To-Lung Ratio Inference on Cardiac CT Scans: The C4R Study 15 Jan 2025 · 0 repositories · arXiv:2501.08902
-
Multimodal Fake News Video Explanation: Dataset, Analysis and Evaluation 15 Jan 2025 · 0 repositories · arXiv:2501.08514
-
SuperSAM: Crafting a SAM Supernetwork via Structured Pruning and Unstructured Parameter Prioritization 15 Jan 2025 · 1 repository · arXiv:2501.08504
-
SwinTExCo: Exemplar-based video colorization using Swin Transformer 15 Jan 2025 · 1 repository
-
The Impact of Big Five Personality Traits on AI Agent Decision-Making in Public Spaces: A Social Simulation Study 15 Jan 2025 · 0 repositories · arXiv:2503.15497
-
A Driver Advisory System Based on Large Language Model for High-speed Train 14 Jan 2025 · 0 repositories · arXiv:2501.07837
-
Active Sampling for Node Attribute Completion on Graphs 14 Jan 2025 · 0 repositories · arXiv:2501.08450
-
ASTRID -- An Automated and Scalable TRIaD for the Evaluation of RAG-based Clinical Question Answering Systems 14 Jan 2025 · 0 repositories · arXiv:2501.08208
-
Comparative Analysis of Efficient Adapter-Based Fine-Tuning of State-of-the-Art Transformer Models 14 Jan 2025 · 0 repositories · arXiv:2501.08271
-
Decision Transformers for RIS-Assisted Systems with Diffusion Model-Based Channel Acquisition 14 Jan 2025 · 0 repositories · arXiv:2501.08007
-
Decoding Interpretable Logic Rules from Neural Networks 14 Jan 2025 · 0 repositories · arXiv:2501.08281
-
Efficient Deep Learning-based Forward Solvers for Brain Tumor Growth Models 14 Jan 2025 · 1 repository · arXiv:2501.08226
-
Eliciting In-context Retrieval and Reasoning for Long-context Large Language Models 14 Jan 2025 · 0 repositories · arXiv:2501.08248
-
EmoNeXt: an Adapted ConvNeXt for Facial Emotion Recognition 14 Jan 2025 · 1 repository · arXiv:2501.08199
-
Exploring Narrative Clustering in Large Language Models: A Layerwise Analysis of BERT 14 Jan 2025 · 0 repositories · arXiv:2501.08053
-
Exploring Robustness of Multilingual LLMs on Real-World Noisy Data 14 Jan 2025 · 1 repository · arXiv:2501.08322
-
Investigating Energy Efficiency and Performance Trade-offs in LLM Inference Across Tasks and DVFS Settings 14 Jan 2025 · 0 repositories · arXiv:2501.08219
-
Large Language Models For Text Classification: Case Study And Comprehensive Review 14 Jan 2025 · 0 repositories · arXiv:2501.08457
-
Optimizing Language Models for Grammatical Acceptability: A Comparative Study of Fine-Tuning Techniques 14 Jan 2025 · 0 repositories · arXiv:2501.07853
-
PokerBench: Training Large Language Models to become Professional Poker Players 14 Jan 2025 · 1 repository · arXiv:2501.08328
-
PSReg: Prior-guided Sparse Mixture of Experts for Point Cloud Registration 14 Jan 2025 · 0 repositories · arXiv:2501.07762
-
ReARTeR: Retrieval-Augmented Reasoning with Trustworthy Process Rewarding 14 Jan 2025 · 1 repository · arXiv:2501.07861
-
Towards Lightweight Time Series Forecasting: a Patch-wise Transformer with Weak Data Enriching 14 Jan 2025 · 0 repositories · arXiv:2501.10448
-
Transforming Indoor Localization: Advanced Transformer Architecture for NLOS Dominated Wireless Environments with Distributed Sensors 14 Jan 2025 · 0 repositories · arXiv:2501.07774
-
UFGraphFR: An attempt at a federated recommendation system based on user text characteristics 14 Jan 2025 · 1 repository · arXiv:2501.08044
-
D3MES: Diffusion Transformer with multihead equivariant self-attention for 3D molecule generation 13 Jan 2025 · 1 repository · arXiv:2501.07077
-
EdgeTAM: On-Device Track Anything Model 13 Jan 2025 · 1 repository · arXiv:2501.07256Syntology official: no sample here; runs from other or unrecorded repositories · 7 ran (of which 2 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 1 violated, 2 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples)
-
Enhancing Retrieval-Augmented Generation: A Study of Best Practices 13 Jan 2025 · 1 repository · arXiv:2501.07391Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Estimating Musical Surprisal in Audio 13 Jan 2025 · 1 repository · arXiv:2501.07474
-
FinerWeb-10BT: Refining Web Data with LLM-Based Line-Level Filtering 13 Jan 2025 · 1 repository · arXiv:2501.07314
-
Future-Conditioned Recommendations with Multi-Objective Controllable Decision Transformer 13 Jan 2025 · 0 repositories · arXiv:2501.07212
-
GPT as a Monte Carlo Language Tree: A Probabilistic Perspective 13 Jan 2025 · 0 repositories · arXiv:2501.07641
-
How GPT learns layer by layer 13 Jan 2025 · 1 repository · arXiv:2501.07108
-
MathReader : Text-to-Speech for Mathematical Documents 13 Jan 2025 · 1 repository · arXiv:2501.07088
-
Parallel Key-Value Cache Fusion for Position Invariant RAG 13 Jan 2025 · 0 repositories · arXiv:2501.07523
-
Scaling Up ESM2 Architectures for Long Protein Sequences Analysis: Long and Quantized Approaches 13 Jan 2025 · 0 repositories · arXiv:2501.07747
-
SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing 13 Jan 2025 · 1 repository · arXiv:2501.07554
-
WebWalker: Benchmarking LLMs in Web Traversal 13 Jan 2025 · 2 repositories · arXiv:2501.07572
-
Better Prompt Compression Without Multi-Layer Perceptrons 12 Jan 2025 · 0 repositories · arXiv:2501.06730
-
DRDT3: Diffusion-Refined Decision Test-Time Training Model 12 Jan 2025 · 0 repositories · arXiv:2501.06718
-
Eliza: A Web3 friendly AI Agent Operating System 12 Jan 2025 · 2 repositories · arXiv:2501.06781