Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 42
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 42 of 190: papers 4,101 to 4,200 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
DepthART: Monocular Depth Estimation as Autoregressive Refinement Task 23 Sep 2024 · 0 repositories · arXiv:2409.15010
-
Designing Pre-training Datasets from Unlabeled Data for EEG Classification with Transformers 23 Sep 2024 · 0 repositories · arXiv:2410.07190
-
Diffusion-based RGB-D Semantic Segmentation with Deformable Attention Transformer 23 Sep 2024 · 0 repositories · arXiv:2409.15117
-
Dual Stream Graph Transformer Fusion Networks for Enhanced Brain Decoding 23 Sep 2024 · 0 repositories · arXiv:2410.07189
-
EDGE-Rec: Efficient and Data-Guided Edge Diffusion For Recommender Systems Graphs 23 Sep 2024 · 0 repositories · arXiv:2409.14689
-
PAPILLON: Efficient and Stealthy Fuzz Testing-Powered Jailbreaks for LLMs 23 Sep 2024 · 1 repository · arXiv:2409.14866Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
Enhancing Scientific Reproducibility Through Automated BioCompute Object Creation Using Retrieval-Augmented Generation from Publications 23 Sep 2024 · 0 repositories · arXiv:2409.15076
-
GEM-RAG: Graphical Eigen Memories For Retrieval Augmented Generation 23 Sep 2024 · 0 repositories · arXiv:2409.15566
-
Generalizing monocular colonoscopy image depth estimation by uncertainty-based global and local fusion network 23 Sep 2024 · 0 repositories · arXiv:2409.15006
-
Generative AI Is Not Ready for Clinical Use in Patient Education for Lower Back Pain Patients, Even With Retrieval-Augmented Generation 23 Sep 2024 · 0 repositories · arXiv:2409.15260
-
HydroVision: LiDAR-Guided Hydrometric Prediction with Vision Transformers and Hybrid Graph Learning 23 Sep 2024 · 0 repositories · arXiv:2409.15213
-
Improving Academic Skills Assessment with NLP and Ensemble Learning 23 Sep 2024 · 0 repositories · arXiv:2409.19013
-
Kriformer: A Novel Spatiotemporal Kriging Approach Based on Graph Transformers 23 Sep 2024 · 0 repositories · arXiv:2409.14906
-
Learning When to Retrieve, What to Rewrite, and How to Respond in Conversational QA 23 Sep 2024 · 0 repositories · arXiv:2409.15515
-
Lessons Learned on Information Retrieval in Electronic Health Records: A Comparison of Embedding Models and Pooling Strategies 23 Sep 2024 · 0 repositories · arXiv:2409.15163
-
Location is Key: Leveraging Large Language Model for Functional Bug Localization in Verilog 23 Sep 2024 · 0 repositories · arXiv:2409.15186
-
M2OST: Many-to-one Regression for Predicting Spatial Transcriptomics from Digital Pathology Images 23 Sep 2024 · 1 repository · arXiv:2409.15092Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 4 harvested samples)
-
MemeCLIP: Leveraging CLIP Representations for Multimodal Meme Classification 23 Sep 2024 · 1 repository · arXiv:2409.14703
-
Micrometer: Micromechanics Transformer for Predicting Mechanical Responses of Heterogeneous Materials 23 Sep 2024 · 0 repositories · arXiv:2410.05281
-
PALLM: Evaluating and Enhancing PALLiative Care Conversations with Large Language Models 23 Sep 2024 · 1 repository · arXiv:2409.15188
-
Privacy Policy Analysis through Prompt Engineering for LLMs 23 Sep 2024 · 0 repositories · arXiv:2409.14879
-
RACER: Rich Language-Guided Failure Recovery Policies for Imitation Learning 23 Sep 2024 · 0 repositories · arXiv:2409.14674
-
Retrieval Augmented Generation (RAG) and Beyond: A Comprehensive Survey on How to Make your LLMs use External Data More Wisely 23 Sep 2024 · 0 repositories · arXiv:2409.14924
-
RoWSFormer: A Robust Watermarking Framework with Swin Transformer for Enhanced Geometric Attack Resilience 23 Sep 2024 · 0 repositories · arXiv:2409.14829
-
Safe Guard: an LLM-agent for Real-time Voice-based Hate Speech Detection in Social Virtual Reality 23 Sep 2024 · 0 repositories · arXiv:2409.15623
-
Scaling Laws of Decoder-Only Models on the Multilingual Machine Translation Task 23 Sep 2024 · 0 repositories · arXiv:2409.15051
-
SDBA: A Stealthy and Long-Lasting Durable Backdoor Attack in Federated Learning 23 Sep 2024 · 1 repository · arXiv:2409.14805
-
TransUKAN:Computing-Efficient Hybrid KAN-Transformer for Enhanced Medical Image Segmentation 23 Sep 2024 · 0 repositories · arXiv:2409.14676
-
Beyond Words: Evaluating Large Language Models in Transportation Planning 22 Sep 2024 · 0 repositories · arXiv:2409.14516
-
Can pre-trained language models generate titles for research papers? 22 Sep 2024 · 1 repository · arXiv:2409.14602
-
Enhancing LLM-based Autonomous Driving Agents to Mitigate Perception Attacks 22 Sep 2024 · 0 repositories · arXiv:2409.14488
-
Evaluating the Quality of Code Comments Generated by Large Language Models for Novice Programmers 22 Sep 2024 · 0 repositories · arXiv:2409.14368
-
Large Model Based Agents: State-of-the-Art, Cooperation Paradigms, Security and Privacy, and Future Trends 22 Sep 2024 · 0 repositories · arXiv:2409.14457
-
LLMs are One-Shot URL Classifiers and Explainers 22 Sep 2024 · 0 repositories · arXiv:2409.14306
-
More Effective LLM Compressed Tokens with Uniformly Spread Position Identifiers and Compression Loss 22 Sep 2024 · 0 repositories · arXiv:2409.14364
-
Patch Ranking: Efficient CLIP by Learning to Rank Local Patches 22 Sep 2024 · 1 repository · arXiv:2409.14607
-
Proof Automation with Large Language Models 22 Sep 2024 · 0 repositories · arXiv:2409.14274
-
Sparse Low-Ranked Self-Attention Transformer for Remaining Useful Lifetime Prediction of Optical Fiber Amplifiers 22 Sep 2024 · 0 repositories · arXiv:2409.14378
-
ChemEval: A Comprehensive Multi-Level Chemical Evaluation for Large Language Models 21 Sep 2024 · 1 repository · arXiv:2409.13989Syntology official (archive's flag): 17 ran · 17 ran (of which 0 constructed an object rather than computing a result; 17 with no instrument failure: 0 honoured, 0 violated, 17 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 18 harvested samples) · 18 pointer-only (licence)
-
SMART-RAG: Selection using Determinantal Matrices for Augmented Retrieval 21 Sep 2024 · 0 repositories · arXiv:2409.13992
-
Can LLMs replace Neil deGrasse Tyson? Evaluating the Reliability of LLMs as Science Communicators 21 Sep 2024 · 1 repository · arXiv:2409.14037
-
QMOS: Enhancing LLMs for Telecommunication with Question Masked loss and Option Shuffling 21 Sep 2024 · 1 repository · arXiv:2409.14175
-
AI Assistants for Spaceflight Procedures: Combining Generative Pre-Trained Transformer and Retrieval-Augmented Generation on Knowledge Graphs With Augmented Reality Cues 21 Sep 2024 · 0 repositories · arXiv:2409.14206
-
Developing a Thailand solar irradiance map using Himawari-8 satellite imageries and deep learning models 21 Sep 2024 · 1 repository · arXiv:2409.16320
-
Drift to Remember 21 Sep 2024 · 0 repositories · arXiv:2409.13997
-
Knowledge in Triples for LLMs: Enhancing Table QA Accuracy with Semantic Extraction 21 Sep 2024 · 0 repositories · arXiv:2409.14192
-
Loop Neural Networks for Parameter Sharing 21 Sep 2024 · 0 repositories · arXiv:2409.14199
-
Window-based Channel Attention for Wavelet-enhanced Learned Image Compression 21 Sep 2024 · 0 repositories · arXiv:2409.14090
-
Contextual Compression in Retrieval-Augmented Generation for Large Language Models: A Survey 20 Sep 2024 · 1 repository · arXiv:2409.13385
-
A Personalised 3D+t Mesh Generative Model for Unveiling Normal Heart Dynamics 20 Sep 2024 · 1 repository · arXiv:2409.13825
-
Aligning Language Models Using Follow-up Likelihood as Reward Signal 20 Sep 2024 · 1 repository · arXiv:2409.13948
-
AVG-LLaVA: A Large Multimodal Model with Adaptive Visual Granularity 20 Sep 2024 · 1 repository · arXiv:2410.02745Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 1 violated, 3 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples) · 1 pointer-only (licence)
-
EMMeTT: Efficient Multimodal Machine Translation Training 20 Sep 2024 · 0 repositories · arXiv:2409.13523
-
Enhancing Large Language Models with Domain-specific Retrieval Augment Generation: A Case Study on Long-form Consumer Health Question Answering in Ophthalmology 20 Sep 2024 · 0 repositories · arXiv:2409.13902
-
FAIR GPT: A virtual consultant for research data management in ChatGPT 20 Sep 2024 · 1 repository · arXiv:2410.07108
-
HUT: A More Computation Efficient Fine-Tuning Method With Hadamard Updated Transformation 20 Sep 2024 · 0 repositories · arXiv:2409.13501
-
Leveraging Knowledge Graphs and LLMs to Support and Monitor Legislative Systems 20 Sep 2024 · 0 repositories · arXiv:2409.13252
-
Localized Gaussians as Self-Attention Weights for Point Clouds Correspondence 20 Sep 2024 · 0 repositories · arXiv:2409.13291
-
Neural-Symbolic Collaborative Distillation: Advancing Small Language Models for Complex Reasoning Tasks 20 Sep 2024 · 1 repository · arXiv:2409.13203Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Prompting Large Language Models for Supporting the Differential Diagnosis of Anemia 20 Sep 2024 · 0 repositories · arXiv:2409.15377
-
ShizishanGPT: An Agricultural Large Language Model Integrating Tools and Resources 20 Sep 2024 · 1 repository · arXiv:2409.13537
-
STOP! Benchmarking Large Language Models with Sensitivity Testing on Offensive Progressions 20 Sep 2024 · 1 repository · arXiv:2409.13843
-
Tackling fluffy clouds: field boundaries detection using time series of S2 and/or S1 imagery 20 Sep 2024 · 1 repository · arXiv:2409.13568
-
TalkMosaic: Interactive PhotoMosaic with Multi-modal LLM Q&A Interactions 20 Sep 2024 · 0 repositories · arXiv:2409.13941
-
ViTGuard: Attention-aware Detection against Adversarial Examples for Vision Transformer 20 Sep 2024 · 0 repositories · arXiv:2409.13828
-
3DTopia-XL: Scaling High-quality 3D Asset Generation via Primitive Diffusion 19 Sep 2024 · 1 repository · arXiv:2409.12957Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
A Novel Perspective for Multi-modal Multi-label Skin Lesion Classification 19 Sep 2024 · 0 repositories · arXiv:2409.12390
-
End-to-end Open-vocabulary Video Visual Relationship Detection using Multi-modal Prompting 19 Sep 2024 · 0 repositories · arXiv:2409.12499
-
Enhancing E-commerce Product Title Translation with Retrieval-Augmented Generation and Large Language Models 19 Sep 2024 · 0 repositories · arXiv:2409.12880
-
Enhancing TinyBERT for Financial Sentiment Analysis Using GPT-Augmented FinBERT Distillation 19 Sep 2024 · 1 repository · arXiv:2409.18999
-
Evaluating Image Hallucination in Text-to-Image Generation with Question-Answering 19 Sep 2024 · 1 repository · arXiv:2409.12784
-
Exploring Large Language Models for Product Attribute Value Identification 19 Sep 2024 · 0 repositories · arXiv:2409.12695
-
Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation 19 Sep 2024 · 2 repositories · arXiv:2409.12941
-
LMT-Net: Lane Model Transformer Network for Automated HD Mapping from Sparse Vehicle Observations 19 Sep 2024 · 0 repositories · arXiv:2409.12409
-
MURI: High-Quality Instruction Tuning Datasets for Low-Resource Languages via Reverse Instructions 19 Sep 2024 · 1 repository · arXiv:2409.12958
-
On the Effectiveness of LLMs for Manual Test Verifications 19 Sep 2024 · 0 repositories · arXiv:2409.12405
-
Prompts Are Programs Too! Understanding How Developers Build Software Containing Prompts 19 Sep 2024 · 0 repositories · arXiv:2409.12447
-
Retrieval-Augmented Test Generation: How Far Are We? 19 Sep 2024 · 0 repositories · arXiv:2409.12682
-
Should RAG Chatbots Forget Unimportant Conversations? Exploring Importance and Forgetting with Psychological Insights 19 Sep 2024 · 1 repository · arXiv:2409.12524
-
Small Language Models are Equation Reasoners 19 Sep 2024 · 0 repositories · arXiv:2409.12393
-
TACO-RL: Task Aware Prompt Compression Optimization with Reinforcement Learning 19 Sep 2024 · 0 repositories · arXiv:2409.13035
-
What Would You Ask When You First Saw a²+b²=c²? Evaluating LLM on Curiosity-Driven Questioning 19 Sep 2024 · 0 repositories · arXiv:2409.17172
-
Data Efficient Acoustic Scene Classification using Teacher-Informed Confusing Class Instruction 18 Sep 2024 · 0 repositories · arXiv:2409.11964
-
DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech 18 Sep 2024 · 0 repositories · arXiv:2409.11835
-
DynaMo: In-Domain Dynamics Pretraining for Visuo-Motor Control 18 Sep 2024 · 0 repositories · arXiv:2409.12192
-
Extract-and-Abstract: Unifying Extractive and Abstractive Summarization within Single Encoder-Decoder Framework 18 Sep 2024 · 0 repositories · arXiv:2409.11827
-
From Lists to Emojis: How Format Bias Affects Model Alignment 18 Sep 2024 · 0 repositories · arXiv:2409.11704
-
Harnessing LLMs for API Interactions: A Framework for Classification and Synthetic Data Generation 18 Sep 2024 · 0 repositories · arXiv:2409.11703
-
MAgICoRe: Multi-Agent, Iterative, Coarse-to-Fine Refinement for Reasoning 18 Sep 2024 · 1 repository · arXiv:2409.12147Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
NT-ViT: Neural Transcoding Vision Transformers for EEG-to-fMRI Synthesis 18 Sep 2024 · 0 repositories · arXiv:2409.11836
-
Recommendation with Generative Models 18 Sep 2024 · 0 repositories · arXiv:2409.15173
-
Reinforcement Learning as an Improvement Heuristic for Real-World Production Scheduling 18 Sep 2024 · 0 repositories · arXiv:2409.11933
-
TART: An Open-Source Tool-Augmented Framework for Explainable Table-based Reasoning 18 Sep 2024 · 1 repository · arXiv:2409.11724
-
Unsupervised Feature Orthogonalization for Learning Distortion-Invariant Representations 18 Sep 2024 · 1 repository · arXiv:2409.12276
-
VERA: Validation and Enhancement for Retrieval Augmented systems 18 Sep 2024 · 0 repositories · arXiv:2409.15364
-
A Unified Framework to Classify Business Activities into International Standard Industrial Classification through Large Language Models for Circular Economy 17 Sep 2024 · 0 repositories · arXiv:2409.18988
-
American Sign Language to Text Translation using Transformer and Seq2Seq with LSTM 17 Sep 2024 · 0 repositories · arXiv:2409.10874
-
Chain-of-Thought Prompting for Speech Translation 17 Sep 2024 · 0 repositories · arXiv:2409.11538
-
Contrasformer: A Brain Network Contrastive Transformer for Neurodegenerative Condition Identification 17 Sep 2024 · 1 repository · arXiv:2409.10944
-
EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer 17 Sep 2024 · 0 repositories · arXiv:2409.10819