Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 36
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 36 of 190: papers 3,501 to 3,600 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Automated Spinal MRI Labelling from Reports Using a Large Language Model 22 Oct 2024 · 1 repository · arXiv:2410.17235
-
Distill-SynthKG: Distilling Knowledge Graph Synthesis Workflow for Improved Coverage and Efficiency 22 Oct 2024 · 0 repositories · arXiv:2410.16597
-
DNAHLM -- DNA sequence and Human Language mixed large language Model 22 Oct 2024 · 1 repository · arXiv:2410.16917
-
Exploring Possibilities of AI-Powered Legal Assistance in Bangladesh through Large Language Modeling 22 Oct 2024 · 1 repository · arXiv:2410.17210
-
Graph Transformers Dream of Electric Flow 22 Oct 2024 · 0 repositories · arXiv:2410.16699
-
In Context Learning and Reasoning for Symbolic Regression with Large Language Models 22 Oct 2024 · 1 repository · arXiv:2410.17448
-
Interchangeable Token Embeddings for Extendable Vocabulary and Alpha-Equivalence 22 Oct 2024 · 0 repositories · arXiv:2410.17161
-
LiNo: Advancing Recursive Residual Decomposition of Linear and Nonlinear Patterns for Robust Time Series Forecasting 22 Oct 2024 · 1 repository · arXiv:2410.17159
-
Methods of improving LLM training stability 22 Oct 2024 · 0 repositories · arXiv:2410.16682
-
Representation Shattering in Transformers: A Synthetic Study with Knowledge Editing 22 Oct 2024 · 0 repositories · arXiv:2410.17194Syntology 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Scattered Forest Search: Smarter Code Space Exploration with LLMs 22 Oct 2024 · 0 repositories · arXiv:2411.05010
-
SmartRAG: Jointly Learn RAG-Related Tasks From the Environment Feedback 22 Oct 2024 · 0 repositories · arXiv:2410.18141
-
An Efficient System for Automatic Map Storytelling -- A Case Study on Historical Maps 21 Oct 2024 · 1 repository · arXiv:2410.15780
-
An Explainable Contrastive-based Dilated Convolutional Network with Transformer for Pediatric Pneumonia Detection 21 Oct 2024 · 0 repositories · arXiv:2410.16143
-
Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count 21 Oct 2024 · 1 repository · arXiv:2410.15787
-
Beyond 2:4: exploring V:N:M sparsity for efficient transformer inference on GPUs 21 Oct 2024 · 0 repositories · arXiv:2410.16135
-
Building A Coding Assistant via the Retrieval-Augmented Language Model 21 Oct 2024 · 1 repository · arXiv:2410.16229
-
CausalGraph2LLM: Evaluating LLMs for Causal Queries 21 Oct 2024 · 1 repository · arXiv:2410.15939
-
Developing Retrieval Augmented Generation (RAG) based LLM Systems from PDFs: An Experience Report 21 Oct 2024 · 1 repository · arXiv:2410.15944
-
Diffusion Transformer Policy 21 Oct 2024 · 1 repository · arXiv:2410.15959
-
Disambiguating Monocular Reconstruction of 3D Clothed Human with Spatial-Temporal Transformer 21 Oct 2024 · 0 repositories · arXiv:2410.16337
-
Generalized Probabilistic Attention Mechanism in Transformers 21 Oct 2024 · 0 repositories · arXiv:2410.15578
-
Generalizing Motion Planners with Mixture of Experts for Autonomous Driving 21 Oct 2024 · 1 repository · arXiv:2410.15774Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Guardians of Discourse: Evaluating LLMs on Multilingual Offensive Language Detection 21 Oct 2024 · 0 repositories · arXiv:2410.15623
-
Improving Neuron-level Interpretability with White-box Language Models 21 Oct 2024 · 0 repositories · arXiv:2410.16443
-
Large Language Models in Computer Science Education: A Systematic Literature Review 21 Oct 2024 · 1 repository · arXiv:2410.16349
-
Leveraging Retrieval-Augmented Generation for Culturally Inclusive Hakka Chatbots: Design Insights and User Perceptions 21 Oct 2024 · 0 repositories · arXiv:2410.15572
-
Modelling Concurrent RTP Flows for End-to-end Predictions of QoS in Real Time Communications 21 Oct 2024 · 0 repositories · arXiv:2410.15846
-
On Creating an English-Thai Code-switched Machine Translation in Medical Domain 21 Oct 2024 · 1 repository · arXiv:2410.16221
-
RAG4ITOps: A Supervised Fine-Tunable and Comprehensive RAG Framework for IT Operations and Maintenance 21 Oct 2024 · 0 repositories · arXiv:2410.15805
-
Reflection-Bench: probing AI intelligence with reflection 21 Oct 2024 · 1 repository · arXiv:2410.16270
-
Students Rather Than Experts: A New AI For Education Pipeline To Model More Human-Like And Personalised Early Adolescences 21 Oct 2024 · 0 repositories · arXiv:2410.15701
-
Tokenization as Finite-State Transduction 21 Oct 2024 · 0 repositories · arXiv:2410.15696
-
Towards a Reliable Offline Personal AI Assistant for Long Duration Spaceflight 21 Oct 2024 · 0 repositories · arXiv:2410.16397
-
Using GPT Models for Qualitative and Quantitative News Analytics in the 2024 US Presidental Election Process 21 Oct 2024 · 0 repositories · arXiv:2410.15884
-
ViMoE: An Empirical Study of Designing Vision Mixture-of-Experts 21 Oct 2024 · 0 repositories · arXiv:2410.15732
-
Weighted Diversified Sampling for Efficient Data-Driven Single-Cell Gene-Gene Interaction Discovery 21 Oct 2024 · 0 repositories · arXiv:2410.15616
-
Who's Who: Large Language Models Meet Knowledge Conflicts in Practice 21 Oct 2024 · 1 repository · arXiv:2410.15737
-
YOLO11 and Vision Transformers based 3D Pose Estimation of Immature Green Fruits in Commercial Apple Orchards for Robotic Thinning 21 Oct 2024 · 0 repositories · arXiv:2410.19846
-
Advancing Gasoline Consumption Forecasting: A Novel Hybrid Model Integrating Transformers, LSTM, and CNN 20 Oct 2024 · 0 repositories · arXiv:2410.16336
-
Back to School: Translation Using Grammar Books 20 Oct 2024 · 1 repository · arXiv:2410.15263Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 6 harvested samples)
-
BRIEF: Bridging Retrieval and Inference for Multi-hop Reasoning via Compression 20 Oct 2024 · 1 repository · arXiv:2410.15277
-
Comparative Analysis of LSTM, GRU, and Transformer Models for Stock Price Prediction 20 Oct 2024 · 0 repositories · arXiv:2411.05790
-
Contextual Augmented Multi-Model Programming (CAMP): A Hybrid Local-Cloud Copilot Framework 20 Oct 2024 · 1 repository · arXiv:2410.15285
-
ConTReGen: Context-driven Tree-structured Retrieval for Open-domain Long-form Text Generation 20 Oct 2024 · 0 repositories · arXiv:2410.15511
-
Do RAG Systems Cover What Matters? Evaluating and Optimizing Responses with Sub-Question Coverage 20 Oct 2024 · 1 repository · arXiv:2410.15531
-
Does ChatGPT Have a Poetic Style? 20 Oct 2024 · 1 repository · arXiv:2410.15299
-
Evaluating Consistencies in LLM responses through a Semantic Clustering of Question Answering 20 Oct 2024 · 0 repositories · arXiv:2410.15440
-
Exploring Social Desirability Response Bias in Large Language Models: Evidence from GPT-4 Simulations 20 Oct 2024 · 0 repositories · arXiv:2410.15442
-
LTPNet Integration of Deep Learning and Environmental Decision Support Systems for Renewable Energy Demand Forecasting 20 Oct 2024 · 0 repositories · arXiv:2410.15286
-
MMDS: A Multimodal Medical Diagnosis System Integrating Image Analysis and Knowledge-based Departmental Consultation 20 Oct 2024 · 0 repositories · arXiv:2410.15403
-
SDP4Bit: Toward 4-bit Communication Quantization in Sharded Data Parallelism for LLM Training 20 Oct 2024 · 0 repositories · arXiv:2410.15526
-
SEA: State-Exchange Attention for High-Fidelity Physics Based Transformers 20 Oct 2024 · 1 repository · arXiv:2410.15495Syntology official (archive's flag): 11 ran · 11 ran (of which 7 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 1 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples)
-
Training Language Models to Critique With Multi-agent Feedback 20 Oct 2024 · 0 repositories · arXiv:2410.15287
-
Unveiling and Consulting Core Experts in Retrieval-Augmented MoE-based LLMs 20 Oct 2024 · 0 repositories · arXiv:2410.15438
-
When Machine Unlearning Meets Retrieval-Augmented Generation (RAG): Keep Secret or Forget Knowledge? 20 Oct 2024 · 0 repositories · arXiv:2410.15267
-
Accelerate Coastal Ocean Circulation Model with AI Surrogate 19 Oct 2024 · 0 repositories · arXiv:2410.14952
-
Bias Amplification: Language Models as Increasingly Biased Media 19 Oct 2024 · 0 repositories · arXiv:2410.15234
-
EPT-1.5 Technical Report 19 Oct 2024 · 0 repositories · arXiv:2410.15076
-
Evaluation Of P300 Speller Performance Using Large Language Models Along With Cross-Subject Training 19 Oct 2024 · 1 repository · arXiv:2410.15161
-
EViT-Unet: U-Net Like Efficient Vision Transformer for Medical Image Segmentation on Mobile and Edge Devices 19 Oct 2024 · 1 repository · arXiv:2410.15036
-
MCCoder: Streamlining Motion Control with LLM-Assisted Code Generation and Rigorous Verification 19 Oct 2024 · 1 repository · arXiv:2410.15154
-
SemiHVision: Enhancing Medical Multimodal Models with a Semi-Human Annotated Dataset and Fine-Tuned Instruction Generation 19 Oct 2024 · 1 repository · arXiv:2410.14948
-
Visual Navigation of Digital Libraries: Retrieval and Classification of Images in the National Library of Norway's Digitised Book Collection 19 Oct 2024 · 1 repository · arXiv:2410.14969
-
Backdoored Retrievers for Prompt Injection Attacks on Retrieval Augmented Generation of Large Language Models 18 Oct 2024 · 0 repositories · arXiv:2410.14479
-
CausalChat: Interactive Causal Model Development and Refinement Using Large Language Models 18 Oct 2024 · 0 repositories · arXiv:2410.14146
-
CELI: Controller-Embedded Language Model Interactions 18 Oct 2024 · 0 repositories · arXiv:2410.14627
-
DFlow: Diverse Dialogue Flow Simulation with Large Language Models 18 Oct 2024 · 0 repositories · arXiv:2410.14853
-
Flame quality monitoring of flare stack based on deep visual features 18 Oct 2024 · 0 repositories · arXiv:2410.19823
-
Good Parenting is all you need -- Multi-agentic LLM Hallucination Mitigation 18 Oct 2024 · 0 repositories · arXiv:2410.14262
-
Improving Vision Transformers by Overlapping Heads in Multi-Head Self-Attention 18 Oct 2024 · 0 repositories · arXiv:2410.14874
-
Mixed Attention Transformer Enhanced Channel Estimation for Extremely Large-Scale MIMO Systems 18 Oct 2024 · 0 repositories · arXiv:2410.14439
-
Novel Development of LLM Driven mCODE Data Model for Improved Clinical Trial Matching to Enable Standardization and Interoperability in Oncology Research 18 Oct 2024 · 0 repositories · arXiv:2410.19826
-
Optimizing Retrieval-Augmented Generation with Elasticsearch for Enhanced Question-Answering Systems 18 Oct 2024 · 0 repositories · arXiv:2410.14167
-
Paths-over-Graph: Knowledge Graph Empowered Large Language Model Reasoning 18 Oct 2024 · 1 repository · arXiv:2410.14211Syntology 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 12 harvested samples) · 2 pointer-only (licence)
-
ELOQ: Resources for Enhancing LLM Detection of Out-of-Scope Questions 18 Oct 2024 · 1 repository · arXiv:2410.14567
-
Real-time Fake News from Adversarial Feedback 18 Oct 2024 · 1 repository · arXiv:2410.14651
-
Rethinking Transformer for Long Contextual Histopathology Whole Slide Image Analysis 18 Oct 2024 · 1 repository · arXiv:2410.14195Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 1 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Self-Satisfied: An end-to-end framework for SAT generation and prediction 18 Oct 2024 · 0 repositories · arXiv:2410.14888
-
Supervised Chain of Thought 18 Oct 2024 · 0 repositories · arXiv:2410.14198
-
TimeSeriesExam: A time series understanding exam 18 Oct 2024 · 1 repository · arXiv:2410.14752Syntology 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Transfer Learning on Transformers for Building Energy Consumption Forecasting -- A Comparative Study 18 Oct 2024 · 0 repositories · arXiv:2410.14107
-
xPerT: Extended Persistence Transformer 18 Oct 2024 · 2 repositories · arXiv:2410.14193
-
How Does Knowledge Selection Help Retrieval Augmented Generation? 17 Oct 2024 · 0 repositories · arXiv:2410.13258
-
Adversarial Testing as a Tool for Interpretability: Length-based Overfitting of Elementary Functions in Transformers 17 Oct 2024 · 0 repositories · arXiv:2410.13802
-
Better to Ask in English: Evaluation of Large Language Models on English, Low-resource and Cross-Lingual Settings 17 Oct 2024 · 0 repositories · arXiv:2410.13153
-
Co-Segmentation without any Pixel-level Supervision with Application to Large-Scale Sketch Classification 17 Oct 2024 · 0 repositories · arXiv:2410.13582
-
D-FINE: Redefine Regression Task in DETRs as Fine-grained Distribution Refinement 17 Oct 2024 · 5 repositories · arXiv:2410.13842Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 14 harvested samples)
-
DurIAN-E 2: Duration Informed Attention Network with Adaptive Variational Autoencoder and Adversarial Learning for Expressive Text-to-Speech Synthesis 17 Oct 2024 · 0 repositories · arXiv:2410.13288
-
Enhancing Generalization in Sparse Mixture of Experts Models: The Case for Increased Expert Activation in Compositional Tasks 17 Oct 2024 · 0 repositories · arXiv:2410.13964
-
Evaluating Self-Generated Documents for Enhancing Retrieval-Augmented Generation with Large Language Models 17 Oct 2024 · 0 repositories · arXiv:2410.13192
-
FaithBench: A Diverse Hallucination Benchmark for Summarization by Modern LLMs 17 Oct 2024 · 2 repositories · arXiv:2410.13210
-
Help Me Identify: Is an LLM+VQA System All We Need to Identify Visual Concepts? 17 Oct 2024 · 1 repository · arXiv:2410.13651
-
Hiformer: Hybrid Frequency Feature Enhancement Inverted Transformer for Long-Term Wind Power Prediction 17 Oct 2024 · 1 repository · arXiv:2410.13303
-
Integrating Temporal Representations for Dynamic Memory Retrieval and Management in Large Language Models 17 Oct 2024 · 0 repositories · arXiv:2410.13553
-
IterSelectTune: An Iterative Training Framework for Efficient Instruction-Tuning Data Selection 17 Oct 2024 · 0 repositories · arXiv:2410.13464
-
Jailbreaking LLM-Controlled Robots 17 Oct 2024 · 0 repositories · arXiv:2410.13691
-
Learning Graph Quantized Tokenizers 17 Oct 2024 · 1 repository · arXiv:2410.13798Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 4 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples) · 3 pointer-only (licence)
-
Looking Inward: Language Models Can Learn About Themselves by Introspection 17 Oct 2024 · 1 repository · arXiv:2410.13787Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
MarineFormer: A Spatio-Temporal Attention Model for USV Navigation in Dynamic Marine Environments 17 Oct 2024 · 0 repositories · arXiv:2410.13973