Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 74
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 74 of 190: papers 7,301 to 7,400 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Empirical Studies of Parameter Efficient Methods for Large Language Models of Code and Knowledge Transfer to R 16 Mar 2024 · 1 repository · arXiv:2405.01553
-
From Melting Pots to Misrepresentations: Exploring Harms in Generative AI 16 Mar 2024 · 0 repositories · arXiv:2403.10776
-
Can Large Language Models Solve Robot Routing? 16 Mar 2024 · 1 repository · arXiv:2403.10795
-
Large language model-powered chatbots for internationalizing student support in higher education 16 Mar 2024 · 0 repositories · arXiv:2403.14702
-
Pointer-Generator Networks for Low-Resource Machine Translation: Don't Copy That! 16 Mar 2024 · 1 repository · arXiv:2403.10963
-
RetMIL: Retentive Multiple Instance Learning for Histopathological Whole Slide Image Classification 16 Mar 2024 · 0 repositories · arXiv:2403.10858
-
Time Series Representation Learning with Supervised Contrastive Temporal Transformer 16 Mar 2024 · 1 repository · arXiv:2403.10787
-
Twin Transformer using Gated Dynamic Learnable Attention mechanism for Fault Detection and Diagnosis in the Tennessee Eastman Process 16 Mar 2024 · 0 repositories · arXiv:2403.10842
-
Understanding Robustness of Visual State Space Models for Image Classification 16 Mar 2024 · 1 repository · arXiv:2403.10935
-
A Thorough Comparison of Cross-Encoders and LLMs for Reranking SPLADE 15 Mar 2024 · 0 repositories · arXiv:2403.10407
-
Application of GPT Language Models for Innovation in Activities in University Teaching 15 Mar 2024 · 0 repositories · arXiv:2403.14694
-
Attention-Enhanced Hybrid Feature Aggregation Network for 3D Brain Tumor Segmentation 15 Mar 2024 · 1 repository · arXiv:2403.09942
-
DRAGIN: Dynamic Retrieval Augmented Generation based on the Information Needs of Large Language Models 15 Mar 2024 · 1 repository · arXiv:2403.10081Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Enhancing LLM Factual Accuracy with RAG to Counter Hallucinations: A Case Study on Domain-Specific Queries in Private Knowledge-Bases 15 Mar 2024 · 1 repository · arXiv:2403.10446Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 13 harvested samples)
-
ExeGPT: Constraint-Aware Resource Scheduling for LLM Inference 15 Mar 2024 · 0 repositories · arXiv:2404.07947
-
FBPT: A Fully Binary Point Transformer 15 Mar 2024 · 0 repositories · arXiv:2403.09998
-
Generative Region-Language Pretraining for Open-Ended Object Detection 15 Mar 2024 · 1 repository · arXiv:2403.10191Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 11 harvested samples) · 4 pointer-only (licence)
-
How Powerful Potential of Attention on Image Restoration? 15 Mar 2024 · 0 repositories · arXiv:2403.10336
-
Knowledge Condensation and Reasoning for Knowledge-based VQA 15 Mar 2024 · 0 repositories · arXiv:2403.10037
-
Magic Tokens: Select Diverse Tokens for Multi-modal Object Re-Identification 15 Mar 2024 · 2 repositories · arXiv:2403.10254Syntology official (archive's flag): 3 ran · 9 ran (of which 5 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples) · 9 pointer-only (licence)
-
MEDPNet: Achieving High-Precision Adaptive Registration for Complex Die Castings 15 Mar 2024 · 0 repositories · arXiv:2403.09996
-
PASTA: Towards Flexible and Efficient HDR Imaging Via Progressively Aggregated Spatio-Temporal Alignment 15 Mar 2024 · 0 repositories · arXiv:2403.10376
-
RAFT: Adapting Language Model to Domain Specific RAG 15 Mar 2024 · 1 repository · arXiv:2403.10131
-
Repoformer: Selective Retrieval for Repository-Level Code Completion 15 Mar 2024 · 0 repositories · arXiv:2403.10059
-
Rough Transformers for Continuous and Efficient Time-Series Modelling 15 Mar 2024 · 0 repositories · arXiv:2403.10288
-
SparseFusion: Efficient Sparse Multi-Modal Fusion Framework for Long-Range 3D Perception 15 Mar 2024 · 0 repositories · arXiv:2403.10036
-
SwinMTL: A Shared Architecture for Simultaneous Depth Estimation and Semantic Segmentation from Monocular Camera Images 15 Mar 2024 · 1 repository · arXiv:2403.10662
-
ViTCN: Vision Transformer Contrastive Network For Reasoning 15 Mar 2024 · 0 repositories · arXiv:2403.09962
-
A Continued Pretrained LLM Approach for Automatic Medical Note Generation 14 Mar 2024 · 0 repositories · arXiv:2403.09057
-
AI on AI: Exploring the Utility of GPT as an Expert Annotator of AI Publications 14 Mar 2024 · 0 repositories · arXiv:2403.09097
-
AraTrust: An Evaluation of Trustworthiness for LLMs in Arabic 14 Mar 2024 · 0 repositories · arXiv:2403.09017
-
Basque and Spanish Counter Narrative Generation: Data Creation and Evaluation 14 Mar 2024 · 0 repositories · arXiv:2403.09159
-
Circuit Transformer: A Transformer That Preserves Logical Equivalence 14 Mar 2024 · 1 repository · arXiv:2403.13838Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
CodeUltraFeedback: An LLM-as-a-Judge Dataset for Aligning Large Language Models to Coding Preferences 14 Mar 2024 · 2 repositories · arXiv:2403.09032Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Evaluating LLMs for Gender Disparities in Notable Persons 14 Mar 2024 · 0 repositories · arXiv:2403.09148
-
Fisher Mask Nodes for Language Model Merging 14 Mar 2024 · 1 repository · arXiv:2403.09891
-
GiT: Towards Generalist Vision Transformer through Universal Language Interface 14 Mar 2024 · 1 repository · arXiv:2403.09394
-
Information Extraction: An application to the domain of hyper-local financial data on developing countries 14 Mar 2024 · 0 repositories · arXiv:2403.09077
-
Komodo: A Linguistic Expedition into Indonesia's Regional Languages 14 Mar 2024 · 0 repositories · arXiv:2403.09362
-
LAMP: A Language Model on the Map 14 Mar 2024 · 1 repository · arXiv:2403.09059
-
Leap: molecular synthesisability scoring with intermediates 14 Mar 2024 · 0 repositories · arXiv:2403.13005
-
Optimistic Verifiable Training by Controlling Hardware Nondeterminism 14 Mar 2024 · 1 repository · arXiv:2403.09603Syntology official (archive's flag): 9 ran · 9 ran (of which 3 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 5 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
RAGGED: Towards Informed Design of Retrieval Augmented Generation Systems 14 Mar 2024 · 1 repository · arXiv:2403.09040Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 14 harvested samples) · 1 pointer-only (licence)
-
Rectifying Demonstration Shortcut in In-Context Learning 14 Mar 2024 · 1 repository · arXiv:2403.09488Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 6 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Retrieval augmented text-to-SQL generation for epidemiological question answering using electronic health records 14 Mar 2024 · 1 repository · arXiv:2403.09226
-
Sabiá-2: A New Generation of Portuguese Large Language Models 14 Mar 2024 · 0 repositories · arXiv:2403.09887
-
SkateFormer: Skeletal-Temporal Transformer for Human Action Recognition 14 Mar 2024 · 1 repository · arXiv:2403.09508Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Switch Diffusion Transformer: Synergizing Denoising Tasks with Sparse Mixture-of-Experts 14 Mar 2024 · 2 repositories · arXiv:2403.09176Syntology official (archive's flag): 15 ran · 15 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 3 honoured, 0 violated, 6 with no contract checked; 6 where Syntology's instrument failed) · 3 unverified (of 18 harvested samples) · 6 pointer-only (licence)
-
VisionGPT-3D: A Generalized Multimodal Agent for Enhanced 3D Vision Understanding 14 Mar 2024 · 0 repositories · arXiv:2403.09530
-
VM-UNET-V2 Rethinking Vision Mamba UNet for Medical Image Segmentation 14 Mar 2024 · 1 repository · arXiv:2403.09157
-
A Multimodal Fusion Network For Student Emotion Recognition Based on Transformer and Tensor Product 13 Mar 2024 · 0 repositories · arXiv:2403.08511
-
Authorship Verification based on the Likelihood Ratio of Grammar Models 13 Mar 2024 · 0 repositories · arXiv:2403.08462
-
Autoregressive Score Generation for Multi-trait Essay Scoring 13 Mar 2024 · 1 repository · arXiv:2403.08332
-
Can Large Language Models Identify Authorship? 13 Mar 2024 · 1 repository · arXiv:2403.08213
-
CAMSIC: Content-aware Masked Image Modeling Transformer for Stereo Image Compression 13 Mar 2024 · 1 repository · arXiv:2403.08505
-
Prompting Large Language Models to Tackle the Full Software Development Lifecycle: A Case Study 13 Mar 2024 · 2 repositories · arXiv:2403.08604Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 3 pointer-only (licence)
-
Distilling Named Entity Recognition Models for Endangered Species from Large Language Models 13 Mar 2024 · 0 repositories · arXiv:2403.15430
-
Do Language Models Care About Text Quality? Evaluating Web-Crawled Corpora Across 11 Languages 13 Mar 2024 · 0 repositories · arXiv:2403.08693
-
Embedded Translations for Low-resource Automated Glossing 13 Mar 2024 · 0 repositories · arXiv:2403.08189
-
Evaluating the Application of Large Language Models to Generate Feedback in Programming Education 13 Mar 2024 · 0 repositories · arXiv:2403.09744
-
Generative Pretrained Structured Transformers: Unsupervised Syntactic Language Models at Scale 13 Mar 2024 · 2 repositories · arXiv:2403.08293Syntology official (archive's flag): 2 ran · 4 ran (of which 2 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples)
-
GPT, Ontology, and CAABAC: A Tripartite Personalized Access Control Model Anchored by Compliance, Context and Attribute 13 Mar 2024 · 0 repositories · arXiv:2403.08264
-
Large Language Models are Contrastive Reasoners 13 Mar 2024 · 1 repository · arXiv:2403.08211
-
OneVOS: Unifying Video Object Segmentation with All-in-One Transformer Framework 13 Mar 2024 · 0 repositories · arXiv:2403.08682
-
P2LHAP:Wearable sensor-based human activity recognition, segmentation and forecast through Patch-to-Label Seq2Seq Transformer 13 Mar 2024 · 0 repositories · arXiv:2403.08214
-
The Garden of Forking Paths: Observing Dynamic Parameters Distribution in Large Language Models 13 Mar 2024 · 0 repositories · arXiv:2403.08739
-
SAP: Corrective Machine Unlearning with Scaled Activation Projection for Label Noise Robustness 13 Mar 2024 · 1 repository · arXiv:2403.08618
-
ViT-CoMer: Vision Transformer with Convolutional Multi-scale Feature Interaction for Dense Predictions 13 Mar 2024 · 2 repositories
-
A Survey of Vision Transformers in Autonomous Driving: Current Trends and Future Directions 12 Mar 2024 · 0 repositories · arXiv:2403.07542
-
Chronos: Learning the Language of Time Series 12 Mar 2024 · 6 repositories · arXiv:2403.07815Syntology official (archive's flag): 8 ran · 23 ran (of which 0 constructed an object rather than computing a result; 22 with no instrument failure: 3 honoured, 1 violated, 18 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (of 28 harvested samples) · 5 pointer-only (licence)
-
Contextual Clarity: Generating Sentences with Transformer Models using Context-Reverso Data 12 Mar 2024 · 1 repository · arXiv:2403.08103
-
Decomposing Disease Descriptions for Enhanced Pathology Detection: A Multi-Aspect Vision-Language Pre-training Framework 12 Mar 2024 · 2 repositories · arXiv:2403.07636Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion 12 Mar 2024 · 2 repositories · arXiv:2403.07865Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
GPT-generated Text Detection: Benchmark Dataset and Tensor-based Detection Method 12 Mar 2024 · 1 repository · arXiv:2403.07321Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
In-context learning enables multimodal large language models to classify cancer pathology images 12 Mar 2024 · 0 repositories · arXiv:2403.07407
-
Investigating the performance of Retrieval-Augmented Generation and fine-tuning for the development of AI-driven knowledge-based systems 12 Mar 2024 · 1 repository · arXiv:2403.09727
-
Learning Correction Errors via Frequency-Self Attention for Blind Image Super-Resolution 12 Mar 2024 · 0 repositories · arXiv:2403.07390
-
MoralBERT: A Fine-Tuned Language Model for Capturing Moral Values in Social Discussions 12 Mar 2024 · 1 repository · arXiv:2403.07678
-
Q-SLAM: Quadric Representations for Monocular SLAM 12 Mar 2024 · 0 repositories · arXiv:2403.08125
-
Reinforced Sequential Decision-Making for Sepsis Treatment: The POSNEGDM Framework with Mortality Classifier and Transformer 12 Mar 2024 · 1 repository · arXiv:2403.07309Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 2 pointer-only (licence)
-
Rethinking ASTE: A Minimalist Tagging Scheme Alongside Contrastive Learning 12 Mar 2024 · 0 repositories · arXiv:2403.07342
-
Rethinking Generative Large Language Model Evaluation for Semantic Comprehension 12 Mar 2024 · 0 repositories · arXiv:2403.07872
-
SIFiD: Reassess Summary Factual Inconsistency Detection with LLM 12 Mar 2024 · 0 repositories · arXiv:2403.07557
-
StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models 12 Mar 2024 · 4 repositories · arXiv:2403.07714Syntology official (archive's flag): 3 ran · 11 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 1 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 15 harvested samples) · 5 pointer-only (licence)
-
Stress index strategy enhanced with financial news sentiment analysis for the equity markets 12 Mar 2024 · 0 repositories · arXiv:2404.00012
-
Textual Knowledge Matters: Cross-Modality Co-Teaching for Generalized Visual Class Discovery 12 Mar 2024 · 1 repository · arXiv:2403.07369Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 3 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
The future of document indexing: GPT and Donut revolutionize table of content processing 12 Mar 2024 · 0 repositories · arXiv:2403.07553
-
Towards a clinically accessible radiology foundation model: open-access and lightweight, with automated evaluation 12 Mar 2024 · 1 repository · arXiv:2403.08002Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Unleashing HyDRa: Hybrid Fusion, Depth Consistency and Radar for Unified 3D Perception 12 Mar 2024 · 1 repository · arXiv:2403.07746
-
ViT-CoMer: Vision Transformer with Convolutional Multi-scale Feature Interaction for Dense Predictions 12 Mar 2024 · 1 repository · arXiv:2403.07392Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
A multi-cohort study on prediction of acute brain dysfunction states using selective state space models 11 Mar 2024 · 0 repositories · arXiv:2403.07201
-
Amharic LLaMA and LLaVA: Multimodal LLMs for Low Resource Languages 11 Mar 2024 · 1 repository · arXiv:2403.06354
-
Development of a Reliable and Accessible Caregiving Language Model (CaLM) 11 Mar 2024 · 0 repositories · arXiv:2403.06857
-
GRITv2: Efficient and Light-weight Social Relation Recognition 11 Mar 2024 · 0 repositories · arXiv:2403.06895
-
Guiding Clinical Reasoning with Large Language Models via Knowledge Seeds 11 Mar 2024 · 0 repositories · arXiv:2403.06609
-
HDRTransDC: High Dynamic Range Image Reconstruction with Transformer Deformation Convolution 11 Mar 2024 · 0 repositories · arXiv:2403.06831
-
Hybrid Human-LLM Corpus Construction and LLM Evaluation for Rare Linguistic Phenomena 11 Mar 2024 · 0 repositories · arXiv:2403.06965
-
In-context Exploration-Exploitation for Reinforcement Learning 11 Mar 2024 · 0 repositories · arXiv:2403.06826
-
Multi-Scale Implicit Transformer with Re-parameterize for Arbitrary-Scale Super-Resolution 11 Mar 2024 · 0 repositories · arXiv:2403.06536
-
Multilingual Turn-taking Prediction Using Voice Activity Projection 11 Mar 2024 · 0 repositories · arXiv:2403.06487