Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 111
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 111 of 190: papers 11,001 to 11,100 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
ClipSitu: Effectively Leveraging CLIP for Conditional Predictions in Situation Recognition 2 Jul 2023 · 1 repository · arXiv:2307.00586
-
Conformer LLMs -- Convolution Augmented Large Language Models 2 Jul 2023 · 0 repositories · arXiv:2307.00461
-
Bidirectional Correlation-Driven Inter-Frame Interaction Transformer for Referring Video Object Segmentation 2 Jul 2023 · 0 repositories · arXiv:2307.00536
-
TensorGPT: Efficient Compression of Large Language Models based on Tensor-Train Decomposition 2 Jul 2023 · 0 repositories · arXiv:2307.00526
-
AutoST: Training-free Neural Architecture Search for Spiking Transformers 1 Jul 2023 · 1 repository · arXiv:2307.00293
-
Effective Matching of Patients to Clinical Trials using Entity Extraction and Neural Re-ranking 1 Jul 2023 · 0 repositories · arXiv:2307.00381
-
Rearrangement Planning for General Part Assembly 1 Jul 2023 · 0 repositories · arXiv:2307.00206
-
How far is Language Model from 100% Few-shot Named Entity Recognition in Medical Domain 1 Jul 2023 · 1 repository · arXiv:2307.00186
-
Learning Content-enhanced Mask Transformer for Domain Generalized Urban-Scene Segmentation 1 Jul 2023 · 1 repository · arXiv:2307.00371
-
More for Less: Compact Convolutional Transformers Enable Robust Medical Image Classification with Limited Data 1 Jul 2023 · 0 repositories · arXiv:2307.00213
-
PM-DETR: Domain Adaptive Prompt Memory for Object Detection with Transformers 1 Jul 2023 · 0 repositories · arXiv:2307.00313
-
Spatial-Temporal Graph Enhanced DETR Towards Multi-Frame 3D Object Detection 1 Jul 2023 · 1 repository · arXiv:2307.00347Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples)
-
Act3D: 3D Feature Field Transformers for Multi-Task Robotic Manipulation 30 Jun 2023 · 2 repositories · arXiv:2306.17817
-
Harnessing LLMs in Curricular Design: Using GPT-4 to Support Authoring of Learning Objectives 30 Jun 2023 · 0 repositories · arXiv:2306.17459
-
HVTSurv: Hierarchical Vision Transformer for Patient-Level Survival Prediction from Whole Slide Image 30 Jun 2023 · 1 repository · arXiv:2306.17373Syntology official (archive's flag): 4 ran · 4 ran (of which 3 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
Large Language Models are Effective Text Rankers with Pairwise Ranking Prompting 30 Jun 2023 · 0 repositories · arXiv:2306.17563
-
Large Language Models (GPT) for automating feedback on programming assignments 30 Jun 2023 · 0 repositories · arXiv:2307.00150
-
Learning to Localize with Attention: from sparse mmWave channel estimates from a single BS to high accuracy 3D location 30 Jun 2023 · 0 repositories · arXiv:2307.00167
-
Meta-Reasoning: Semantics-Symbol Deconstruction for Large Language Models 30 Jun 2023 · 1 repository · arXiv:2306.17820
-
Preference Ranking Optimization for Human Alignment 30 Jun 2023 · 1 repository · arXiv:2306.17492
-
SPAE: Semantic Pyramid AutoEncoder for Multimodal Generation with Frozen LLMs 30 Jun 2023 · 0 repositories · arXiv:2306.17842
-
SpATr: MoCap 3D Human Action Recognition based on Spiral Auto-encoder and Transformer Network 30 Jun 2023 · 1 repository · arXiv:2306.17574
-
Stay on topic with Classifier-Free Guidance 30 Jun 2023 · 0 repositories · arXiv:2306.17806
-
SummQA at MEDIQA-Chat 2023:In-Context Learning with GPT-4 for Medical Summarization 30 Jun 2023 · 1 repository · arXiv:2306.17384
-
The Shaped Transformer: Attention Models in the Infinite Depth-and-Width Limit 30 Jun 2023 · 0 repositories · arXiv:2306.17759
-
Towards Improving the Performance of Pre-Trained Speech Models for Low-Resource Languages Through Lateral Inhibition 30 Jun 2023 · 0 repositories · arXiv:2306.17792
-
Transformers in Healthcare: A Survey 30 Jun 2023 · 0 repositories · arXiv:2307.00067
-
A Formal Perspective on Byte-Pair Encoding 29 Jun 2023 · 1 repository · arXiv:2306.16837
-
A negation detection assessment of GPTs: analysis with the xNot360 dataset 29 Jun 2023 · 0 repositories · arXiv:2306.16638
-
Answer Mining from a Pool of Images: Towards Retrieval-Based Visual Question Answering 29 Jun 2023 · 1 repository · arXiv:2306.16713Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples) · 6 pointer-only (licence)
-
Benchmarking Large Language Model Capabilities for Conditional Generation 29 Jun 2023 · 0 repositories · arXiv:2306.16793
-
CMATH: Can Your Language Model Pass Chinese Elementary School Math Test? 29 Jun 2023 · 0 repositories · arXiv:2306.16636
-
Generative AI for Programming Education: Benchmarking ChatGPT, GPT-4, and Human Tutors 29 Jun 2023 · 0 repositories · arXiv:2306.17156
-
LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding 29 Jun 2023 · 2 repositories · arXiv:2306.17107Syntology official (archive's flag): 1 ran · 6 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 6 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 2 pointer-only (licence)
-
LyricWhiz: Robust Multilingual Zero-shot Lyrics Transcription by Whispering to ChatGPT 29 Jun 2023 · 1 repository · arXiv:2306.17103
-
MNISQ: A Large-Scale Quantum Circuit Dataset for Machine Learning on/for Quantum Computers in the NISQ era 29 Jun 2023 · 1 repository · arXiv:2306.16627
-
Multi-source Semantic Graph-based Multimodal Sarcasm Explanation Generation 29 Jun 2023 · 1 repository · arXiv:2306.16650
-
UMASS_BioNLP at MEDIQA-Chat 2023: Can LLMs generate high-quality synthetic note-oriented doctor-patient conversations? 29 Jun 2023 · 1 repository · arXiv:2306.16931Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Pareto Optimal Learning for Estimating Large Language Model Errors 28 Jun 2023 · 0 repositories · arXiv:2306.16564
-
Chatlaw: A Multi-Agent Collaborative Legal Assistant with Knowledge Graph Enhanced Mixture-of-Experts Large Language Model 28 Jun 2023 · 1 repository · arXiv:2306.16092
-
Inferring the Goals of Communicating Agents from Actions and Instructions 28 Jun 2023 · 0 repositories · arXiv:2306.16207
-
Is ChatGPT a Biomedical Expert? -- Exploring the Zero-Shot Performance of Current GPT Models in Biomedical Tasks 28 Jun 2023 · 1 repository · arXiv:2306.16108
-
Leveraging GPT-4 for Food Effect Summarization to Enhance Product-Specific Guidance Development via Iterative Prompting 28 Jun 2023 · 0 repositories · arXiv:2306.16275
-
Mass Spectra Prediction with Structural Motif-based Graph Neural Networks 28 Jun 2023 · 0 repositories · arXiv:2306.16085
-
C²Former: Calibrated and Complementary Transformer for RGB-Infrared Object Detection 28 Jun 2023 · 2 repositories · arXiv:2306.16175
-
SkillNet-X: A Multilingual Multitask Model with Sparsely Activated Skills 28 Jun 2023 · 0 repositories · arXiv:2306.16176
-
Taqyim: Evaluating Arabic NLP Tasks Using ChatGPT Models 28 Jun 2023 · 1 repository · arXiv:2306.16322
-
The 2nd Place Solution for 2023 Waymo Open Sim Agents Challenge 28 Jun 2023 · 0 repositories · arXiv:2306.15914
-
CellViT: Vision Transformers for Precise Cell Segmentation and Classification 27 Jun 2023 · 3 repositories · arXiv:2306.15350Syntology community repositories only · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Evaluating GPT-3.5 and GPT-4 on Grammatical Error Correction for Brazilian Portuguese 27 Jun 2023 · 0 repositories · arXiv:2306.15788
-
FedET: A Communication-Efficient Federated Class-Incremental Learning Framework Based on Enhanced Transformer 27 Jun 2023 · 0 repositories · arXiv:2306.15347
-
HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide Resolution 27 Jun 2023 · 4 repositories · arXiv:2306.15794Syntology official (archive's flag): 5 ran · 17 ran (of which 11 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 4 where Syntology's instrument failed) · 11 unverified (of 28 harvested samples) · 1 pointer-only (licence)
-
LeanDojo: Theorem Proving with Retrieval-Augmented Language Models 27 Jun 2023 · 3 repositories · arXiv:2306.15626Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples) · 1 pointer-only (licence)
-
Novel Hybrid-Learning Algorithms for Improved Millimeter-Wave Imaging Systems 27 Jun 2023 · 1 repository · arXiv:2306.15341
-
SparseOptimizer: Sparsify Language Models through Moreau-Yosida Regularization and Accelerate via Compiler Co-design 27 Jun 2023 · 0 repositories · arXiv:2306.15656
-
Style-transfer based Speech and Audio-visual Scene Understanding for Robot Action Sequence Acquisition from Videos 27 Jun 2023 · 0 repositories · arXiv:2306.15644
-
Taming Detection Transformers for Medical Object Detection 27 Jun 2023 · 0 repositories · arXiv:2306.15472
-
Towards predicting Pedestrian Evacuation Time and Density from Floorplans using a Vision Transformer 27 Jun 2023 · 1 repository · arXiv:2306.15318
-
Variational latent discrete representation for time series modelling 27 Jun 2023 · 0 repositories · arXiv:2306.15282
-
CST-YOLO: A Novel Method for Blood Cell Detection Based on Improved YOLOv7 and CNN-Swin Transformer 26 Jun 2023 · 1 repository · arXiv:2306.14590
-
DNABERT-2: Efficient Foundation Model and Benchmark For Multi-Species Genome 26 Jun 2023 · 6 repositories · arXiv:2306.15006Syntology community repositories only · 13 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 1 honoured, 0 violated, 8 with no contract checked; 4 where Syntology's instrument failed) · 10 unverified (of 23 harvested samples) · 1 pointer-only (licence)
-
Exploring the Robustness of Large Language Models for Solving Programming Problems 26 Jun 2023 · 0 repositories · arXiv:2306.14583
-
FeSViBS: Federated Split Learning of Vision Transformer with Block Sampling 26 Jun 2023 · 1 repository · arXiv:2306.14638
-
Large Multimodal Models: Notes on CVPR 2023 Tutorial 26 Jun 2023 · 0 repositories · arXiv:2306.14895
-
LM4HPC: Towards Effective Language Model Application in High-Performance Computing 26 Jun 2023 · 0 repositories · arXiv:2306.14979
-
LongCoder: A Long-Range Pre-trained Language Model for Code Completion 26 Jun 2023 · 1 repository · arXiv:2306.14893Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
ParameterNet: Parameters Are All You Need 26 Jun 2023 · 0 repositories · arXiv:2306.14525
-
Supervised Pretraining Can Learn In-Context Reinforcement Learning 26 Jun 2023 · 0 repositories · arXiv:2306.14892
-
ViNT: A Foundation Model for Visual Navigation 26 Jun 2023 · 1 repository · arXiv:2306.14846
-
Adaptive Window Pruning for Efficient Local Motion Deblurring 25 Jun 2023 · 0 repositories · arXiv:2306.14268
-
G-STO: Sequential Main Shopping Intention Detection via Graph-Regularized Stochastic Transformer 25 Jun 2023 · 0 repositories · arXiv:2306.14314
-
Interactive Design by Integrating a Large Pre-Trained Language Model and Building Information Modeling 25 Jun 2023 · 0 repositories · arXiv:2306.14165
-
Let's Do a Thought Experiment: Using Counterfactuals to Improve Moral Reasoning 25 Jun 2023 · 0 repositories · arXiv:2306.14308
-
Multi-Scale Cross Contrastive Learning for Semi-Supervised Medical Image Segmentation 25 Jun 2023 · 0 repositories · arXiv:2306.14293
-
Revolutionizing Cyber Threat Detection with Large Language Models: A privacy-preserving BERT-based Lightweight Model for IoT/IIoT Devices 25 Jun 2023 · 0 repositories · arXiv:2306.14263
-
Steganographic Capacity of Deep Learning Models 25 Jun 2023 · 0 repositories · arXiv:2306.17189
-
Action Q-Transformer: Visual Explanation in Deep Reinforcement Learning with Encoder-Decoder Model using Action Query 24 Jun 2023 · 0 repositories · arXiv:2306.13879
-
Can GPT-4 Support Analysis of Textual Data in Tasks Requiring Highly Specialized Domain Expertise? 24 Jun 2023 · 0 repositories · arXiv:2306.13906
-
Emotion Flip Reasoning in Multiparty Conversations 24 Jun 2023 · 0 repositories · arXiv:2306.13959
-
Fusing Multimodal Signals on Hyper-complex Space for Extreme Abstractive Text Summarization (TL;DR) of Scientific Contents 24 Jun 2023 · 1 repository · arXiv:2306.13968
-
Is Pre-training Truly Better Than Meta-Learning? 24 Jun 2023 · 0 repositories · arXiv:2306.13841
-
Large Language Models as Sous Chefs: Revising Recipes with GPT-3 24 Jun 2023 · 1 repository · arXiv:2306.13986
-
Large Sequence Models for Sequential Decision-Making: A Survey 24 Jun 2023 · 0 repositories · arXiv:2306.13945
-
On the Uses of Large Language Models to Interpret Ambiguous Cyberattack Descriptions 24 Jun 2023 · 0 repositories · arXiv:2306.14062
-
Waypoint Transformer: Reinforcement Learning via Supervised Learning with Intermediate Targets 24 Jun 2023 · 0 repositories · arXiv:2306.14069
-
Abstractive Text Summarization for Resumes With Cutting Edge NLP Transformers and LSTM 23 Jun 2023 · 0 repositories · arXiv:2306.13315
-
Bridging the Performance Gap between DETR and R-CNN for Graphical Object Detection in Document Images 23 Jun 2023 · 0 repositories · arXiv:2306.13526
-
Cross-Language Speech Emotion Recognition Using Multimodal Dual Attention Transformers 23 Jun 2023 · 0 repositories · arXiv:2306.13804
-
Efficient Online Processing with Deep Neural Networks 23 Jun 2023 · 1 repository · arXiv:2306.13474
-
On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes 23 Jun 2023 · 0 repositories · arXiv:2306.13649
-
Incorporating Graph Information in Transformer-based AMR Parsing 23 Jun 2023 · 1 repository · arXiv:2306.13467
-
LLM-Assisted Content Analysis: Using Large Language Models to Support Deductive Coding 23 Jun 2023 · 0 repositories · arXiv:2306.14924
-
Retrieval-Pretrained Transformer: Long-range Language Modeling with Self-retrieval 23 Jun 2023 · 1 repository · arXiv:2306.13421Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
ProRes: Exploring Degradation-aware Visual Prompt for Universal Image Restoration 23 Jun 2023 · 1 repository · arXiv:2306.13653Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Resume Information Extraction via Post-OCR Text Processing 23 Jun 2023 · 0 repositories · arXiv:2306.13775
-
Swin-Free: Achieving Better Cross-Window Attention and Efficiency with Size-varying Window 23 Jun 2023 · 0 repositories · arXiv:2306.13776
-
System-Level Natural Language Feedback 23 Jun 2023 · 1 repository · arXiv:2306.13588Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
The Double Helix inside the NLP Transformer 23 Jun 2023 · 0 repositories · arXiv:2306.13817
-
Upscaling Global Hourly GPP with Temporal Fusion Transformer (TFT) 23 Jun 2023 · 0 repositories · arXiv:2306.13815
-
Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale 23 Jun 2023 · 1 repository · arXiv:2306.15687Syntology 10 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 3 honoured, 3 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples) · 3 pointer-only (licence)