Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 83
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 83 of 190: papers 8,201 to 8,300 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Validation, Robustness, and Accuracy of Perturbation-Based Sensitivity Analysis Methods for Time-Series Deep Learning Models 29 Jan 2024 · 0 repositories · arXiv:2401.16521
-
Byte Pair Encoding Is All You Need For Automatic Bengali Speech Recognition 28 Jan 2024 · 0 repositories · arXiv:2401.15532
-
Identifying and Improving Disability Bias in GPT-Based Resume Screening 28 Jan 2024 · 0 repositories · arXiv:2402.01732
-
PRE: A Peer Review Based Large Language Model Evaluator 28 Jan 2024 · 0 repositories · arXiv:2401.15641
-
SCTransNet: Spatial-channel Cross Transformer Network for Infrared Small Target Detection 28 Jan 2024 · 1 repository · arXiv:2401.15583Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
A New Method for Vehicle Logo Recognition Based on Swin Transformer 27 Jan 2024 · 0 repositories · arXiv:2401.15458
-
Baichuan2-Sum: Instruction Finetune Baichuan2-7B Model for Dialogue Summarization 27 Jan 2024 · 0 repositories · arXiv:2401.15496
-
ConvoSense: Overcoming Monotonous Commonsense Inferences for Conversational AI 27 Jan 2024 · 1 repository · arXiv:2401.15471
-
DataFrame QA: A Universal LLM Framework on DataFrame Question Answering Without Data Exposure 27 Jan 2024 · 0 repositories · arXiv:2401.15463
-
Enhancing Large Language Model Performance To Answer Questions and Extract Information More Accurately 27 Jan 2024 · 0 repositories · arXiv:2402.01722
-
Equipping Language Models with Tool Use Capability for Tabular Data Analysis in Finance 27 Jan 2024 · 0 repositories · arXiv:2401.15328
-
Fortifying Ethical Boundaries in AI: Advanced Strategies for Enhancing Security in Large Language Models 27 Jan 2024 · 0 repositories · arXiv:2402.01725
-
Improving Medical Reasoning through Retrieval and Self-Reflection with Retrieval-Augmented Large Language Models 27 Jan 2024 · 1 repository · arXiv:2401.15269
-
MiTU-Net: A fine-tuned U-Net with SegFormer backbone for segmenting pubic symphysis-fetal head 27 Jan 2024 · 1 repository · arXiv:2401.15513
-
MultiHop-RAG: Benchmarking Retrieval-Augmented Generation for Multi-Hop Queries 27 Jan 2024 · 2 repositories · arXiv:2401.15391Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
ParaTransCNN: Parallelized TransCNN Encoder for Medical Image Segmentation 27 Jan 2024 · 1 repository · arXiv:2401.15307
-
Prompting Diverse Ideas: Increasing AI Idea Variance 27 Jan 2024 · 0 repositories · arXiv:2402.01727
-
Transformer-based Clipped Contrastive Quantization Learning for Unsupervised Image Retrieval 27 Jan 2024 · 0 repositories · arXiv:2401.15362
-
A Korean Legal Judgment Prediction Dataset for Insurance Disputes 26 Jan 2024 · 0 repositories · arXiv:2401.14654
-
Adaptive Point Transformer 26 Jan 2024 · 0 repositories · arXiv:2401.14845
-
CascadedGaze: Efficiency in Global Context Extraction for Image Restoration 26 Jan 2024 · 1 repository · arXiv:2401.15235Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 11 harvested samples) · 4 pointer-only (licence)
-
ChemDFM: A Large Language Foundation Model for Chemistry 26 Jan 2024 · 1 repository · arXiv:2401.14818Syntology 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
DAM: Diffusion Activation Maximization for 3D Global Explanations 26 Jan 2024 · 1 repository · arXiv:2401.14938
-
Deep Learning with Tabular Data: A Self-supervised Approach 26 Jan 2024 · 1 repository · arXiv:2401.15238
-
Endowing Protein Language Models with Structural Knowledge 26 Jan 2024 · 1 repository · arXiv:2401.14819
-
Enhancing Diagnostic Accuracy through Multi-Agent Conversations: Using Large Language Models to Mitigate Cognitive Bias 26 Jan 2024 · 0 repositories · arXiv:2401.14589
-
Evaluation of LLM Chatbots for OSINT-based Cyber Threat Awareness 26 Jan 2024 · 0 repositories · arXiv:2401.15127
-
FedGT: Federated Node Classification with Scalable Graph Transformer 26 Jan 2024 · 0 repositories · arXiv:2401.15203
-
From Blurry to Brilliant Detection: YOLOv5-Based Aerial Object Detection with Super Resolution 26 Jan 2024 · 0 repositories · arXiv:2401.14661
-
From GPT-4 to Gemini and Beyond: Assessing the Landscape of MLLMs on Generalizability, Trustworthiness and Causality through Four Modalities 26 Jan 2024 · 0 repositories · arXiv:2401.15071
-
From RAG to QA-RAG: Integrating Generative AI for Pharmaceutical Regulatory Compliance Process 26 Jan 2024 · 1 repository · arXiv:2402.01717
-
GeoDecoder: Empowering Multimodal Map Understanding 26 Jan 2024 · 0 repositories · arXiv:2401.15118
-
Health Text Simplification: An Annotated Corpus for Digestive Cancer Education and Novel Strategies for Reinforcement Learning 26 Jan 2024 · 1 repository · arXiv:2401.15043
-
Inferring Data Preconditions from Deep Learning Models for Trustworthy Prediction in Deployment 26 Jan 2024 · 1 repository · arXiv:2401.14628
-
LYT-NET: Lightweight YUV Transformer-based Network for Low-light Image Enhancement 26 Jan 2024 · 2 repositories · arXiv:2401.15204
-
On the generalization capacity of neural networks during generic multimodal reasoning 26 Jan 2024 · 1 repository · arXiv:2401.15030Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
PL-FSCIL: Harnessing the Power of Prompts for Few-Shot Class-Incremental Learning 26 Jan 2024 · 1 repository · arXiv:2401.14807
-
Scalable Qualitative Coding with LLMs: Chain-of-Thought Reasoning Matches Human Performance in Some Hermeneutic Tasks 26 Jan 2024 · 0 repositories · arXiv:2401.15170
-
The Power of Noise: Redefining Retrieval for RAG Systems 26 Jan 2024 · 3 repositories · arXiv:2401.14887Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Topology-Aware Exploration of Energy-Based Models Equilibrium: Toric QC-LDPC Codes and Hyperbolic MET QC-LDPC Codes 26 Jan 2024 · 1 repository · arXiv:2401.14749
-
A comparative study of zero-shot inference with large language models and supervised modeling in breast cancer pathology classification 25 Jan 2024 · 0 repositories · arXiv:2401.13887
-
(Chat)GPT v BERT: Dawn of Justice for Semantic Change Detection 25 Jan 2024 · 1 repository · arXiv:2401.14040
-
DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence 25 Jan 2024 · 1 repository · arXiv:2401.14196Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 10 harvested samples) · 1 pointer-only (licence)
-
Evaluating GPT-3.5's Awareness and Summarization Abilities for European Constitutional Texts with Shared Topics 25 Jan 2024 · 0 repositories · arXiv:2401.14524
-
Investigate-Consolidate-Exploit: A General Strategy for Inter-Task Agent Self-Evolution 25 Jan 2024 · 0 repositories · arXiv:2401.13996
-
LLM on FHIR -- Demystifying Health Records 25 Jan 2024 · 0 repositories · arXiv:2402.01711
-
LongHealth: A Question Answering Benchmark with Long Clinical Documents 25 Jan 2024 · 1 repository · arXiv:2401.14490Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Prompting Large Language Models for Zero-Shot Clinical Prediction with Structured Longitudinal Electronic Health Record Data 25 Jan 2024 · 1 repository · arXiv:2402.01713Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 1 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
TrICy: Trigger-guided Data-to-text Generation with Intent aware Attention-Copy 25 Jan 2024 · 0 repositories · arXiv:2402.01714
-
Unmasking and Quantifying Racial Bias of Large Language Models in Medical Report Generation 25 Jan 2024 · 0 repositories · arXiv:2401.13867
-
VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech 25 Jan 2024 · 0 repositories · arXiv:2401.14321
-
WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models 25 Jan 2024 · 2 repositories · arXiv:2401.13919Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Zero-shot Sequential Neuro-symbolic Reasoning for Automatically Generating Architecture Schematic Designs 25 Jan 2024 · 0 repositories · arXiv:2402.00052
-
ZS4C: Zero-Shot Synthesis of Compilable Code for Incomplete Code Snippets using LLMs 25 Jan 2024 · 0 repositories · arXiv:2401.14279
-
A Unified Approach to Emotion Detection and Task-Oriented Dialogue Modeling 24 Jan 2024 · 1 repository · arXiv:2401.13789
-
Automated Root Causing of Cloud Incidents using In-Context Learning with GPT-4 24 Jan 2024 · 0 repositories · arXiv:2401.13810
-
Can GPT-3.5 Generate and Code Discharge Summaries? 24 Jan 2024 · 1 repository · arXiv:2401.13512Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Fine-Grained Stateful Knowledge Exploration: A Novel Paradigm for Integrating Knowledge Graphs with Large Language Models 24 Jan 2024 · 1 repository · arXiv:2401.13444
-
ConTextual: Evaluating Context-Sensitive Text-Rich Visual Reasoning in Large Multimodal Models 24 Jan 2024 · 1 repository · arXiv:2401.13311Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Discovering Mathematical Formulas from Data via GPT-guided Monte Carlo Tree Search 24 Jan 2024 · 0 repositories · arXiv:2401.14424
-
Evaluation of General Large Language Models in Contextually Assessing Semantic Concepts Extracted from Adult Critical Care Electronic Health Record Notes 24 Jan 2024 · 0 repositories · arXiv:2401.13588
-
Graph Guided Question Answer Generation for Procedural Question-Answering 24 Jan 2024 · 0 repositories · arXiv:2401.13594
-
How Good is ChatGPT at Face Biometrics? A First Look into Recognition, Soft Biometrics, and Explainability 24 Jan 2024 · 1 repository · arXiv:2401.13641
-
Inadequacy of common stochastic neural networks for reliable clinical decision support 24 Jan 2024 · 0 repositories · arXiv:2401.13657
-
Graph Diffusion Transformers for Multi-Conditional Molecular Generation 24 Jan 2024 · 1 repository · arXiv:2401.13858Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Language-Guided World Models: A Model-Based Approach to AI Control 24 Jan 2024 · 0 repositories · arXiv:2402.01695
-
Learning Representations for Clustering via Partial Information Discrimination and Cross-Level Interaction 24 Jan 2024 · 1 repository · arXiv:2401.13503
-
Research about the Ability of LLM in the Tamper-Detection Area 24 Jan 2024 · 0 repositories · arXiv:2401.13504
-
LPNL: Scalable Link Prediction with Large Language Models 24 Jan 2024 · 0 repositories · arXiv:2401.13227
-
SegMamba: Long-range Sequential Modeling Mamba For 3D Medical Image Segmentation 24 Jan 2024 · 1 repository · arXiv:2401.13560Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Segment Any Cell: A SAM-based Auto-prompting Fine-tuning Framework for Nuclei Segmentation 24 Jan 2024 · 0 repositories · arXiv:2401.13220
-
TAT-LLM: A Specialized Language Model for Discrete Reasoning over Tabular and Textual Data 24 Jan 2024 · 0 repositories · arXiv:2401.13223
-
ARGS: Alignment as Reward-Guided Search 23 Jan 2024 · 1 repository · arXiv:2402.01694Syntology official (archive's flag): 5 ran · 5 ran (of which 1 constructed an object rather than computing a result; 4 with no instrument failure: 3 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
EL-VIT: Probing Vision Transformer with Interactive Visualization 23 Jan 2024 · 0 repositories · arXiv:2401.12666
-
Exploration and Improvement of Nerf-based 3D Scene Editing Techniques 23 Jan 2024 · 0 repositories · arXiv:2401.12456
-
KAM-CoT: Knowledge Augmented Multimodal Chain-of-Thoughts Reasoning 23 Jan 2024 · 0 repositories · arXiv:2401.12863
-
MAST: Video Polyp Segmentation with a Mixture-Attention Siamese Transformer 23 Jan 2024 · 1 repository · arXiv:2401.12439
-
Meta-Prompting: Enhancing Language Models with Task-Agnostic Scaffolding 23 Jan 2024 · 1 repository · arXiv:2401.12954Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples) · 4 pointer-only (licence)
-
On the Efficacy of Text-Based Input Modalities for Action Anticipation 23 Jan 2024 · 0 repositories · arXiv:2401.12972
-
Quality of Answers of Generative Large Language Models vs Peer Patients for Interpreting Lab Test Results for Lay Patients: Evaluation Study 23 Jan 2024 · 0 repositories · arXiv:2402.01693
-
Revolutionizing Retrieval-Augmented Generation with Enhanced PDF Structure Recognition 23 Jan 2024 · 0 repositories · arXiv:2401.12599
-
TroVE: Inducing Verifiable and Efficient Toolboxes for Solving Programmatic Tasks 23 Jan 2024 · 1 repository · arXiv:2401.12869Syntology official (archive's flag): 16 ran · 16 ran (of which 0 constructed an object rather than computing a result; 16 with no instrument failure: 0 honoured, 0 violated, 16 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 19 harvested samples) · 19 pointer-only (licence)
-
APT: Adaptive Pruning and Tuning Pretrained Language Models for Efficient Training and Inference 22 Jan 2024 · 1 repository · arXiv:2401.12200Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 1 pointer-only (licence)
-
BETA: Binarized Energy-Efficient Transformer Accelerator at the Edge 22 Jan 2024 · 0 repositories · arXiv:2401.11851
-
Codebook-enabled Generative End-to-end Semantic Communication Powered by Transformer 22 Jan 2024 · 0 repositories · arXiv:2402.16868
-
Enhancing In-context Learning via Linear Probe Calibration 22 Jan 2024 · 1 repository · arXiv:2401.12406Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Evaluation of QCNN-LSTM for Disability Forecasting in Multiple Sclerosis Using Sequential Multisequence MRI 22 Jan 2024 · 0 repositories · arXiv:2401.12132
-
Friends Across Time: Multi-Scale Action Segmentation Transformer for Surgical Phase Recognition 22 Jan 2024 · 0 repositories · arXiv:2401.11644
-
Investigating Large Language Models for Financial Causality Detection in Multilingual Setup 22 Jan 2024 · 0 repositories
-
LKFormer: Large Kernel Transformer for Infrared Image Super-Resolution 22 Jan 2024 · 1 repository · arXiv:2401.11859
-
MsSVT++: Mixed-scale Sparse Voxel Transformer with Center Voting for 3D Object Detection 22 Jan 2024 · 0 repositories · arXiv:2401.11718
-
OnDev-LCT: On-Device Lightweight Convolutional Transformers towards federated learning 22 Jan 2024 · 0 repositories · arXiv:2401.11652
-
P2DT: Mitigating Forgetting in task-incremental Learning with progressive prompt Decision Transformer 22 Jan 2024 · 0 repositories · arXiv:2401.11666
-
ATFusion: An Alternate Cross-Attention Transformer Network for Infrared and Visible Image Fusion 22 Jan 2024 · 0 repositories · arXiv:2401.11675
-
Revolutionizing Finance with LLMs: An Overview of Applications and Insights 22 Jan 2024 · 0 repositories · arXiv:2401.11641
-
Speak It Out: Solving Symbol-Related Problems with Symbol-to-Language Conversion for Language Models 22 Jan 2024 · 1 repository · arXiv:2401.11725
-
SuperCLUE-Math6: Graded Multi-Step Math Reasoning Benchmark for LLMs in Chinese 22 Jan 2024 · 1 repository · arXiv:2401.11819
-
Parsimony or Capability? Decomposition Delivers Both in Long-term Time Series Forecasting 22 Jan 2024 · 0 repositories · arXiv:2401.11929
-
The Right Model for the Job: An Evaluation of Legal Multi-Label Classification Baselines 22 Jan 2024 · 0 repositories · arXiv:2401.11852
-
Adversarial Augmentation Training Makes Action Recognition Models More Robust to Realistic Video Distribution Shifts 21 Jan 2024 · 1 repository · arXiv:2401.11406