Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 66
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 66 of 190: papers 6,501 to 6,600 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Utilizing GPT to Enhance Text Summarization: A Strategy to Minimize Hallucinations 7 May 2024 · 0 repositories · arXiv:2405.04039
-
Vision Mamba: A Comprehensive Survey and Taxonomy 7 May 2024 · 1 repository · arXiv:2405.04404
-
xLSTM: Extended Long Short-Term Memory 7 May 2024 · 5 repositories · arXiv:2405.04517Syntology community repositories only · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples)
-
AlphaMath Almost Zero: Process Supervision without Process 6 May 2024 · 1 repository · arXiv:2405.03553
-
Anchored Answers: Unravelling Positional Bias in GPT-2's Multiple-Choice Questions 6 May 2024 · 1 repository · arXiv:2405.03205
-
Characterizing the Dilemma of Performance and Index Size in Billion-Scale Vector Search and Breaking It with Second-Tier Memory 6 May 2024 · 0 repositories · arXiv:2405.03267
-
Class-relevant Patch Embedding Selection for Few-Shot Image Classification 6 May 2024 · 0 repositories · arXiv:2405.03722
-
Compressing Long Context for Enhancing RAG with AMR-based Concept Distillation 6 May 2024 · 0 repositories · arXiv:2405.03085
-
CRA5: Extreme Compression of ERA5 for Portable Global Climate and Weather Research via an Efficient Variational Transformer 6 May 2024 · 1 repository · arXiv:2405.03376
-
Dual Relation Mining Network for Zero-Shot Learning 6 May 2024 · 0 repositories · arXiv:2405.03613
-
Enhancing DETRs Variants through Improved Content Query and Similar Query Aggregation 6 May 2024 · 0 repositories · arXiv:2405.03318
-
ERAGent: Enhancing Retrieval-Augmented Language Models with Improved Accuracy, Efficiency, and Personalization 6 May 2024 · 1 repository · arXiv:2405.06683
-
Exploring the Frontiers of Softmax: Provable Optimization, Applications in Diffusion Model, and Beyond 6 May 2024 · 0 repositories · arXiv:2405.03251
-
GREEN: Generative Radiology Report Evaluation and Error Notation 6 May 2024 · 0 repositories · arXiv:2405.03595
-
Hire Me or Not? Examining Language Model's Behavior with Occupation Attributes 6 May 2024 · 1 repository · arXiv:2405.06687Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
Intra-task Mutual Attention based Vision Transformer for Few-Shot Learning 6 May 2024 · 0 repositories · arXiv:2405.03109
-
Investigating Personalized Driving Behaviors in Dilemma Zones: Analysis and Prediction of Stop-or-Go Decisions 6 May 2024 · 0 repositories · arXiv:2405.03873
-
Large Language Models Reveal Information Operation Goals, Tactics, and Narrative Frames 6 May 2024 · 1 repository · arXiv:2405.03688
-
MAmmoTH2: Scaling Instructions from the Web 6 May 2024 · 0 repositories · arXiv:2405.03548
-
Modality Prompts for Arbitrary Modality Salient Object Detection 6 May 2024 · 0 repositories · arXiv:2405.03351
-
ReCycle: Fast and Efficient Long Time Series Forecasting with Residual Cyclic Transformers 6 May 2024 · 1 repository · arXiv:2405.03429
-
Salient Object Detection From Arbitrary Modalities 6 May 2024 · 1 repository · arXiv:2405.03352
-
SocialFormer: Social Interaction Modeling with Edge-enhanced Heterogeneous Graph Transformers for Trajectory Prediction 6 May 2024 · 0 repositories · arXiv:2405.03809
-
Transformer-based RGB-T Tracking with Channel and Spatial Feature Fusion 6 May 2024 · 1 repository · arXiv:2405.03177
-
Transformer models as an efficient replacement for statistical test suites to evaluate the quality of random numbers 6 May 2024 · 0 repositories · arXiv:2405.03904
-
Can Large Language Models Make the Grade? An Empirical Study Evaluating LLMs Ability to Mark Short Answer Questions in K-12 Education 5 May 2024 · 0 repositories · arXiv:2405.02985
-
E-TSL: A Continuous Educational Turkish Sign Language Dataset with Baseline Methods 5 May 2024 · 0 repositories · arXiv:2405.02984
-
Graph as Point Set 5 May 2024 · 1 repository · arXiv:2405.02795Syntology official (archive's flag): 20 ran · 20 ran (of which 15 constructed an object rather than computing a result; 16 with no instrument failure: 1 honoured, 0 violated, 15 with no contract checked; 4 where Syntology's instrument failed) · 7 unverified (of 27 harvested samples) · 27 pointer-only (licence)
-
Labeling supervised fine-tuning data with the scaling law 5 May 2024 · 2 repositories · arXiv:2405.02817
-
IceFormer: Accelerated Inference with Long-Sequence Transformers on CPUs 5 May 2024 · 0 repositories · arXiv:2405.02842
-
Leveraging Lecture Content for Improved Feedback: Explorations with GPT-4 and Retrieval Augmented Generation 5 May 2024 · 0 repositories · arXiv:2405.06681
-
Multi-hop graph transformer network for 3D human pose estimation 5 May 2024 · 0 repositories · arXiv:2405.03055
-
NegativePrompt: Leveraging Psychology for Large Language Models Enhancement via Negative Emotional Stimuli 5 May 2024 · 1 repository · arXiv:2405.02814Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Overconfidence is Key: Verbalized Uncertainty Evaluation in Large Language and Vision-Language Models 5 May 2024 · 0 repositories · arXiv:2405.02917
-
Stochastic RAG: End-to-End Retrieval-Augmented Generation through Expected Utility Maximization 5 May 2024 · 0 repositories · arXiv:2405.02816
-
Unraveling the Dominance of Large Language Models Over Transformer Models for Bangla Natural Language Inference: A Comprehensive Study 5 May 2024 · 1 repository · arXiv:2405.02937
-
A Combination of BERT and Transformer for Vietnamese Spelling Correction 4 May 2024 · 0 repositories · arXiv:2405.02573
-
Assessing Adversarial Robustness of Large Language Models: An Empirical Study 4 May 2024 · 0 repositories · arXiv:2405.02764
-
Boosting 3D Neuron Segmentation with 2D Vision Transformer Pre-trained on Natural Images 4 May 2024 · 0 repositories · arXiv:2405.02686
-
Open-SQL Framework: Enhancing Text-to-SQL on Open-source Large Language Models 4 May 2024 · 0 repositories · arXiv:2405.06674
-
PropertyGPT: LLM-driven Formal Verification of Smart Contracts through Retrieval-Augmented Property Generation 4 May 2024 · 1 repository · arXiv:2405.02580
-
ViTALS: Vision Transformer for Action Localization in Surgical Nephrectomy 4 May 2024 · 0 repositories · arXiv:2405.02571
-
Automating the Enterprise with Foundation Models 3 May 2024 · 1 repository · arXiv:2405.03710Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
Comparative Analysis of Retrieval Systems in the Real World 3 May 2024 · 0 repositories · arXiv:2405.02048
-
CVTGAD: Simplified Transformer with Cross-View Attention for Unsupervised Graph-level Anomaly Detection 3 May 2024 · 1 repository · arXiv:2405.02359Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Evaluating Large Language Models for Structured Science Summarization in the Open Research Knowledge Graph 3 May 2024 · 0 repositories · arXiv:2405.02105
-
Exploring Combinatorial Problem Solving with Large Language Models: A Case Study on the Travelling Salesman Problem Using GPT-3.5 Turbo 3 May 2024 · 0 repositories · arXiv:2405.01997
-
Attribution in Scientific Literature: New Benchmark and Methods 3 May 2024 · 0 repositories · arXiv:2405.02228
-
Single and Multi-Hop Question-Answering Datasets for Reticular Chemistry with GPT-4-Turbo 3 May 2024 · 1 repository · arXiv:2405.02128
-
SatSwinMAE: Efficient Autoencoding for Multiscale Time-series Satellite Imagery 3 May 2024 · 0 repositories · arXiv:2405.02512
-
Technical report on target classification in SAR track 3 May 2024 · 0 repositories · arXiv:2405.02361
-
A Survey on Large Language Models for Critical Societal Domains: Finance, Healthcare, and Law 2 May 2024 · 1 repository · arXiv:2405.01769
-
Bayesian Optimization with LLM-Based Acquisition Functions for Natural Language Preference Elicitation 2 May 2024 · 0 repositories · arXiv:2405.00981
-
CrossMPT: Cross-attention Message-Passing Transformer for Error Correcting Codes 2 May 2024 · 0 repositories · arXiv:2405.01033
-
Early Transformers: A study on Efficient Training of Transformer Models through Early-Bird Lottery Tickets 2 May 2024 · 0 repositories · arXiv:2405.02353
-
How Can I Get It Right? Using GPT to Rephrase Incorrect Trainee Responses 2 May 2024 · 0 repositories · arXiv:2405.00970
-
Investigating Wit, Creativity, and Detectability of Large Language Models in Domain-Specific Writing Style Adaptation of Reddit's Showerthoughts 2 May 2024 · 1 repository · arXiv:2405.01660
-
Less is More: on the Over-Globalizing Problem in Graph Transformers 2 May 2024 · 1 repository · arXiv:2405.01102Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Leverage Multi-source Traffic Demand Data Fusion with Transformer Model for Urban Parking Prediction 2 May 2024 · 0 repositories · arXiv:2405.01055
-
MiniGPT-3D: Efficiently Aligning 3D Point Clouds with Large Language Models using 2D Priors 2 May 2024 · 1 repository · arXiv:2405.01413
-
Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models 2 May 2024 · 1 repository · arXiv:2405.01535
-
Reinforcement Learning for Edit-Based Non-Autoregressive Neural Machine Translation 2 May 2024 · 0 repositories · arXiv:2405.01280
-
The Effectiveness of LLMs as Annotators: A Comparative Overview and Empirical Analysis of Direct Representation 2 May 2024 · 0 repositories · arXiv:2405.01299
-
The Role of Model Architecture and Scale in Predicting Molecular Properties: Insights from Fine-Tuning RoBERTa, BART, and LLaMA 2 May 2024 · 2 repositories · arXiv:2405.00949
-
Transformers Fusion across Disjoint Samples for Hyperspectral Image Classification 2 May 2024 · 0 repositories · arXiv:2405.01095
-
UQA: Corpus for Urdu Question Answering 2 May 2024 · 3 repositories · arXiv:2405.01458
-
WildChat: 1M ChatGPT Interaction Logs in the Wild 2 May 2024 · 0 repositories · arXiv:2405.01470
-
Brighteye: Glaucoma Screening with Color Fundus Photographs based on Vision Transformer 1 May 2024 · 1 repository · arXiv:2405.00857
-
CourseAssist: Pedagogically Appropriate AI Tutor for Computer Science Education 1 May 2024 · 0 repositories · arXiv:2407.10246
-
DAM: A Universal Dual Attention Mechanism for Multimodal Timeseries Cryptocurrency Trend Forecasting 1 May 2024 · 1 repository · arXiv:2405.00522
-
How Can I Improve? Using GPT to Highlight the Desired and Undesired Parts of Open-ended Responses 1 May 2024 · 0 repositories · arXiv:2405.00291
-
ICU Bloodstream Infection Prediction: A Transformer-Based Approach for EHR Analysis 1 May 2024 · 1 repository · arXiv:2405.00819
-
Integrating A.I. in Higher Education: Protocol for a Pilot Study with 'SAMCares: An Adaptive Learning Hub' 1 May 2024 · 1 repository · arXiv:2405.00330
-
Opinion Mining Using Pre-Trained Large Language Models: Identifying the Type, Polarity, Intensity, Expression, and Source of Private States 1 May 2024 · 1 repository
-
Self-Play Preference Optimization for Language Model Alignment 1 May 2024 · 1 repository · arXiv:2405.00675Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
Transformer-based Reasoning for Learning Evolutionary Chain of Events on Temporal Knowledge Graph 1 May 2024 · 1 repository · arXiv:2405.00352
-
A Framework for Leveraging Human Computation Gaming to Enhance Knowledge Graphs for Accuracy Critical Generative AI Applications 30 Apr 2024 · 0 repositories · arXiv:2404.19729
-
Aspect and Opinion Term Extraction Using Graph Attention Network 30 Apr 2024 · 0 repositories · arXiv:2404.19260
-
Automatic Cardiac Pathology Recognition in Echocardiography Images Using Higher Order Dynamic Mode Decomposition and a Vision Transformer for Small Datasets 30 Apr 2024 · 0 repositories · arXiv:2404.19579
-
Better & Faster Large Language Models via Multi-token Prediction 30 Apr 2024 · 1 repository · arXiv:2404.19737
-
Can Large Language Models put 2 and 2 together? Probing for Entailed Arithmetical Relationships 30 Apr 2024 · 0 repositories · arXiv:2404.19432
-
CLIP-Mamba: CLIP Pretrained Mamba Models with OOD and Hessian Evaluation 30 Apr 2024 · 1 repository · arXiv:2404.19394
-
Constrained Decoding for Secure Code Generation 30 Apr 2024 · 2 repositories · arXiv:2405.00218Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 13 harvested samples) · 1 pointer-only (licence)
-
Do Large Language Models Understand Conversational Implicature -- A case study with a chinese sitcom 30 Apr 2024 · 1 repository · arXiv:2404.19509Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 1 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Extending Llama-3's Context Ten-Fold Overnight 30 Apr 2024 · 1 repository · arXiv:2404.19553
-
Graphical Reasoning: LLM-based Semi-Open Relation Extraction 30 Apr 2024 · 1 repository · arXiv:2405.00216
-
Harmonic LLMs are Trustworthy 30 Apr 2024 · 0 repositories · arXiv:2404.19708
-
Neuro-Vision to Language: Enhancing Brain Recording-based Visual Reconstruction and Language Interaction 30 Apr 2024 · 0 repositories · arXiv:2404.19438
-
Octopus v4: Graph of language models 30 Apr 2024 · 0 repositories · arXiv:2404.19296
-
PANGeA: Procedural Artificial Narrative using Generative AI for Turn-Based Video Games 30 Apr 2024 · 0 repositories · arXiv:2404.19721
-
RepEval: Effective Text Evaluation with LLM Representation 30 Apr 2024 · 1 repository · arXiv:2404.19563Syntology official: no sample here; runs from other or unrecorded repositories · 18 ran (of which 0 constructed an object rather than computing a result; 15 with no instrument failure: 2 honoured, 0 violated, 13 with no contract checked; 3 where Syntology's instrument failed) · 9 unverified (of 27 harvested samples) · 6 pointer-only (licence)
-
Seeing Through the Clouds: Cloud Gap Imputation with Prithvi Foundation Model 30 Apr 2024 · 1 repository · arXiv:2404.19609
-
SPAFIT: Stratified Progressive Adaptation Fine-tuning for Pre-trained Large Language Models 30 Apr 2024 · 0 repositories · arXiv:2405.00201
-
Towards a Search Engine for Machines: Unified Ranking for Multiple Retrieval-Augmented Large Language Models 30 Apr 2024 · 1 repository · arXiv:2405.00175
-
TuBA: Cross-Lingual Transferability of Backdoor Attacks in LLMs with Instruction Tuning 30 Apr 2024 · 0 repositories · arXiv:2404.19597
-
Transformer-Enhanced Motion Planner: Attention-Guided Sampling for State-Specific Decision Making 30 Apr 2024 · 0 repositories · arXiv:2404.19403
-
Automated Construction of Theme-specific Knowledge Graphs 29 Apr 2024 · 0 repositories · arXiv:2404.19146
-
Can GPT-4 do L2 analytic assessment? 29 Apr 2024 · 0 repositories · arXiv:2404.18557
-
Capabilities of Gemini Models in Medicine 29 Apr 2024 · 0 repositories · arXiv:2404.18416
-
CVTN: Cross Variable and Temporal Integration for Time Series Forecasting 29 Apr 2024 · 0 repositories · arXiv:2404.18730