Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 92
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 92 of 190: papers 9,101 to 9,200 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
AcademicGPT: Empowering Academic Research 21 Nov 2023 · 0 repositories · arXiv:2311.12315
-
Advancing Transformer Architecture in Long-Context Large Language Models: A Comprehensive Survey 21 Nov 2023 · 1 repository · arXiv:2311.12351
-
ALPHA: AnomaLous Physiological Health Assessment Using Large Language Models 21 Nov 2023 · 1 repository · arXiv:2311.12524
-
GeoLocator: a location-integrated large multimodal model for inferring geo-privacy 21 Nov 2023 · 0 repositories · arXiv:2311.13018
-
AudioLog: LLMs-Powered Long Audio Logging with Hybrid Token-Semantic Contrastive Learning 21 Nov 2023 · 1 repository · arXiv:2311.12371
-
Causality is all you need 21 Nov 2023 · 0 repositories · arXiv:2311.12307
-
Descriptor and Word Soups: Overcoming the Parameter Efficiency Accuracy Tradeoff for Out-of-Distribution Few-shot Learning 21 Nov 2023 · 1 repository · arXiv:2311.13612
-
Extracting Definienda in Mathematical Scholarly Articles with Transformers 21 Nov 2023 · 2 repositories · arXiv:2311.12448
-
From Classification to Clinical Insights: Towards Analyzing and Reasoning About Mobile and Behavioral Health Data With Large Language Models 21 Nov 2023 · 1 repository · arXiv:2311.13063
-
GAIA: a benchmark for General AI Assistants 21 Nov 2023 · 2 repositories · arXiv:2311.12983
-
GPT4Motion: Scripting Physical Motions in Text-to-Video Generation via Blender-Oriented GPT Planning 21 Nov 2023 · 0 repositories · arXiv:2311.12631
-
HoVer-UNet: Accelerating HoVerNet with UNet-based multi-class nuclei segmentation via knowledge distillation 21 Nov 2023 · 1 repository · arXiv:2311.12553
-
Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable Tasks 21 Nov 2023 · 1 repository · arXiv:2311.12997Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
HPCNeuroNet: Advancing Neuromorphic Audio Signal Processing with Transformer-Enhanced Spiking Neural Networks 21 Nov 2023 · 0 repositories · arXiv:2311.12449
-
IEKM: A Model Incorporating External Keyword Matrices 21 Nov 2023 · 0 repositories · arXiv:2311.12310
-
Improving Source-Free Target Adaptation with Vision Transformers Leveraging Domain Representation Images 21 Nov 2023 · 0 repositories · arXiv:2311.12589
-
Interpretation of the Transformer and Improvement of the Extractor 21 Nov 2023 · 1 repository · arXiv:2311.12678
-
InterPrompt: Interpretable Prompting for Interrelated Interpersonal Risk Factors in Reddit Posts 21 Nov 2023 · 0 repositories · arXiv:2311.12404
-
Learning to Compute Gröbner Bases 21 Nov 2023 · 2 repositories · arXiv:2311.12904Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified (of 18 harvested samples) · 17 pointer-only (licence)
-
Learning to Optimise Wind Farms with Graph Transformers 21 Nov 2023 · 0 repositories · arXiv:2311.12750
-
Long-MIL: Scaling Long Contextual Multiple Instance Learning for Histopathology Whole Slide Image Analysis 21 Nov 2023 · 0 repositories · arXiv:2311.12885
-
Oasis: Data Curation and Assessment System for Pretraining of Large Language Models 21 Nov 2023 · 1 repository · arXiv:2311.12537
-
A novel transformer-based approach for soil temperature prediction 20 Nov 2023 · 0 repositories · arXiv:2311.11626
-
Assessing Prompt Injection Risks in 200+ Custom GPTs 20 Nov 2023 · 1 repository · arXiv:2311.11538
-
Correlated Attention in Transformers for Multivariate Time Series 20 Nov 2023 · 0 repositories · arXiv:2311.11959
-
Decoupled DETR For Few-shot Object Detection 20 Nov 2023 · 0 repositories · arXiv:2311.11570
-
Disentangling Structure and Appearance in ViT Feature Space 20 Nov 2023 · 0 repositories · arXiv:2311.12193
-
Evil Geniuses: Delving into the Safety of LLM-based Agents 20 Nov 2023 · 1 repository · arXiv:2311.11855Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Generating Valid and Natural Adversarial Examples with Large Language Models 20 Nov 2023 · 0 repositories · arXiv:2311.11861
-
GPQA: A Graduate-Level Google-Proof Q&A Benchmark 20 Nov 2023 · 3 repositories · arXiv:2311.12022Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples)
-
Towards Human-Level Text Coding with LLMs: The Case of Fatherhood Roles in Public Policy Documents 20 Nov 2023 · 1 repository · arXiv:2311.11844
-
LiDAR-HMR: 3D Human Mesh Recovery from LiDAR 20 Nov 2023 · 2 repositories · arXiv:2311.11971
-
MemoryCompanion: A Smart Healthcare Solution to Empower Efficient Alzheimer's Care Via Unleashing Generative AI 20 Nov 2023 · 0 repositories · arXiv:2311.14730
-
Meta Prompting for AI Systems 20 Nov 2023 · 1 repository · arXiv:2311.11482Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
MGCT: Mutual-Guided Cross-Modality Transformer for Survival Outcome Prediction using Integrative Histopathology-Genomic Features 20 Nov 2023 · 1 repository · arXiv:2311.11659
-
PMP-Swin: Multi-Scale Patch Message Passing Swin Transformer for Retinal Disease Classification 20 Nov 2023 · 0 repositories · arXiv:2311.11669
-
Refactoring Programs Using Large Language Models with Few-Shot Examples 20 Nov 2023 · 0 repositories · arXiv:2311.11690
-
Unveiling the Power of Self-Attention for Shipping Cost Prediction: The Rate Card Transformer 20 Nov 2023 · 1 repository · arXiv:2311.11694
-
Which AI Technique Is Better to Classify Requirements? An Experiment with SVM, LSTM, and ChatGPT 20 Nov 2023 · 1 repository · arXiv:2311.11547
-
Inspecting Explainability of Transformer Models with Additional Statistical Information 19 Nov 2023 · 0 repositories · arXiv:2311.11378
-
Spot the Bot: Distinguishing Human-Written and Bot-Generated Texts Using Clustering and Information Theory Techniques 19 Nov 2023 · 0 repositories · arXiv:2311.11441
-
Behavior Optimized Image Generation 18 Nov 2023 · 0 repositories · arXiv:2311.10995
-
Bit Cipher -- A Simple yet Powerful Word Representation System that Integrates Efficiently with Language Models 18 Nov 2023 · 0 repositories · arXiv:2311.11012
-
Partially Randomizing Transformer Weights for Dialogue Response Diversity 18 Nov 2023 · 0 repositories · arXiv:2311.10943
-
Structure-Aware Sparse-View X-ray 3D Reconstruction 18 Nov 2023 · 2 repositories · arXiv:2311.10959Syntology official (archive's flag): 2 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
Vashantor: A Large-scale Multilingual Benchmark Dataset for Automated Translation of Bangla Regional Dialects to Bangla Language 18 Nov 2023 · 1 repository · arXiv:2311.11142
-
Visual AI and Linguistic Intelligence Through Steerability and Composability 18 Nov 2023 · 0 repositories · arXiv:2312.12383
-
Advancements in Generative AI: A Comprehensive Review of GANs, GPT, Autoencoders, Diffusion Model, and Transformers 17 Nov 2023 · 0 repositories · arXiv:2311.10242
-
Bias A-head? Analyzing Bias in Transformer-Based Language Model Attention Heads 17 Nov 2023 · 0 repositories · arXiv:2311.10395
-
Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2 17 Nov 2023 · 3 repositories · arXiv:2311.10702
-
DynaPipe: Optimizing Multi-task Training through Dynamic Pipelines 17 Nov 2023 · 2 repositories · arXiv:2311.10418
-
EduQuick: A Dataset Toward Evaluating Summarization of Informal Educational Content for Social Media 17 Nov 2023 · 0 repositories
-
Multi-entity Video Transformers for Fine-Grained Video Representation Learning 17 Nov 2023 · 1 repository · arXiv:2311.10873
-
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers 17 Nov 2023 · 0 repositories · arXiv:2311.10642
-
Semi-supervised ViT knowledge distillation network with style transfer normalization for colorectal liver metastases survival prediction 17 Nov 2023 · 0 repositories · arXiv:2311.10305
-
TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in LLMs through Translation-Assisted Chain-of-Thought Processes 17 Nov 2023 · 1 repository · arXiv:2311.10797
-
ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems 16 Nov 2023 · 1 repository · arXiv:2311.09476Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 15 harvested samples)
-
BLT: Can Large Language Models Handle Basic Legal Text? 16 Nov 2023 · 1 repository · arXiv:2311.09693
-
Co-data Learning for Bayesian Additive Regression Trees 16 Nov 2023 · 1 repository · arXiv:2311.09997
-
Prompt-based Pseudo-labeling Strategy for Sample-Efficient Semi-Supervised Extractive Summarization 16 Nov 2023 · 0 repositories · arXiv:2311.09559
-
Event Causality Is Key to Computational Story Understanding 16 Nov 2023 · 1 repository · arXiv:2311.09648Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Fumbling in Babel: An Investigation into ChatGPT's Language Identification Ability 16 Nov 2023 · 0 repositories · arXiv:2311.09696
-
GEE! Grammar Error Explanation with Large Language Models 16 Nov 2023 · 1 repository · arXiv:2311.09517
-
Generative AI for Hate Speech Detection: Evaluation and Findings 16 Nov 2023 · 0 repositories · arXiv:2311.09993
-
HuatuoGPT-II, One-stage Training for Medical Adaption of LLMs 16 Nov 2023 · 1 repository · arXiv:2311.09774Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 14 harvested samples) · 14 pointer-only (licence)
-
Human Still Wins over LLM: An Empirical Study of Active Learning on Domain-Specific Annotation Tasks 16 Nov 2023 · 0 repositories · arXiv:2311.09825
-
INTERVENOR: Prompting the Coding Ability of Large Language Models with the Interactive Chain of Repair 16 Nov 2023 · 1 repository · arXiv:2311.09868
-
Investigating Data Contamination in Modern Benchmarks for Large Language Models 16 Nov 2023 · 0 repositories · arXiv:2311.09783
-
FinanceMath: Knowledge-Intensive Math Reasoning in Finance Domains 16 Nov 2023 · 1 repository · arXiv:2311.09797Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 15 harvested samples) · 15 pointer-only (licence)
-
Large Language Models for Propaganda Span Annotation 16 Nov 2023 · 1 repository · arXiv:2311.09812
-
MARformer: An Efficient Metal Artifact Reduction Transformer for Dental CBCT Images 16 Nov 2023 · 0 repositories · arXiv:2311.09590
-
ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code 16 Nov 2023 · 1 repository · arXiv:2311.09835Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Multi-View Spectrogram Transformer for Respiratory Sound Classification 16 Nov 2023 · 1 repository · arXiv:2311.09655
-
Neural-Logic Human-Object Interaction Detection 16 Nov 2023 · 1 repository · arXiv:2311.09817Syntology 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
On Evaluating the Integration of Reasoning and Action in LLM Agents with Database Question Answering 16 Nov 2023 · 0 repositories · arXiv:2311.09721
-
On Retrieval Augmentation and the Limitations of Language Model Training 16 Nov 2023 · 0 repositories · arXiv:2311.09615
-
Predictive Minds: LLMs As Atypical Active Inference Agents 16 Nov 2023 · 0 repositories · arXiv:2311.10215
-
ConceptPsy:A Benchmark Suite with Conceptual Comprehensiveness in Psychology 16 Nov 2023 · 0 repositories · arXiv:2311.09861
-
Reducing Privacy Risks in Online Self-Disclosures with Language Models 16 Nov 2023 · 0 repositories · arXiv:2311.09538
-
Self-Contradictory Reasoning Evaluation and Detection 16 Nov 2023 · 1 repository · arXiv:2311.09603Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Structured Chemistry Reasoning with Large Language Models 16 Nov 2023 · 1 repository · arXiv:2311.09656Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
SurvTimeSurvival: Survival Analysis On The Patient With Multiple Visits/Records 16 Nov 2023 · 1 repository · arXiv:2311.09854
-
Chemist-X: Large Language Model-empowered Agent for Reaction Condition Recommendation in Chemical Synthesis 16 Nov 2023 · 0 repositories · arXiv:2311.10776
-
Towards Autonomous Hypothesis Verification via Language Models with Minimal Guidance 16 Nov 2023 · 0 repositories · arXiv:2311.09706
-
UnifiedVisionGPT: Streamlining Vision-Oriented AI through Generalized Multimodal Framework 16 Nov 2023 · 1 repository · arXiv:2311.10125
-
Wildfire Smoke Detection with Cross Contrast Patch Embedding 16 Nov 2023 · 1 repository · arXiv:2311.10116
-
Can Large Language Models Follow Concept Annotation Guidelines? A Case Study on Scientific and Financial Domains 15 Nov 2023 · 1 repository · arXiv:2311.08704
-
Comparing Generalization in Learning with Limited Numbers of Exemplars: Transformer vs. RNN in Attractor Dynamics 15 Nov 2023 · 0 repositories · arXiv:2311.10763
-
Contrastive Transformer Learning with Proximity Data Generation for Text-Based Person Search 15 Nov 2023 · 1 repository · arXiv:2311.09084Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Deep Group Interest Modeling of Full Lifelong User Behaviors for CTR Prediction 15 Nov 2023 · 0 repositories · arXiv:2311.10764
-
Degradation Estimation Recurrent Neural Network with Local and Non-Local Priors for Compressive Spectral Imaging 15 Nov 2023 · 1 repository · arXiv:2311.08808
-
DISTA: Denoising Spiking Transformer with intrinsic plasticity and spatiotemporal attention 15 Nov 2023 · 0 repositories · arXiv:2311.09376
-
Enhancing Machine Translation through Advanced In-Context Learning: A Methodological Strategy for GPT-4 Improvement 15 Nov 2023 · 0 repositories · arXiv:2311.10765
-
Evaluating Gender Bias in the Translation of Gender-Neutral Languages into English 15 Nov 2023 · 0 repositories · arXiv:2311.08836
-
Factcheck-Bench: Fine-Grained Evaluation Benchmark for Automatic Fact-checkers 15 Nov 2023 · 2 repositories · arXiv:2311.09000Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Generalizable Imitation Learning Through Pre-Trained Representations 15 Nov 2023 · 0 repositories · arXiv:2311.09350
-
GENEVA: GENErating and Visualizing branching narratives using LLMs 15 Nov 2023 · 0 repositories · arXiv:2311.09213
-
I Was Blind but Now I See: Implementing Vision-Enabled Dialogue in Social Robots 15 Nov 2023 · 1 repository · arXiv:2311.08957
-
Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts 15 Nov 2023 · 0 repositories · arXiv:2311.09127
-
Llamas Know What GPTs Don't Show: Surrogate Models for Confidence Estimation 15 Nov 2023 · 0 repositories · arXiv:2311.08877