Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 85
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 85 of 190: papers 8,401 to 8,500 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Killer Apps: Low-Speed, Large-Scale AI Weapons 14 Jan 2024 · 0 repositories · arXiv:2402.01663
-
Learning to be Homo Economicus: Can an LLM Learn Preferences from Choice 14 Jan 2024 · 0 repositories · arXiv:2401.07345
-
MapGPT: Map-Guided Prompting with Adaptive Path Planning for Vision-and-Language Navigation 14 Jan 2024 · 0 repositories · arXiv:2401.07314
-
Streamlining the Selection Phase of Systematic Literature Reviews (SLRs) Using AI-Enabled GPT-4 Assistant API 14 Jan 2024 · 0 repositories · arXiv:2402.18582
-
A Novel Multi-Stage Prompting Approach for Language Agnostic MCQ Generation using GPT 13 Jan 2024 · 1 repository · arXiv:2401.07098
-
Assessing Large Language Models in Mechanical Engineering Education: A Study on Mechanics-Focused Conceptual Understanding 13 Jan 2024 · 0 repositories · arXiv:2401.12983
-
Bridging the Preference Gap between Retrievers and LLMs 13 Jan 2024 · 0 repositories · arXiv:2401.06954
-
Combining Confidence Elicitation and Sample-based Methods for Uncertainty Quantification in Misinformation Mitigation 13 Jan 2024 · 0 repositories · arXiv:2401.08694
-
GoMatching: A Simple Baseline for Video Text Spotting via Long and Short Term Matching 13 Jan 2024 · 1 repository · arXiv:2401.07080
-
Knowledge Distillation of Black-Box Large Language Models 13 Jan 2024 · 0 repositories · arXiv:2401.07013
-
Transformer for Object Re-Identification: A Survey 13 Jan 2024 · 1 repository · arXiv:2401.06960Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
A Survey on the Applications of Frontier AI, Foundation Models, and Large Language Models to Intelligent Transportation Systems 12 Jan 2024 · 0 repositories · arXiv:2401.06831
-
Adapting Large Language Models for Document-Level Machine Translation 12 Jan 2024 · 0 repositories · arXiv:2401.06468
-
Comparing GPT-4 and Open-Source Language Models in Misinformation Mitigation 12 Jan 2024 · 0 repositories · arXiv:2401.06920
-
Domain Adaptation for Time series Transformers using One-step fine-tuning 12 Jan 2024 · 0 repositories · arXiv:2401.06524
-
Few-Shot Detection of Machine-Generated Text using Style Representations 12 Jan 2024 · 1 repository · arXiv:2401.06712Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
Fine-grained Hallucination Detection and Editing for Language Models 12 Jan 2024 · 0 repositories · arXiv:2401.06855
-
Human-AI Collaborative Essay Scoring: A Dual-Process Framework with LLMs 12 Jan 2024 · 1 repository · arXiv:2401.06431
-
Health-LLM: Large Language Models for Health Prediction via Wearable Sensor Data 12 Jan 2024 · 1 repository · arXiv:2401.06866Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs 12 Jan 2024 · 2 repositories · arXiv:2401.06373Syntology official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 2 pointer-only (licence)
-
Hyper-STTN: Social Group-aware Spatial-Temporal Transformer Network for Human Trajectory Prediction with Hypergraph Reasoning 12 Jan 2024 · 0 repositories · arXiv:2401.06344
-
Intention Analysis Makes LLMs A Good Jailbreak Defender 12 Jan 2024 · 1 repository · arXiv:2401.06561Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Mapping Transformer Leveraged Embeddings for Cross-Lingual Document Representation 12 Jan 2024 · 1 repository · arXiv:2401.06583
-
Mission: Impossible Language Models 12 Jan 2024 · 1 repository · arXiv:2401.06416Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 12 harvested samples)
-
PersianMind: A Cross-Lingual Persian-English Large Language Model 12 Jan 2024 · 0 repositories · arXiv:2401.06466
-
PizzaCommonSense: Learning to Model Commonsense Reasoning about Intermediate Steps in Cooking Recipes 12 Jan 2024 · 1 repository · arXiv:2401.06930
-
Video Super-Resolution Transformer with Masked Inter&Intra-Frame Attention 12 Jan 2024 · 1 repository · arXiv:2401.06312
-
Autocompletion of Chief Complaints in the Electronic Health Records using Large Language Models 11 Jan 2024 · 0 repositories · arXiv:2401.06088
-
Brain Tumor Radiogenomic Classification 11 Jan 2024 · 0 repositories · arXiv:2401.09471
-
Efficient Image Deblurring Networks based on Diffusion Models 11 Jan 2024 · 1 repository · arXiv:2401.05907Syntology official (archive's flag): 14 ran · 14 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 2 violated, 8 with no contract checked; 4 where Syntology's instrument failed) · 3 unverified (of 17 harvested samples) · 17 pointer-only (licence)
-
Evidence to Generate (E2G): A Single-agent Two-step Prompting for Context Grounded and Retrieval Augmented Reasoning 11 Jan 2024 · 0 repositories · arXiv:2401.05787
-
Investigating Data Contamination for Pre-training Language Models 11 Jan 2024 · 0 repositories · arXiv:2401.06059
-
Masked Attribute Description Embedding for Cloth-Changing Person Re-identification 11 Jan 2024 · 1 repository · arXiv:2401.05646
-
Mutation-based Consistency Testing for Evaluating the Code Understanding Capability of LLMs 11 Jan 2024 · 0 repositories · arXiv:2401.05940
-
Prompt-based mental health screening from social media text 11 Jan 2024 · 0 repositories · arXiv:2401.05912
-
Surface Normal Estimation with Transformers 11 Jan 2024 · 0 repositories · arXiv:2401.05745
-
The Benefits of a Concise Chain of Thought on Problem-Solving in Large Language Models 11 Jan 2024 · 1 repository · arXiv:2401.05618Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Transforming Image Super-Resolution: A ConvFormer-based Efficient Approach 11 Jan 2024 · 1 repository · arXiv:2401.05633
-
AdvMT: Adversarial Motion Transformer for Long-term Human Motion Prediction 10 Jan 2024 · 0 repositories · arXiv:2401.05018
-
Deep learning in motion deblurring: current status, benchmarks and future prospects 10 Jan 2024 · 1 repository · arXiv:2401.05055
-
AutoAct: Automatic Agent Learning from Scratch for QA via Self-Planning 10 Jan 2024 · 1 repository · arXiv:2401.05268Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples)
-
CADgpt: Harnessing Natural Language Processing for 3D Modelling to Enhance Computer-Aided Design Workflows 10 Jan 2024 · 0 repositories · arXiv:2401.05476
-
Can AI Write Classical Chinese Poetry like Humans? An Empirical Study Inspired by Turing Test 10 Jan 2024 · 0 repositories · arXiv:2401.04952
-
Diffusion-based Pose Refinement and Muti-hypothesis Generation for 3D Human Pose Estimaiton 10 Jan 2024 · 1 repository · arXiv:2401.04921
-
Efficient Fine-Tuning with Domain Adaptation for Privacy-Preserving Vision Transformer 10 Jan 2024 · 0 repositories · arXiv:2401.05126
-
I am a Strange Dataset: Metalinguistic Tests for Language Models 10 Jan 2024 · 1 repository · arXiv:2401.05300
-
InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks 10 Jan 2024 · 1 repository · arXiv:2401.05507Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Knowledge-aware Graph Transformer for Pedestrian Trajectory Prediction 10 Jan 2024 · 0 repositories · arXiv:2401.04872
-
Knowledge Sharing in Manufacturing using Large Language Models: User Evaluation and Model Benchmarking 10 Jan 2024 · 0 repositories · arXiv:2401.05200
-
Leveraging Print Debugging to Improve Code Generation in Large Language Models 10 Jan 2024 · 0 repositories · arXiv:2401.05319
-
Can Active Label Correction Improve LLM-based Modular AI Systems? 10 Jan 2024 · 0 repositories · arXiv:2401.05467
-
Monte Carlo Tree Search for Recipe Generation using GPT-2 10 Jan 2024 · 0 repositories · arXiv:2401.05199
-
Motion Guided Token Compression for Efficient Masked Video Modeling 10 Jan 2024 · 0 repositories · arXiv:2402.18577
-
Reinforcement Learning for Optimizing RAG for Domain Chatbots 10 Jan 2024 · 0 repositories · arXiv:2401.06800
-
SPT: Spectral Transformer for Red Giant Stars Age and Mass Estimation 10 Jan 2024 · 0 repositories · arXiv:2401.04900
-
An Assessment on Comprehending Mental Health through Large Language Models 9 Jan 2024 · 0 repositories · arXiv:2401.04592
-
Arabic Text Diacritization In The Age Of Transfer Learning: Token Classification Is All You Need 9 Jan 2024 · 0 repositories · arXiv:2401.04848
-
DebugBench: Evaluating Debugging Capability of Large Language Models 9 Jan 2024 · 1 repository · arXiv:2401.04621Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 8 harvested samples)
-
DedustNet: A Frequency-dominated Swin Transformer-based Wavelet Network for Agricultural Dust Removal 9 Jan 2024 · 0 repositories · arXiv:2401.04750
-
DepressionEmo: A novel dataset for multilabel classification of depression emotions 9 Jan 2024 · 1 repository · arXiv:2401.04655
-
Fighting Fire with Fire: Adversarial Prompting to Generate a Misinformation Detection Dataset 9 Jan 2024 · 0 repositories · arXiv:2401.04481
-
Informed AI Regulation: Comparing the Ethical Frameworks of Leading LLM Chatbots Using an Ethics-Based Audit to Assess Moral Reasoning and Normative Values 9 Jan 2024 · 1 repository · arXiv:2402.01651
-
Iterative Feedback Network for Unsupervised Point Cloud Registration 9 Jan 2024 · 1 repository · arXiv:2401.04357
-
Phishing Website Detection through Multi-Model Analysis of HTML Content 9 Jan 2024 · 0 repositories · arXiv:2401.04820
-
Setting the Record Straight on Transformer Oversmoothing 9 Jan 2024 · 0 repositories · arXiv:2401.04301
-
Skin Cancer Segmentation and Classification Using Vision Transformer for Automatic Analysis in Dermatoscopy-based Non-invasive Digital System 9 Jan 2024 · 0 repositories · arXiv:2401.04746
-
T-PRIME: Transformer-based Protocol Identification for Machine-learning at the Edge 9 Jan 2024 · 1 repository · arXiv:2401.04837
-
WaveletFormerNet: A Transformer-based Wavelet Network for Real-world Non-homogeneous and Dense Fog Removal 9 Jan 2024 · 0 repositories · arXiv:2401.04550
-
A Philosophical Introduction to Language Models -- Part I: Continuity With Classic Debates 8 Jan 2024 · 0 repositories · arXiv:2401.03910
-
Advancing Spatial Reasoning in Large Language Models: An In-Depth Evaluation and Enhancement Using the StepGame Benchmark 8 Jan 2024 · 1 repository · arXiv:2401.03991
-
Can Large Language Models Beat Wall Street? Unveiling the Potential of AI in Stock Selection 8 Jan 2024 · 0 repositories · arXiv:2401.03737
-
Distortions in Judged Spatial Relations in Large Language Models 8 Jan 2024 · 0 repositories · arXiv:2401.04218
-
Efficient Multiscale Multimodal Bottleneck Transformer for Audio-Video Classification 8 Jan 2024 · 0 repositories · arXiv:2401.04023
-
Efficient Selective Audio Masked Multimodal Bottleneck Transformer for Audio-Video Classification 8 Jan 2024 · 0 repositories · arXiv:2401.04154
-
GloTSFormer: Global Video Text Spotting Transformer 8 Jan 2024 · 1 repository · arXiv:2401.03694
-
Gramformer: Learning Crowd Counting via Graph-Modulated Transformer 8 Jan 2024 · 1 repository · arXiv:2401.03870Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Advancing bioinformatics with large language models: components, applications and perspectives 8 Jan 2024 · 0 repositories · arXiv:2401.04155
-
LF-ViT: Reducing Spatial Redundancy in Vision Transformer for Efficient Image Recognition 8 Jan 2024 · 1 repository · arXiv:2402.00033Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 3 where Syntology's instrument failed) · 7 unverified (of 18 harvested samples) · 18 pointer-only (licence)
-
LLM4PLC: Harnessing Large Language Models for Verifiable Programming of PLCs in Industrial Control Systems 8 Jan 2024 · 1 repository · arXiv:2401.05443
-
MARG: Multi-Agent Review Generation for Scientific Papers 8 Jan 2024 · 1 repository · arXiv:2401.04259Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples)
-
Mixtral of Experts 8 Jan 2024 · 6 repositories · arXiv:2401.04088Syntology 5 ran (of which 5 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; every one of the 5 samples that ran constructed an object rather than computing a result (of 5 harvested samples)
-
MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts 8 Jan 2024 · 1 repository · arXiv:2401.04081
-
MS-DETR: Efficient DETR Training with Mixed Supervision 8 Jan 2024 · 1 repository · arXiv:2401.03989Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 1 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples) · 3 pointer-only (licence)
-
Why Solving Multi-agent Path Finding with Large Language Model has not Succeeded Yet 8 Jan 2024 · 0 repositories · arXiv:2401.03630
-
Can generative AI and ChatGPT outperform humans on cognitive-demanding problem-solving tasks in science? 7 Jan 2024 · 0 repositories · arXiv:2401.15081
-
EAT: Self-Supervised Pre-Training with Efficient Audio Transformer 7 Jan 2024 · 1 repository · arXiv:2401.03497Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 17 harvested samples) · 3 pointer-only (licence)
-
Escalation Risks from Language Models in Military and Diplomatic Decision-Making 7 Jan 2024 · 1 repository · arXiv:2401.03408Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
InFoBench: Evaluating Instruction Following Ability in Large Language Models 7 Jan 2024 · 1 repository · arXiv:2401.03601Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
On Leveraging Large Language Models for Enhancing Entity Resolution: A Cost-efficient Approach 7 Jan 2024 · 0 repositories · arXiv:2401.03426
-
RoBERTurk: Adjusting RoBERTa for Turkish 7 Jan 2024 · 0 repositories · arXiv:2401.03515
-
See360: Novel Panoramic View Interpolation 7 Jan 2024 · 1 repository · arXiv:2401.03431
-
The NPU-ASLP-LiAuto System Description for Visual Speech Recognition in CNVSRC 2023 7 Jan 2024 · 2 repositories · arXiv:2401.06788
-
CharPoet: A Chinese Classical Poetry Generation System Based on Token-free LLM 7 Jan 2024 · 0 repositories · arXiv:2401.03512
-
Towards Effective Multiple-in-One Image Restoration: A Sequential and Prompt Learning Strategy 7 Jan 2024 · 2 repositories · arXiv:2401.03379Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 1 violated, 8 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
Exploring Defeasibility in Causal Reasoning 6 Jan 2024 · 0 repositories · arXiv:2401.03183
-
Multimodal Informative ViT: Information Aggregation and Distribution for Hyperspectral and LiDAR Classification 6 Jan 2024 · 1 repository · arXiv:2401.03179
-
PIXAR: Auto-Regressive Language Modeling in Pixel Space 6 Jan 2024 · 0 repositories · arXiv:2401.03321
-
PosDiffNet: Positional Neural Diffusion for Point Cloud Registration in a Large Field of View with Perturbations 6 Jan 2024 · 0 repositories · arXiv:2401.03167
-
Realism in Action: Anomaly-Aware Diagnosis of Brain Tumors from Medical Images Using YOLOv8 and DeiT 6 Jan 2024 · 0 repositories · arXiv:2401.03302
-
SecureReg: Combining NLP and MLP for Enhanced Detection of Malicious Domain Name Registrations 6 Jan 2024 · 0 repositories · arXiv:2401.03196