Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 79
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 79 of 190: papers 7,801 to 7,900 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Evaluation of ChatGPT's Smart Contract Auditing Capabilities Based on Chain of Thought 19 Feb 2024 · 0 repositories · arXiv:2402.12023
-
FeB4RAG: Evaluating Federated Search in the Context of Retrieval Augmented Generation 19 Feb 2024 · 0 repositories · arXiv:2402.11891
-
FiT: Flexible Vision Transformer for Diffusion Model 19 Feb 2024 · 2 repositories · arXiv:2402.12376Syntology official (archive's flag): 15 ran · 15 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 1 honoured, 0 violated, 10 with no contract checked; 4 where Syntology's instrument failed) · 2 unverified (of 17 harvested samples) · 3 pointer-only (licence)
-
Graph-Based Retriever Captures the Long Tail of Biomedical Knowledge 19 Feb 2024 · 0 repositories · arXiv:2402.12352
-
GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations 19 Feb 2024 · 2 repositories · arXiv:2402.12348
-
IMBUE: Improving Interpersonal Effectiveness through Simulation and Just-in-time Feedback with Human-Language Model Interaction 19 Feb 2024 · 0 repositories · arXiv:2402.12556
-
Is Open-Source There Yet? A Comparative Study on Commercial and Open-Source LLMs in Their Ability to Label Chest X-Ray Reports 19 Feb 2024 · 0 repositories · arXiv:2402.12298
-
Key ingredients for effective zero-shot cross-lingual knowledge transfer in generative tasks 19 Feb 2024 · 0 repositories · arXiv:2402.12279
-
Locality-Sensitive Hashing-Based Efficient Point Transformer with Applications in High-Energy Physics 19 Feb 2024 · 1 repository · arXiv:2402.12535Syntology official (archive's flag): 12 ran · 12 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 11 where Syntology's instrument failed) · 3 unverified (of 15 harvested samples)
-
Mafin: Enhancing Black-Box Embeddings with Model Augmented Fine-Tuning 19 Feb 2024 · 0 repositories · arXiv:2402.12177
-
Enabling Weak LLMs to Judge Response Reliability via Meta Ranking 19 Feb 2024 · 0 repositories · arXiv:2402.12146
-
Cofca: A Step-Wise Counterfactual Multi-hop QA benchmark 19 Feb 2024 · 0 repositories · arXiv:2402.11924
-
Perceiving Longer Sequences With Bi-Directional Cross-Attention Transformers 19 Feb 2024 · 2 repositories · arXiv:2402.12138Syntology official (archive's flag): 16 ran · 17 ran (of which 11 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 6 where Syntology's instrument failed) · 11 unverified (of 28 harvested samples) · 27 pointer-only (licence)
-
Query-Based Adversarial Prompt Generation 19 Feb 2024 · 2 repositories · arXiv:2402.12329
-
Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models 19 Feb 2024 · 1 repository · arXiv:2402.12336Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 6 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples)
-
Shallow Synthesis of Knowledge in GPT-Generated Texts: A Case Study in Automatic Related Work Composition 19 Feb 2024 · 0 repositories · arXiv:2402.12255
-
SPML: A DSL for Defending Language Models Against Prompt Attacks 19 Feb 2024 · 0 repositories · arXiv:2402.11755
-
Standardize: Aligning Language Models with Expert-Defined Standards for Content Generation 19 Feb 2024 · 1 repository · arXiv:2402.12593
-
Stealing the Invisible: Unveiling Pre-Trained CNN Models through Adversarial Examples and Timing Side-Channels 19 Feb 2024 · 0 repositories · arXiv:2402.11953
-
Stick to your Role! Stability of Personal Values Expressed in Large Language Models 19 Feb 2024 · 0 repositories · arXiv:2402.14846
-
CiMNet: Towards Joint Optimization for DNN Architecture and Configuration for Compute-In-Memory Hardware 19 Feb 2024 · 0 repositories · arXiv:2402.11780
-
What Evidence Do Language Models Find Convincing? 19 Feb 2024 · 1 repository · arXiv:2402.11782Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified (of 14 harvested samples)
-
Your Large Language Model is Secretly a Fairness Proponent and You Should Prompt it Like One 19 Feb 2024 · 0 repositories · arXiv:2402.12150
-
A Curious Case of Searching for the Correlation between Training Data and Adversarial Robustness of Transformer Textual Models 18 Feb 2024 · 1 repository · arXiv:2402.11469
-
Can Deception Detection Go Deeper? Dataset, Evaluation, and Benchmark for Deception Reasoning 18 Feb 2024 · 0 repositories · arXiv:2402.11432
-
Decoding News Narratives: A Critical Analysis of Large Language Models in Framing Detection 18 Feb 2024 · 0 repositories · arXiv:2402.11621
-
DictLLM: Harnessing Key-Value Data Structures with Large Language Models for Enhanced Medical Diagnostics 18 Feb 2024 · 0 repositories · arXiv:2402.11481
-
EventRL: Enhancing Event Extraction with Outcome Supervision for Large Language Models 18 Feb 2024 · 1 repository · arXiv:2402.11430
-
FactPICO: Factuality Evaluation for Plain Language Summarization of Medical Evidence 18 Feb 2024 · 1 repository · arXiv:2402.11456Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
GNNavi: Navigating the Information Flow in Large Language Models by Graph Neural Network 18 Feb 2024 · 1 repository · arXiv:2402.11709
-
KMMLU: Measuring Massive Multitask Language Understanding in Korean 18 Feb 2024 · 0 repositories · arXiv:2402.11548
-
Learning From Failure: Integrating Negative Examples when Fine-tuning Large Language Models as Agents 18 Feb 2024 · 1 repository · arXiv:2402.11651Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
LongAgent: Scaling Language Models to 128k Context through Multi-Agent Collaboration 18 Feb 2024 · 1 repository · arXiv:2402.11550
-
MSynFD: Multi-hop Syntax aware Fake News Detection 18 Feb 2024 · 0 repositories · arXiv:2402.14834
-
Multi-dimensional Evaluation of Empathetic Dialog Responses 18 Feb 2024 · 0 repositories · arXiv:2402.11409
-
Multi-Task Inference: Can Large Language Models Follow Multiple Instructions at Once? 18 Feb 2024 · 1 repository · arXiv:2402.11597Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement 18 Feb 2024 · 1 repository · arXiv:2402.11436Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Ploutos: Towards interpretable stock movement prediction with financial large language model 18 Feb 2024 · 0 repositories · arXiv:2403.00782
-
Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning 18 Feb 2024 · 0 repositories · arXiv:2402.11690
-
Boosting of Thoughts: Trial-and-Error Problem Solving with Large Language Models 17 Feb 2024 · 2 repositories · arXiv:2402.11140
-
Detecting a Proxy for Potential Comorbid ADHD in People Reporting Anxiety Symptoms from Social Media Data 17 Feb 2024 · 0 repositories · arXiv:2403.05561
-
Exploring ChatGPT for Next-generation Information Retrieval: Opportunities and Challenges 17 Feb 2024 · 0 repositories · arXiv:2402.11203
-
GenDec: A robust generative Question-decomposition method for Multi-hop reasoning 17 Feb 2024 · 0 repositories · arXiv:2402.11166
-
Reasoning before Comparison: LLM-Enhanced Semantic Similarity Metrics for Domain Specialized Text Analysis 17 Feb 2024 · 0 repositories · arXiv:2402.11398
-
ReViT: Enhancing Vision Transformers Feature Diversity with Attention Residual Connections 17 Feb 2024 · 1 repository · arXiv:2402.11301Syntology official (archive's flag): 22 ran · 22 ran (of which 0 constructed an object rather than computing a result; 20 with no instrument failure: 0 honoured, 0 violated, 20 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 24 harvested samples) · 1 pointer-only (licence)
-
ZeroG: Investigating Cross-dataset Zero-shot Transferability in Graphs 17 Feb 2024 · 1 repository · arXiv:2402.11235Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Assessing the Reasoning Abilities of ChatGPT in the Context of Claim Verification 16 Feb 2024 · 0 repositories · arXiv:2402.10735
-
Can LLMs Speak For Diverse People? Tuning LLMs via Debate to Generate Controllable Controversial Statements 16 Feb 2024 · 1 repository · arXiv:2402.10614Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Can Separators Improve Chain-of-Thought Prompting? 16 Feb 2024 · 0 repositories · arXiv:2402.10645
-
ContiFormer: Continuous-Time Transformer for Irregular Time Series Modeling 16 Feb 2024 · 1 repository · arXiv:2402.10635Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 12 harvested samples)
-
Disordered-DABS: A Benchmark for Dynamic Aspect-Based Summarization in Disordered Texts 16 Feb 2024 · 1 repository · arXiv:2402.10554
-
Dynamic Patch-aware Enrichment Transformer for Occluded Person Re-Identification 16 Feb 2024 · 0 repositories · arXiv:2402.10435
-
Emoji Driven Crypto Assets Market Reactions 16 Feb 2024 · 0 repositories · arXiv:2402.10481
-
FinTral: A Family of GPT-4 Level Multimodal Financial Large Language Models 16 Feb 2024 · 0 repositories · arXiv:2402.10986
-
German Text Simplification: Finetuning Large Language Models with Semi-Synthetic Data 16 Feb 2024 · 1 repository · arXiv:2402.10675
-
How Reliable Are Automatic Evaluation Methods for Instruction-Tuned LLMs? 16 Feb 2024 · 0 repositories · arXiv:2402.10770
-
In Search of Needles in a 11M Haystack: Recurrent Memory Finds What LLMs Miss 16 Feb 2024 · 2 repositories · arXiv:2402.10790
-
Inference to the Best Explanation in Large Language Models 16 Feb 2024 · 0 repositories · arXiv:2402.10767
-
When "Competency" in Reasoning Opens the Door to Vulnerability: Jailbreaking LLMs via Novel Complex Ciphers 16 Feb 2024 · 1 repository · arXiv:2402.10601Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Large Language Models as Zero-shot Dialogue State Tracker through Function Calling 16 Feb 2024 · 1 repository · arXiv:2402.10466Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Large Language Models Fall Short: Understanding Complex Relationships in Detective Narratives 16 Feb 2024 · 0 repositories · arXiv:2402.11051
-
Linear Transformers with Learnable Kernel Functions are Better In-Context Models 16 Feb 2024 · 2 repositories · arXiv:2402.10644Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
LinkNER: Linking Local Named Entity Recognition Models to Large Language Models using Uncertainty 16 Feb 2024 · 1 repository · arXiv:2402.10573
-
LLMs in the Heart of Differential Testing: A Case Study on a Medical Rule Engine 16 Feb 2024 · 0 repositories · arXiv:2404.03664
-
Cultural Commonsense Knowledge for Intercultural Dialogues 16 Feb 2024 · 0 repositories · arXiv:2402.10689
-
Network Formation and Dynamics Among Multi-LLMs 16 Feb 2024 · 1 repository · arXiv:2402.10659
-
PRISE: LLM-Style Sequence Compression for Learning Temporal Action Abstractions in Control 16 Feb 2024 · 1 repository · arXiv:2402.10450Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 11 harvested samples)
-
ToolSword: Unveiling Safety Issues of Large Language Models in Tool Learning Across Three Stages 16 Feb 2024 · 1 repository · arXiv:2402.10753
-
Universal Prompt Optimizer for Safe Text-to-Image Generation 16 Feb 2024 · 1 repository · arXiv:2402.10882
-
Where is the answer? Investigating Positional Bias in Language Model Knowledge Extraction 16 Feb 2024 · 1 repository · arXiv:2402.12170
-
Weak-Mamba-UNet: Visual Mamba Makes CNN and ViT Work Better for Scribble-based Medical Image Segmentation 16 Feb 2024 · 2 repositories · arXiv:2402.10887Syntology official (archive's flag): 4 ran · 5 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 4 pointer-only (licence)
-
A StrongREJECT for Empty Jailbreaks 15 Feb 2024 · 2 repositories · arXiv:2402.10260Syntology official: harvested, nothing ran · 0 ran · 4 unverified (of 4 harvested samples)
-
An Analysis of Language Frequency and Error Correction for Esperanto 15 Feb 2024 · 0 repositories · arXiv:2402.09696
-
Efficient Prompt Optimization Through the Lens of Best Arm Identification 15 Feb 2024 · 0 repositories · arXiv:2402.09723
-
Camouflage is all you need: Evaluating and Enhancing Language Model Robustness Against Camouflage Adversarial Attacks 15 Feb 2024 · 0 repositories · arXiv:2402.09874
-
Data Engineering for Scaling Language Models to 128K Context 15 Feb 2024 · 1 repository · arXiv:2402.10171Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Fine-tuning Large Language Model (LLM) Artificial Intelligence Chatbots in Ophthalmology and LLM-based evaluation using GPT-4 15 Feb 2024 · 0 repositories · arXiv:2402.10083
-
GPT-4's assessment of its performance in a USMLE-based case study 15 Feb 2024 · 0 repositories · arXiv:2402.09654
-
Grounding Language Model with Chunking-Free In-Context Retrieval 15 Feb 2024 · 0 repositories · arXiv:2402.09760
-
Improving Non-autoregressive Machine Translation with Error Exposure and Consistency Regularization 15 Feb 2024 · 0 repositories · arXiv:2402.09725
-
Large Language Models for Forecasting and Anomaly Detection: A Systematic Literature Review 15 Feb 2024 · 0 repositories · arXiv:2402.10350
-
NYCTALE: Neuro-Evidence Transformer for Adaptive and Personalized Lung Nodule Invasiveness Prediction 15 Feb 2024 · 0 repositories · arXiv:2402.10066
-
OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset 15 Feb 2024 · 1 repository · arXiv:2402.10176
-
PAL: Proxy-Guided Black-Box Attack on Large Language Models 15 Feb 2024 · 1 repository · arXiv:2402.09674
-
Prompt-Based Bias Calibration for Better Zero/Few-Shot Learning of Language Models 15 Feb 2024 · 0 repositories · arXiv:2402.10353
-
ProtChatGPT: Towards Understanding Proteins with Large Language Models 15 Feb 2024 · 0 repositories · arXiv:2402.09649
-
Spike-driven Transformer V2: Meta Spiking Neural Network Architecture Inspiring the Design of Next-generation Neuromorphic Chips 15 Feb 2024 · 1 repository · arXiv:2404.03663Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 1 honoured, 0 violated, 8 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
The Butterfly Effect of Model Editing: Few Edits Can Trigger Large Language Models Collapse 15 Feb 2024 · 1 repository · arXiv:2402.09656
-
Unlocking Structure Measuring: Introducing PDD, an Automatic Metric for Positional Discourse Coherence 15 Feb 2024 · 1 repository · arXiv:2402.10175
-
X-lifecycle Learning for Cloud Incident Management using LLMs 15 Feb 2024 · 0 repositories · arXiv:2404.03662
-
API Pack: A Massive Multi-Programming Language Dataset for API Call Generation 14 Feb 2024 · 1 repository · arXiv:2402.09615
-
AQA-Bench: An Interactive Benchmark for Evaluating LLMs' Sequential Reasoning Ability 14 Feb 2024 · 1 repository · arXiv:2402.09404
-
Bidirectional Generative Pre-training for Improving Healthcare Time-series Representation Learning 14 Feb 2024 · 1 repository · arXiv:2402.09558
-
Changes by Butterflies: Farsighted Forecasting with Group Reservoir Transformer 14 Feb 2024 · 0 repositories · arXiv:2402.09573
-
Context Composing for Full Line Code Completion 14 Feb 2024 · 0 repositories · arXiv:2402.09230
-
Emerging Opportunities of Using Large Language Models for Translation Between Drug Molecules and Indications 14 Feb 2024 · 0 repositories · arXiv:2402.09588
-
FGeo-TP: A Language Model-Enhanced Solver for Geometry Problems 14 Feb 2024 · 0 repositories · arXiv:2402.09047
-
HEAL-ViT: Vision Transformers on a spherical mesh for medium-range weather forecasting 14 Feb 2024 · 0 repositories · arXiv:2403.17016
-
L3GO: Language Agents with Chain-of-3D-Thoughts for Generating Unconventional Objects 14 Feb 2024 · 0 repositories · arXiv:2402.09052
-
Leveraging Large Language Models for Enhanced NLP Task Performance through Knowledge Distillation and Optimized Training Strategies 14 Feb 2024 · 0 repositories · arXiv:2402.09282