Methods › Natural Language Processing › Subword Segmentation › BPE › Papers, page 38
Byte Pair Encoding
BPE
Papers archive 2025-07-28
archive papers tagged: 18,975 · with a code link: 8,675 · where Syntology ran a sample: 2,895 (2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,895 of 18,975 tagged: 2,443 with a run with no instrument failure, 452 where every run was a failure of Syntology's instrument)
Page 38 of 190: papers 3,701 to 3,800 of 18,975, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Performance Evaluation of Deep Learning and Transformer Models Using Multimodal Data for Breast Cancer Classification 14 Oct 2024 · 0 repositories · arXiv:2410.10146
-
Performance in a dialectal profiling task of LLMs for varieties of Brazilian Portuguese 14 Oct 2024 · 0 repositories · arXiv:2410.10991
-
Rethinking Legal Judgement Prediction in a Realistic Scenario in the Era of Large Language Models 14 Oct 2024 · 1 repository · arXiv:2410.10542
-
SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers 14 Oct 2024 · 2 repositories · arXiv:2410.10629
-
SLaNC: Static LayerNorm Calibration 14 Oct 2024 · 0 repositories · arXiv:2410.10553
-
STACKFEED: Structured Textual Actor-Critic Knowledge Base Editing with FeedBack 14 Oct 2024 · 0 repositories · arXiv:2410.10584
-
The Ingredients for Robotic Diffusion Transformers 14 Oct 2024 · 0 repositories · arXiv:2410.10088
-
Towards Better Multi-head Attention via Channel-wise Sample Permutation 14 Oct 2024 · 1 repository · arXiv:2410.10914Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Transforming Game Play: A Comparative Study of DCQN and DTQN Architectures in Reinforcement Learning 14 Oct 2024 · 0 repositories · arXiv:2410.10660
-
Transparent Networks for Multivariate Time Series 14 Oct 2024 · 1 repository · arXiv:2410.10535
-
VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents 14 Oct 2024 · 1 repository · arXiv:2410.10594Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
What Does It Mean to Be a Transformer? Insights from a Theoretical Hessian Analysis 14 Oct 2024 · 1 repository · arXiv:2410.10986Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
3DS: Decomposed Difficulty Data Selection's Case Study on LLM Medical Domain Adaptation 13 Oct 2024 · 0 repositories · arXiv:2410.10901
-
A Comparative Study of PDF Parsing Tools Across Diverse Document Categories 13 Oct 2024 · 0 repositories · arXiv:2410.09871
-
Can In-context Learning Really Generalize to Out-of-distribution Tasks? 13 Oct 2024 · 0 repositories · arXiv:2410.09695
-
Data Adaptive Few-shot Multi Label Segmentation with Foundation Model 13 Oct 2024 · 0 repositories · arXiv:2410.09759
-
Dualformer: Controllable Fast and Slow Thinking by Learning with Randomized Reasoning Traces 13 Oct 2024 · 0 repositories · arXiv:2410.09918
-
EasyJudge: an Easy-to-use Tool for Comprehensive Response Evaluation of LLMs 13 Oct 2024 · 1 repository · arXiv:2410.09775
-
Empowering Dysarthric Speech: Leveraging Advanced LLMs for Accurate Speech Correction and Multimodal Emotion Analysis 13 Oct 2024 · 0 repositories · arXiv:2410.12867
-
Evaluating Gender Bias of LLMs in Making Morality Judgements 13 Oct 2024 · 0 repositories · arXiv:2410.09992
-
HARDMath: A Benchmark Dataset for Challenging Problems in Applied Mathematics 13 Oct 2024 · 1 repository · arXiv:2410.09988
-
Honest AI: Fine-Tuning "Small" Language Models to Say "I Don't Know", and Reducing Hallucination in RAG 13 Oct 2024 · 0 repositories · arXiv:2410.09699
-
InterMask: 3D Human Interaction Generation via Collaborative Masked Modelling 13 Oct 2024 · 1 repository · arXiv:2410.10010
-
Investigating Implicit Bias in Large Language Models: A Large-Scale Study of Over 50 LLMs 13 Oct 2024 · 0 repositories · arXiv:2410.12864
-
Learning to Rank for Multiple Retrieval-Augmented Models through Iterative Utility Maximization 13 Oct 2024 · 0 repositories · arXiv:2410.09942
-
M2M-Gen: A Multimodal Framework for Automated Background Music Generation in Japanese Manga Using Large Language Models 13 Oct 2024 · 0 repositories · arXiv:2410.09928
-
Single Ground Truth Is Not Enough: Add Linguistic Variability to Aspect-based Sentiment Analysis Evaluation 13 Oct 2024 · 0 repositories · arXiv:2410.09807
-
Diabetic retinopathy image classification method based on GreenBen data augmentation 12 Oct 2024 · 0 repositories · arXiv:2410.09444
-
EG-SpikeFormer: Eye-Gaze Guided Transformer on Spiking Neural Networks for Medical Image Analysis 12 Oct 2024 · 0 repositories · arXiv:2410.09674
-
Extended Japanese Commonsense Morality Dataset with Masked Token and Label Enhancement 12 Oct 2024 · 0 repositories · arXiv:2410.09564
-
GPTON: Generative Pre-trained Transformers enhanced with Ontology Narration for accurate annotation of biological data 12 Oct 2024 · 0 repositories · arXiv:2410.10899
-
Improving 3D Finger Traits Recognition via Generalizable Neural Rendering 12 Oct 2024 · 0 repositories · arXiv:2410.09582
-
\llinstruct: An Instruction-tuned model for English Language Proficiency Assessments 12 Oct 2024 · 0 repositories · arXiv:2410.09314
-
Looped ReLU MLPs May Be All You Need as Practical Programmable Computers 12 Oct 2024 · 0 repositories · arXiv:2410.09375
-
Token Pruning using a Lightweight Background Aware Vision Transformer 12 Oct 2024 · 0 repositories · arXiv:2410.09324
-
Toward General Instruction-Following Alignment for Retrieval-Augmented Generation 12 Oct 2024 · 1 repository · arXiv:2410.09584Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
A Methodology for Evaluating RAG Systems: A Case Study On Configuration Dependency Validation 11 Oct 2024 · 1 repository · arXiv:2410.08801
-
AttnGCG: Enhancing Jailbreaking Attacks on LLMs with Attention Manipulation 11 Oct 2024 · 1 repository · arXiv:2410.09040
-
CoTCoNet: An Optimized Coupled Transformer-Convolutional Network with an Adaptive Graph Reconstruction for Leukemia Detection 11 Oct 2024 · 0 repositories · arXiv:2410.08797
-
DAT: Dialogue-Aware Transformer with Modality-Group Fusion for Human Engagement Estimation 11 Oct 2024 · 1 repository · arXiv:2410.08470
-
DeBiFormer: Vision Transformer with Deformable Agent Bi-level Routing Attention 11 Oct 2024 · 1 repository · arXiv:2410.08582
-
Developing a Pragmatic Benchmark for Assessing Korean Legal Language Understanding in Large Language Models 11 Oct 2024 · 1 repository · arXiv:2410.08731
-
Encoding Agent Trajectories as Representations with Sequence Transformers 11 Oct 2024 · 0 repositories · arXiv:2410.09204
-
Enhancing Long Context Performance in LLMs Through Inner Loop Query Mechanism 11 Oct 2024 · 0 repositories · arXiv:2410.12859
-
Fine-Tuning In-House Large Language Models to Infer Differential Diagnosis from Radiology Reports 11 Oct 2024 · 0 repositories · arXiv:2410.09234
-
HorGait: A Hybrid Model for Accurate Gait Recognition in LiDAR Point Cloud Planar Projections 11 Oct 2024 · 0 repositories · arXiv:2410.08454
-
Humanity in AI: Detecting the Personality of Large Language Models 11 Oct 2024 · 0 repositories · arXiv:2410.08545
-
Hypothesis-only Biases in Large Language Model-Elicited Natural Language Inference 11 Oct 2024 · 0 repositories · arXiv:2410.08996
-
JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework 11 Oct 2024 · 1 repository · arXiv:2410.12855Syntology 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples)
-
L3Cube-MahaSum: A Comprehensive Dataset and BART Models for Abstractive Text Summarization in Marathi 11 Oct 2024 · 1 repository · arXiv:2410.09184
-
Large Language Models for Medical OSCE Assessment: A Novel Approach to Transcript Analysis 11 Oct 2024 · 0 repositories · arXiv:2410.12858
-
Observing the Southern US Culture of Honor Using Large-Scale Social Media Analysis 11 Oct 2024 · 0 repositories · arXiv:2410.13887
-
oRetrieval Augmented Generation for 10 Large Language Models and its Generalizability in Assessing Medical Fitness 11 Oct 2024 · 0 repositories · arXiv:2410.08431
-
pLDDT-Predictor: High-speed Protein Screening Using Transformer and ESM2 11 Oct 2024 · 1 repository · arXiv:2410.21283
-
Retriever-and-Memory: Towards Adaptive Note-Enhanced Retrieval-Augmented Generation 11 Oct 2024 · 1 repository · arXiv:2410.08821Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Scaling Gaussian Processes for Learning Curve Prediction via Latent Kronecker Structure 11 Oct 2024 · 0 repositories · arXiv:2410.09239
-
SocialGaze: Improving the Integration of Human Social Norms in Large Language Models 11 Oct 2024 · 1 repository · arXiv:2410.08698
-
StructRAG: Boosting Knowledge Intensive Reasoning of LLMs via Inference-time Hybrid Information Structurization 11 Oct 2024 · 1 repository · arXiv:2410.08815
-
SuperCorrect: Supervising and Correcting Language Models with Error-Driven Insights 11 Oct 2024 · 2 repositories · arXiv:2410.09008Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Synth-SONAR: Sonar Image Synthesis with Enhanced Diversity and Realism via Dual Diffusion Models and GPT Prompting 11 Oct 2024 · 1 repository · arXiv:2410.08612
-
ViT3D Alignment of LLaMA3: 3D Medical Image Report Generation 11 Oct 2024 · 0 repositories · arXiv:2410.08588
-
Adam Exploits ℓ_∞-geometry of Loss Landscape via Coordinate-wise Adaptivity 10 Oct 2024 · 1 repository · arXiv:2410.08198Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples)
-
Benchmarking Agentic Workflow Generation 10 Oct 2024 · 1 repository · arXiv:2410.07869Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Can Looped Transformers Learn to Implement Multi-step Gradient Descent for In-context Learning? 10 Oct 2024 · 0 repositories · arXiv:2410.08292
-
Diversity of Thought Elicits Stronger Reasoning Capabilities in Multi-Agent Debate Frameworks 10 Oct 2024 · 0 repositories · arXiv:2410.12853
-
Do You Know What You Are Talking About? Characterizing Query-Knowledge Relevance For Reliable Retrieval Augmented Generation 10 Oct 2024 · 0 repositories · arXiv:2410.08320
-
Explainability of Deep Neural Networks for Brain Tumor Detection 10 Oct 2024 · 1 repository · arXiv:2410.07613
-
FLIER: Few-shot Language Image Models Embedded with Latent Representations 10 Oct 2024 · 0 repositories · arXiv:2410.07648
-
News Reporter: A Multi-lingual LLM Framework for Broadcast T.V News 10 Oct 2024 · 0 repositories · arXiv:2410.07520
-
No Free Lunch: Retrieval-Augmented Generation Undermines Fairness in LLMs, Even for Vigilant Users 10 Oct 2024 · 0 repositories · arXiv:2410.07589
-
Offline Inverse Constrained Reinforcement Learning for Safe-Critical Decision Making in Healthcare 10 Oct 2024 · 0 repositories · arXiv:2410.07525
-
PLaMo-100B: A Ground-Up Language Model Designed for Japanese Proficiency 10 Oct 2024 · 0 repositories · arXiv:2410.07563
-
Pretraining Graph Transformers with Atom-in-a-Molecule Quantum Properties for Improved ADMET Modeling 10 Oct 2024 · 1 repository · arXiv:2410.08024
-
Prompt Engineering a Schizophrenia Chatbot: Utilizing a Multi-Agent Approach for Enhanced Compliance with Prompt Instructions 10 Oct 2024 · 0 repositories · arXiv:2410.12848
-
RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation 10 Oct 2024 · 1 repository · arXiv:2410.07864Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples)
-
Rescriber: Smaller-LLM-Powered User-Led Data Minimization for LLM-Based Chatbots 10 Oct 2024 · 0 repositories · arXiv:2410.11876
-
SNN-PAR: Energy Efficient Pedestrian Attribute Recognition via Spiking Neural Networks 10 Oct 2024 · 1 repository · arXiv:2410.07857
-
SPA: 3D Spatial-Awareness Enables Effective Embodied Representation 10 Oct 2024 · 1 repository · arXiv:2410.08208Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
Teaching-Inspired Integrated Prompting Framework: A Novel Approach for Enhancing Reasoning in Large Language Models 10 Oct 2024 · 1 repository · arXiv:2410.08068Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
The Rise of AI-Generated Content in Wikipedia 10 Oct 2024 · 1 repository · arXiv:2410.08044Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Think Beyond Size: Adaptive Prompting for More Effective Reasoning 10 Oct 2024 · 0 repositories · arXiv:2410.08130
-
Thought2Text: Text Generation from EEG Signal using Large Language Models (LLMs) 10 Oct 2024 · 1 repository · arXiv:2410.07507Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text 10 Oct 2024 · 1 repository · arXiv:2410.07590
-
VibeCheck: Discover and Quantify Qualitative Differences in Large Language Models 10 Oct 2024 · 1 repository · arXiv:2410.12851Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
Adaptive High-Frequency Transformer for Diverse Wildlife Re-Identification 9 Oct 2024 · 1 repository · arXiv:2410.06977
-
Astute RAG: Overcoming Imperfect Retrieval Augmentation and Knowledge Conflicts for Large Language Models 9 Oct 2024 · 0 repositories · arXiv:2410.07176
-
AutoFeedback: An LLM-based Framework for Efficient and Accurate API Request Generation 9 Oct 2024 · 0 repositories · arXiv:2410.06943
-
Can Transformers Reason Logically? A Study in SAT Solving 9 Oct 2024 · 0 repositories · arXiv:2410.07432
-
Capturing Bias Diversity in LLMs 9 Oct 2024 · 0 repositories · arXiv:2410.12839
-
Cluster-wise Graph Transformer with Dual-granularity Kernelized Attention 9 Oct 2024 · 1 repository · arXiv:2410.06746Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 2 honoured, 0 violated, 8 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 14 harvested samples) · 14 pointer-only (licence)
-
Detecting Bias and Enhancing Diagnostic Accuracy in Large Language Models for Healthcare 9 Oct 2024 · 0 repositories · arXiv:2410.06566
-
Do great minds think alike? Investigating Human-AI Complementarity in Question Answering with CAIMIRA 9 Oct 2024 · 0 repositories · arXiv:2410.06524
-
ETA: Evaluating Then Aligning Safety of Vision Language Models at Inference Time 9 Oct 2024 · 1 repository · arXiv:2410.06625Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 1 violated, 5 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching 9 Oct 2024 · 1 repository · arXiv:2410.06885Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples)
-
Generative Model for Less-Resourced Language with 1 billion parameters 9 Oct 2024 · 0 repositories · arXiv:2410.06898
-
Improving Data Efficiency via Curating LLM-Driven Rating Systems 9 Oct 2024 · 0 repositories · arXiv:2410.10877
-
Investigating Cost-Efficiency of LLM-Generated Training Data for Conversational Semantic Frame Analysis 9 Oct 2024 · 0 repositories · arXiv:2410.06550
-
Large Language Models as Code Executors: An Exploratory Study 9 Oct 2024 · 0 repositories · arXiv:2410.06667
-
LLM Self-Correction with DeCRIM: Decompose, Critique, and Refine for Enhanced Following of Instructions with Multiple Constraints 9 Oct 2024 · 0 repositories · arXiv:2410.06458
-
MaD-Scientist: AI-based Scientist solving Convection-Diffusion-Reaction Equations Using Massive PINN-Based Prior Data 9 Oct 2024 · 0 repositories · arXiv:2410.06442