Methods › General › Attention Modules › Multi-Head Attention › Papers, page 39
Multi-Head Attention
Papers archive 2025-07-28
archive papers tagged: 24,855 · with a code link: 11,214 · where Syntology ran a sample: 3,454 (2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,454 of 24,855 tagged: 2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument)
Page 39 of 249: papers 3,801 to 3,900 of 24,855, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
The Nature of Mathematical Modeling and Probabilistic Optimization Engineering in Generative AI 24 Oct 2024 · 0 repositories · arXiv:2410.18441
-
ToolFlow: Boosting LLM Tool-Calling Through Natural and Coherent Dialogue Synthesis 24 Oct 2024 · 0 repositories · arXiv:2410.18447
-
Understanding Players as if They Are Talking to the Game in a Customized Language: A Pilot Study 24 Oct 2024 · 0 repositories · arXiv:2410.18605
-
Where Am I and What Will I See: An Auto-Regressive Model for Spatial Localization and View Prediction 24 Oct 2024 · 0 repositories · arXiv:2410.18962
-
A Methodology for Transformer Ratio Adjustment in Small-Size Rotary Transformers 23 Oct 2024 · 0 repositories · arXiv:2410.18217
-
ALTA: Compiler-Based Analysis of Transformers 23 Oct 2024 · 1 repository · arXiv:2410.18077Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
An Adaptive Framework for Generating Systematic Explanatory Answer in Online Q&A Platforms 23 Oct 2024 · 1 repository · arXiv:2410.17694
-
Anomaly Resilient Temporal QoS Prediction using Hypergraph Convoluted Transformer Network 23 Oct 2024 · 0 repositories · arXiv:2410.17762
-
Beyond Position: the emergence of wavelet-like properties in Transformers 23 Oct 2024 · 0 repositories · arXiv:2410.18067
-
CLR-Bench: Evaluating Large Language Models in College-level Reasoning 23 Oct 2024 · 0 repositories · arXiv:2410.17558
-
Differentially Private Learning Needs Better Model Initialization and Self-Distillation 23 Oct 2024 · 1 repository · arXiv:2410.17566
-
Federated Transformer: Multi-Party Vertical Federated Learning on Practical Fuzzily Linked Data 23 Oct 2024 · 1 repository · arXiv:2410.17986Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
FIPER: Generalizable Factorized Features for Robust Low-Level Vision Models 23 Oct 2024 · 0 repositories · arXiv:2410.18083
-
From PDFs to Structured Data: Utilizing LLM Analysis in Sports Database Management 23 Oct 2024 · 0 repositories · arXiv:2410.17619
-
Future Token Prediction -- Causal Language Modelling with Per-Token Semantic State Vector for Multi-Token Prediction 23 Oct 2024 · 0 repositories · arXiv:2410.18160
-
Gazelle: An Instruction Dataset for Arabic Writing Assistance 23 Oct 2024 · 0 repositories · arXiv:2410.18163
-
Leveraging the Domain Adaptation of Retrieval Augmented Generation Models for Question Answering and Reducing Hallucination 23 Oct 2024 · 0 repositories · arXiv:2410.17783
-
Lightweight Neural App Control 23 Oct 2024 · 0 repositories · arXiv:2410.17883
-
Small Singular Values Matter: A Random Matrix Analysis of Transformer Models 23 Oct 2024 · 0 repositories · arXiv:2410.17770
-
LongRAG: A Dual-Perspective Retrieval-Augmented Generation Paradigm for Long-Context Question Answering 23 Oct 2024 · 1 repository · arXiv:2410.18050Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
MCUBERT: Memory-Efficient BERT Inference on Commodity Microcontrollers 23 Oct 2024 · 0 repositories · arXiv:2410.17957
-
MiLoRA: Efficient Mixture of Low-Rank Adaptation for Large Language Models Fine-tuning 23 Oct 2024 · 0 repositories · arXiv:2410.18035
-
Multi-scale feature reconstruction network for industrial anomaly detection 23 Oct 2024 · 1 repository
-
OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation 23 Oct 2024 · 1 repository · arXiv:2410.17799
-
SimRAG: Self-Improving Retrieval-Augmented Generation for Adapting Large Language Models to Specialized Domains 23 Oct 2024 · 0 repositories · arXiv:2410.17952
-
TabDPT: Scaling Tabular Foundation Models 23 Oct 2024 · 1 repository · arXiv:2410.18164Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
TAGE: Trustworthy Attribute Group Editing for Stable Few-shot Image Generation 23 Oct 2024 · 0 repositories · arXiv:2410.17855
-
Value Residual Learning For Alleviating Attention Concentration In Transformers 23 Oct 2024 · 1 repository · arXiv:2410.17897Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (of 12 harvested samples) · 1 pointer-only (licence)
-
Which Client is Reliable?: A Reliable and Personalized Prompt-based Federated Learning for Medical Image Question Answering 23 Oct 2024 · 0 repositories · arXiv:2410.17484
-
A Bayesian Perspective on the Maximum Score Problem 22 Oct 2024 · 0 repositories · arXiv:2410.17153
-
A Statistical Analysis of LLMs' Self-Evaluation Using Proverbs 22 Oct 2024 · 0 repositories · arXiv:2410.16640
-
An Eye for an AI: Evaluating GPT-4o's Visual Perception Skills and Geometric Reasoning Skills Using Computer Graphics Questions 22 Oct 2024 · 0 repositories · arXiv:2410.16991
-
Assessment of Transformer-Based Encoder-Decoder Model for Human-Like Summarization 22 Oct 2024 · 0 repositories · arXiv:2410.16842
-
Audio-to-Score Conversion Model Based on Whisper methodology 22 Oct 2024 · 0 repositories · arXiv:2410.17209
-
Automated Spinal MRI Labelling from Reports Using a Large Language Model 22 Oct 2024 · 1 repository · arXiv:2410.17235
-
Dhoroni: Exploring Bengali Climate Change and Environmental Views with a Multi-Perspective News Dataset and Natural Language Processing 22 Oct 2024 · 1 repository · arXiv:2410.17225
-
DI-MaskDINO: A Joint Object Detection and Instance Segmentation Model 22 Oct 2024 · 1 repository · arXiv:2410.16707
-
Distill-SynthKG: Distilling Knowledge Graph Synthesis Workflow for Improved Coverage and Efficiency 22 Oct 2024 · 0 repositories · arXiv:2410.16597
-
DNAHLM -- DNA sequence and Human Language mixed large language Model 22 Oct 2024 · 1 repository · arXiv:2410.16917
-
Exploring Possibilities of AI-Powered Legal Assistance in Bangladesh through Large Language Modeling 22 Oct 2024 · 1 repository · arXiv:2410.17210
-
Graph Transformers Dream of Electric Flow 22 Oct 2024 · 0 repositories · arXiv:2410.16699
-
In Context Learning and Reasoning for Symbolic Regression with Large Language Models 22 Oct 2024 · 1 repository · arXiv:2410.17448
-
Interchangeable Token Embeddings for Extendable Vocabulary and Alpha-Equivalence 22 Oct 2024 · 0 repositories · arXiv:2410.17161
-
LiNo: Advancing Recursive Residual Decomposition of Linear and Nonlinear Patterns for Robust Time Series Forecasting 22 Oct 2024 · 1 repository · arXiv:2410.17159
-
Methods of improving LLM training stability 22 Oct 2024 · 0 repositories · arXiv:2410.16682
-
Representation Shattering in Transformers: A Synthetic Study with Knowledge Editing 22 Oct 2024 · 0 repositories · arXiv:2410.17194Syntology 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Scattered Forest Search: Smarter Code Space Exploration with LLMs 22 Oct 2024 · 0 repositories · arXiv:2411.05010
-
SmartRAG: Jointly Learn RAG-Related Tasks From the Environment Feedback 22 Oct 2024 · 0 repositories · arXiv:2410.18141
-
Tracing the Development of the Virtual Particle Concept Using Semantic Change Detection 22 Oct 2024 · 1 repository · arXiv:2410.16855
-
An Efficient System for Automatic Map Storytelling -- A Case Study on Historical Maps 21 Oct 2024 · 1 repository · arXiv:2410.15780
-
An Explainable Contrastive-based Dilated Convolutional Network with Transformer for Pediatric Pneumonia Detection 21 Oct 2024 · 0 repositories · arXiv:2410.16143
-
Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count 21 Oct 2024 · 1 repository · arXiv:2410.15787
-
Beyond 2:4: exploring V:N:M sparsity for efficient transformer inference on GPUs 21 Oct 2024 · 0 repositories · arXiv:2410.16135
-
Building A Coding Assistant via the Retrieval-Augmented Language Model 21 Oct 2024 · 1 repository · arXiv:2410.16229
-
CausalGraph2LLM: Evaluating LLMs for Causal Queries 21 Oct 2024 · 1 repository · arXiv:2410.15939
-
Deep Learning and Data Augmentation for Detecting Self-Admitted Technical Debt 21 Oct 2024 · 1 repository · arXiv:2410.15804
-
Developing Retrieval Augmented Generation (RAG) based LLM Systems from PDFs: An Experience Report 21 Oct 2024 · 1 repository · arXiv:2410.15944
-
Diffusion Transformer Policy 21 Oct 2024 · 1 repository · arXiv:2410.15959
-
Disambiguating Monocular Reconstruction of 3D Clothed Human with Spatial-Temporal Transformer 21 Oct 2024 · 0 repositories · arXiv:2410.16337
-
Domain-Adaptive Pre-training of Self-Supervised Foundation Models for Medical Image Classification in Gastrointestinal Endoscopy 21 Oct 2024 · 1 repository · arXiv:2410.21302
-
Enabling Energy-Efficient Deployment of Large Language Models on Memristor Crossbar: A Synergy of Large and Small 21 Oct 2024 · 0 repositories · arXiv:2410.15977
-
Exploring Pretraining via Active Forgetting for Improving Cross Lingual Transfer for Decoder Language Models 21 Oct 2024 · 0 repositories · arXiv:2410.16168
-
Generalized Probabilistic Attention Mechanism in Transformers 21 Oct 2024 · 0 repositories · arXiv:2410.15578
-
Generalizing Motion Planners with Mixture of Experts for Autonomous Driving 21 Oct 2024 · 1 repository · arXiv:2410.15774Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
GReFEL: Geometry-Aware Reliable Facial Expression Learning under Bias and Imbalanced Data Distribution 21 Oct 2024 · 0 repositories · arXiv:2410.15927
-
Guardians of Discourse: Evaluating LLMs on Multilingual Offensive Language Detection 21 Oct 2024 · 0 repositories · arXiv:2410.15623
-
Improving Neuron-level Interpretability with White-box Language Models 21 Oct 2024 · 0 repositories · arXiv:2410.16443
-
Large Body Language Models 21 Oct 2024 · 0 repositories · arXiv:2410.16533
-
Large Language Models in Computer Science Education: A Systematic Literature Review 21 Oct 2024 · 1 repository · arXiv:2410.16349
-
Leveraging Retrieval-Augmented Generation for Culturally Inclusive Hakka Chatbots: Design Insights and User Perceptions 21 Oct 2024 · 0 repositories · arXiv:2410.15572
-
LightFusionRec: Lightweight Transformers-Based Cross-Domain Recommendation Model 21 Oct 2024 · 0 repositories · arXiv:2410.15656
-
Modelling Concurrent RTP Flows for End-to-end Predictions of QoS in Real Time Communications 21 Oct 2024 · 0 repositories · arXiv:2410.15846
-
Natural GaLore: Accelerating GaLore for memory-efficient LLM Training and Fine-tuning 21 Oct 2024 · 1 repository · arXiv:2410.16029
-
On Creating an English-Thai Code-switched Machine Translation in Medical Domain 21 Oct 2024 · 1 repository · arXiv:2410.16221
-
RAG4ITOps: A Supervised Fine-Tunable and Comprehensive RAG Framework for IT Operations and Maintenance 21 Oct 2024 · 0 repositories · arXiv:2410.15805
-
Reflection-Bench: probing AI intelligence with reflection 21 Oct 2024 · 1 repository · arXiv:2410.16270
-
SeisLM: a Foundation Model for Seismic Waveforms 21 Oct 2024 · 1 repository · arXiv:2410.15765
-
Students Rather Than Experts: A New AI For Education Pipeline To Model More Human-Like And Personalised Early Adolescences 21 Oct 2024 · 0 repositories · arXiv:2410.15701
-
Towards a Reliable Offline Personal AI Assistant for Long Duration Spaceflight 21 Oct 2024 · 0 repositories · arXiv:2410.16397
-
Using GPT Models for Qualitative and Quantitative News Analytics in the 2024 US Presidental Election Process 21 Oct 2024 · 0 repositories · arXiv:2410.15884
-
ViMoE: An Empirical Study of Designing Vision Mixture-of-Experts 21 Oct 2024 · 0 repositories · arXiv:2410.15732
-
Weighted Diversified Sampling for Efficient Data-Driven Single-Cell Gene-Gene Interaction Discovery 21 Oct 2024 · 0 repositories · arXiv:2410.15616
-
Who's Who: Large Language Models Meet Knowledge Conflicts in Practice 21 Oct 2024 · 1 repository · arXiv:2410.15737
-
YOLO11 and Vision Transformers based 3D Pose Estimation of Immature Green Fruits in Commercial Apple Orchards for Robotic Thinning 21 Oct 2024 · 0 repositories · arXiv:2410.19846
-
Advancing Gasoline Consumption Forecasting: A Novel Hybrid Model Integrating Transformers, LSTM, and CNN 20 Oct 2024 · 0 repositories · arXiv:2410.16336
-
Back to School: Translation Using Grammar Books 20 Oct 2024 · 1 repository · arXiv:2410.15263Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 6 harvested samples)
-
BRIEF: Bridging Retrieval and Inference for Multi-hop Reasoning via Compression 20 Oct 2024 · 1 repository · arXiv:2410.15277
-
Comparative Analysis of LSTM, GRU, and Transformer Models for Stock Price Prediction 20 Oct 2024 · 0 repositories · arXiv:2411.05790
-
Contextual Augmented Multi-Model Programming (CAMP): A Hybrid Local-Cloud Copilot Framework 20 Oct 2024 · 1 repository · arXiv:2410.15285
-
ConTReGen: Context-driven Tree-structured Retrieval for Open-domain Long-form Text Generation 20 Oct 2024 · 0 repositories · arXiv:2410.15511
-
Do RAG Systems Cover What Matters? Evaluating and Optimizing Responses with Sub-Question Coverage 20 Oct 2024 · 1 repository · arXiv:2410.15531
-
Does ChatGPT Have a Poetic Style? 20 Oct 2024 · 1 repository · arXiv:2410.15299
-
Evaluating Consistencies in LLM responses through a Semantic Clustering of Question Answering 20 Oct 2024 · 0 repositories · arXiv:2410.15440
-
Exploring Social Desirability Response Bias in Large Language Models: Evidence from GPT-4 Simulations 20 Oct 2024 · 0 repositories · arXiv:2410.15442
-
LTPNet Integration of Deep Learning and Environmental Decision Support Systems for Renewable Energy Demand Forecasting 20 Oct 2024 · 0 repositories · arXiv:2410.15286
-
MMDS: A Multimodal Medical Diagnosis System Integrating Image Analysis and Knowledge-based Departmental Consultation 20 Oct 2024 · 0 repositories · arXiv:2410.15403
-
SDP4Bit: Toward 4-bit Communication Quantization in Sharded Data Parallelism for LLM Training 20 Oct 2024 · 0 repositories · arXiv:2410.15526
-
SEA: State-Exchange Attention for High-Fidelity Physics Based Transformers 20 Oct 2024 · 1 repository · arXiv:2410.15495Syntology official (archive's flag): 11 ran · 11 ran (of which 7 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 1 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples)
-
Training Language Models to Critique With Multi-agent Feedback 20 Oct 2024 · 0 repositories · arXiv:2410.15287
-
Unveiling and Consulting Core Experts in Retrieval-Augmented MoE-based LLMs 20 Oct 2024 · 0 repositories · arXiv:2410.15438