Methods › General › Attention Modules › Multi-Head Attention › Papers, page 87
Multi-Head Attention
Papers archive 2025-07-28
archive papers tagged: 24,855 · with a code link: 11,214 · where Syntology ran a sample: 3,454 (2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,454 of 24,855 tagged: 2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument)
Page 87 of 249: papers 8,601 to 8,700 of 24,855, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models 26 Feb 2024 · 1 repository · arXiv:2402.16438
-
Layer-wise Regularized Dropout for Neural Language Models 26 Feb 2024 · 0 repositories · arXiv:2402.16361
-
MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs 26 Feb 2024 · 0 repositories · arXiv:2402.16352
-
Predicting Sustainable Development Goals Using Course Descriptions -- from LLMs to Conventional Foundation Models 26 Feb 2024 · 0 repositories · arXiv:2402.16420
-
QASE Enhanced PLMs: Improved Control in Text Generation for MRC 26 Feb 2024 · 0 repositories · arXiv:2403.04771
-
Retrieval Augmented Generation Systems: Automatic Dataset Creation, Evaluation and Boolean Agent Setup 26 Feb 2024 · 1 repository · arXiv:2403.00820
-
Building Flexible Machine Learning Models for Scientific Computing at Scale 25 Feb 2024 · 0 repositories · arXiv:2402.16014
-
ChatMusician: Understanding and Generating Music Intrinsically with LLM 25 Feb 2024 · 1 repository · arXiv:2402.16153Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Cross-Resolution Land Cover Classification Using Outdated Products and Transformers 25 Feb 2024 · 1 repository · arXiv:2402.16001
-
Deep Learning Approaches for Improving Question Answering Systems in Hepatocellular Carcinoma Research 25 Feb 2024 · 0 repositories · arXiv:2402.16038
-
Exploiting Regional Information Transformer for Single Image Deraining 25 Feb 2024 · 1 repository · arXiv:2402.16033
-
DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers 25 Feb 2024 · 1 repository · arXiv:2402.16914Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples)
-
EHRNoteQA: An LLM Benchmark for Real-World Clinical Practice Using Discharge Summaries 25 Feb 2024 · 1 repository · arXiv:2402.16040Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Emotion Classification in Short English Texts using Deep Learning Techniques 25 Feb 2024 · 0 repositories · arXiv:2402.16034
-
Exploring the Power of Pure Attention Mechanisms in Blind Room Parameter Estimation 25 Feb 2024 · 0 repositories · arXiv:2402.16003
-
From Text to Transformation: A Comprehensive Review of Large Language Models' Versatility 25 Feb 2024 · 0 repositories · arXiv:2402.16142
-
Knowledge Fusion of Chat LLMs: A Preliminary Technical Report 25 Feb 2024 · 2 repositories · arXiv:2402.16107
-
GraphWiz: An Instruction-Following Language Model for Graph Problems 25 Feb 2024 · 1 repository · arXiv:2402.16029Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 15 harvested samples) · 15 pointer-only (licence)
-
Hitting "Probe"rty with Non-Linearity, and More 25 Feb 2024 · 0 repositories · arXiv:2402.16168
-
One-stage Prompt-based Continual Learning 25 Feb 2024 · 0 repositories · arXiv:2402.16189
-
PIDformer: Transformer Meets Control Theory 25 Feb 2024 · 0 repositories · arXiv:2402.15989
-
StochCA: A Novel Approach for Exploiting Pretrained Models with Cross-Attention 25 Feb 2024 · 1 repository · arXiv:2402.16092
-
Text Understanding and Generation Using Transformer Models for Intelligent E-commerce Recommendations 25 Feb 2024 · 0 repositories · arXiv:2402.16035
-
Bridging the Gap between 2D and 3D Visual Question Answering: A Fusion Approach for 3D VQA 24 Feb 2024 · 1 repository · arXiv:2402.15933Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 1 honoured, 0 violated, 7 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
Enhancing Cloud-Based Large Language Model Processing with Elasticsearch and Transformer Models 24 Feb 2024 · 0 repositories · arXiv:2403.00807
-
How Do Humans Write Code? Large Models Do It the Same Way Too 24 Feb 2024 · 1 repository · arXiv:2402.15729
-
TV-SAM: Increasing Zero-Shot Segmentation Performance on Multimodal Medical Images Using GPT-4 Generated Descriptive Prompts Without Human Annotation 24 Feb 2024 · 1 repository · arXiv:2402.15759
-
LLMs Can Defend Themselves Against Jailbreaking in a Practical Manner: A Vision Paper 24 Feb 2024 · 0 repositories · arXiv:2402.15727
-
Look Before You Leap: Problem Elaboration Prompting Improves Mathematical Reasoning in Large Language Models 24 Feb 2024 · 0 repositories · arXiv:2402.15764
-
MATHWELL: Generating Educational Math Word Problems Using Teacher Annotations 24 Feb 2024 · 2 repositories · arXiv:2402.15861
-
MultiContrievers: Analysis of Dense Retrieval Representations 24 Feb 2024 · 1 repository · arXiv:2402.15925
-
Predicting Outcomes in Video Games with Long Short Term Memory Networks 24 Feb 2024 · 1 repository · arXiv:2402.15923
-
PRP: Propagating Universal Perturbations to Attack Large Language Model Guard-Rails 24 Feb 2024 · 0 repositories · arXiv:2402.15911
-
Res-VMamba: Fine-Grained Food Category Visual Classification Using Selective State Space Models with Deep Residual Learning 24 Feb 2024 · 1 repository · arXiv:2402.15761Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 3 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
SemEval-2024 Task 8: Weighted Layer Averaging RoBERTa for Black-Box Machine-Generated Text Detection 24 Feb 2024 · 1 repository · arXiv:2402.15873
-
A Data-Centric Approach To Generate Faithful and High Quality Patient Summaries with Large Language Models 23 Feb 2024 · 1 repository · arXiv:2402.15422Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
A First Look at GPT Apps: Landscape and Vulnerability 23 Feb 2024 · 0 repositories · arXiv:2402.15105
-
Advancing Parameter Efficiency in Fine-tuning via Representation Editing 23 Feb 2024 · 2 repositories · arXiv:2402.15179Syntology official (archive's flag): 1 ran · 2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; every one of the 2 samples that ran constructed an object rather than computing a result (of 2 harvested samples) · 1 pointer-only (licence)
-
Attention-aware Semantic Communications for Collaborative Inference 23 Feb 2024 · 1 repository · arXiv:2404.07217
-
AttributionBench: How Hard is Automatic Attribution Evaluation? 23 Feb 2024 · 1 repository · arXiv:2402.15089
-
Constraint Latent Space Matters: An Anti-anomalous Waveform Transformation Solution from Photoplethysmography to Arterial Blood Pressure 23 Feb 2024 · 0 repositories · arXiv:2402.17780
-
Descripción automática de secciones delgadas de rocas: una aplicación Web 23 Feb 2024 · 0 repositories · arXiv:2402.15039
-
Dual Encoder: Exploiting the Potential of Syntactic and Semantic for Aspect Sentiment Triplet Extraction 23 Feb 2024 · 0 repositories · arXiv:2402.15370
-
Evaluating the Performance of ChatGPT for Spam Email Detection 23 Feb 2024 · 0 repositories · arXiv:2402.15537
-
Executing Natural Language-Described Algorithms with Large Language Models: An Investigation 23 Feb 2024 · 1 repository · arXiv:2403.00795
-
Fiducial Focus Augmentation for Facial Landmark Detection 23 Feb 2024 · 0 repositories · arXiv:2402.15044
-
Fine-tuning Large Language Models for Domain-specific Machine Translation 23 Feb 2024 · 0 repositories · arXiv:2402.15061
-
How (un)ethical are instruction-centric responses of LLMs? Unveiling the vulnerabilities of safety guardrails to harmful queries 23 Feb 2024 · 1 repository · arXiv:2402.15302
-
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions 23 Feb 2024 · 0 repositories · arXiv:2402.15055
-
MambaIR: A Simple Baseline for Image Restoration with State-Space Model 23 Feb 2024 · 2 repositories · arXiv:2402.15648Syntology official (archive's flag): 8 ran · 14 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 5 where Syntology's instrument failed) · 1 unverified (of 15 harvested samples)
-
Multimodal Transformer With a Low-Computational-Cost Guarantee 23 Feb 2024 · 0 repositories · arXiv:2402.15096
-
LLMs as Meta-Reviewers' Assistants: A Case Study 23 Feb 2024 · 1 repository · arXiv:2402.15589
-
Spatially-Aware Transformer for Embodied Agents 23 Feb 2024 · 1 repository · arXiv:2402.15160Syntology official (archive's flag): 3 ran · 3 ran (of which 1 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 1 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
State Space Models for Event Cameras 23 Feb 2024 · 2 repositories · arXiv:2402.15584Syntology official (archive's flag): 1 ran · 17 ran (of which 0 constructed an object rather than computing a result; 16 with no instrument failure: 0 honoured, 0 violated, 16 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 19 harvested samples) · 8 pointer-only (licence)
-
The Good and The Bad: Exploring Privacy Issues in Retrieval-Augmented Generation (RAG) 23 Feb 2024 · 1 repository · arXiv:2402.16893Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
The Impact of LoRA on the Emergence of Clusters in Transformers 23 Feb 2024 · 1 repository · arXiv:2402.15415Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
ToMBench: Benchmarking Theory of Mind in Large Language Models 23 Feb 2024 · 1 repository · arXiv:2402.15052Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Towards Efficient Active Learning in NLP via Pretrained Representations 23 Feb 2024 · 0 repositories · arXiv:2402.15613
-
How Do Nonlinear Transformers Learn and Generalize in In-Context Learning? 23 Feb 2024 · 0 repositories · arXiv:2402.15607
-
Transformers are Expressive, But Are They Expressive Enough for Regression? 23 Feb 2024 · 1 repository · arXiv:2402.15478Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
2D Matryoshka Sentence Embeddings 22 Feb 2024 · 1 repository · arXiv:2402.14776
-
A Transformer Model for Boundary Detection in Continuous Sign Language 22 Feb 2024 · 0 repositories · arXiv:2402.14720
-
Annotation and Classification of Relevant Clauses in Terms-and-Conditions Contracts 22 Feb 2024 · 1 repository · arXiv:2402.14457
-
Assessing generalization capability of text ranking models in Polish 22 Feb 2024 · 0 repositories · arXiv:2402.14318
-
BeTAIL: Behavior Transformer Adversarial Imitation Learning from Human Racing Gameplay 22 Feb 2024 · 0 repositories · arXiv:2402.14194
-
Can Large Language Models Detect Misinformation in Scientific News Reporting? 22 Feb 2024 · 0 repositories · arXiv:2402.14268
-
Compression Robust Synthetic Speech Detection Using Patched Spectrogram Transformer 22 Feb 2024 · 0 repositories · arXiv:2402.14205
-
Copilot Evaluation Harness: Evaluating LLM-Guided Software Programming 22 Feb 2024 · 0 repositories · arXiv:2402.14261
-
COPR: Continual Human Preference Learning via Optimal Policy Regularization 22 Feb 2024 · 0 repositories · arXiv:2402.14228
-
GATE X-E : A Challenge Set for Gender-Fair Translations from Weakly-Gendered Languages 22 Feb 2024 · 0 repositories · arXiv:2402.14277
-
Hint-before-Solving Prompting: Guiding LLMs to Effectively Utilize Encoded Knowledge 22 Feb 2024 · 1 repository · arXiv:2402.14310
-
HINT: High-quality INPainting Transformer with Mask-Aware Encoding and Enhanced Attention 22 Feb 2024 · 1 repository · arXiv:2402.14185
-
In-Context Learning of a Linear Transformer Block: Benefits of the MLP Component and One-Step GD Initialization 22 Feb 2024 · 0 repositories · arXiv:2402.14951
-
Is ChatGPT More Empathetic than Humans? 22 Feb 2024 · 1 repository · arXiv:2403.05572
-
Is ChatGPT the Future of Causal Text Mining? A Comprehensive Evaluation and Analysis 22 Feb 2024 · 0 repositories · arXiv:2402.14484
-
KoCoSa: Korean Context-aware Sarcasm Detection Dataset 22 Feb 2024 · 1 repository · arXiv:2402.14428
-
Middleware for LLMs: Tools Are Instrumental for Language Agents in Complex Environments 22 Feb 2024 · 2 repositories · arXiv:2402.14672Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment 22 Feb 2024 · 1 repository · arXiv:2402.14968
-
Multi-HMR: Multi-Person Whole-Body Human Mesh Recovery in a Single Shot 22 Feb 2024 · 1 repository · arXiv:2402.14654Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement 22 Feb 2024 · 1 repository · arXiv:2402.14658Syntology 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Path Planning based on 2D Object Bounding-box 22 Feb 2024 · 0 repositories · arXiv:2402.14933
-
Personalized Behavior-Aware Transformer for Multi-Behavior Sequential Recommendation 22 Feb 2024 · 1 repository · arXiv:2402.14473
-
RoboScript: Code Generation for Free-Form Manipulation Tasks across Real and Simulation 22 Feb 2024 · 0 repositories · arXiv:2402.14623
-
Self-supervised Visualisation of Medical Image Datasets 22 Feb 2024 · 1 repository · arXiv:2402.14566
-
Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs 22 Feb 2024 · 1 repository · arXiv:2402.14903Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Towards Understanding Counseling Conversations: Domain Knowledge and Large Language Models 22 Feb 2024 · 0 repositories · arXiv:2402.14200
-
Transferring BERT Capabilities from High-Resource to Low-Resource Languages Using Vocabulary Matching 22 Feb 2024 · 0 repositories · arXiv:2402.14408
-
Two-stage Cytopathological Image Synthesis for Augmenting Cervical Abnormality Screening 22 Feb 2024 · 0 repositories · arXiv:2402.14707
-
Uncertainty-Aware Evaluation for Vision-Language Models 22 Feb 2024 · 1 repository · arXiv:2402.14418Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 1 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 9 unverified (of 17 harvested samples)
-
Whose LLM is it Anyway? Linguistic Comparison and LLM Attribution for GPT-3.5, GPT-4 and Bard 22 Feb 2024 · 0 repositories · arXiv:2402.14533
-
ActiveRAG: Autonomously Knowledge Assimilation and Accommodation through Retrieval-Augmented Agents 21 Feb 2024 · 1 repository · arXiv:2402.13547
-
An Evaluation of Large Language Models in Bioinformatics Research 21 Feb 2024 · 0 repositories · arXiv:2402.13714
-
An Explainable Transformer-based Model for Phishing Email Detection: A Large Language Model Approach 21 Feb 2024 · 0 repositories · arXiv:2402.13871
-
Are LLMs Effective Negotiators? Systematic Evaluation of the Multifaceted Capabilities of LLMs in Negotiation Dialogues 21 Feb 2024 · 1 repository · arXiv:2402.13550Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Beyond A*: Better Planning with Transformers via Search Dynamics Bootstrapping 21 Feb 2024 · 1 repository · arXiv:2402.14083Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 14 harvested samples) · 14 pointer-only (licence)
-
Beyond Hate Speech: NLP's Challenges and Opportunities in Uncovering Dehumanizing Language 21 Feb 2024 · 0 repositories · arXiv:2402.13818
-
CriticEval: Evaluating Large Language Model as Critic 21 Feb 2024 · 2 repositories · arXiv:2402.13764
-
Data-driven Discovery with Large Generative Models 21 Feb 2024 · 0 repositories · arXiv:2402.13610
-
Do Efficient Transformers Really Save Computation? 21 Feb 2024 · 0 repositories · arXiv:2402.13934
-
Driving Generative Agents With Their Personality 21 Feb 2024 · 0 repositories · arXiv:2402.14879