Methods › General › Attention Modules › Multi-Head Attention › Papers, page 78
Multi-Head Attention
Papers archive 2025-07-28
archive papers tagged: 24,855 · with a code link: 11,214 · where Syntology ran a sample: 3,454 (2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,454 of 24,855 tagged: 2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument)
Page 78 of 249: papers 7,701 to 7,800 of 24,855, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Remembering Transformer for Continual Learning 11 Apr 2024 · 0 repositories · arXiv:2404.07518
-
Rumour Evaluation with Very Large Language Models 11 Apr 2024 · 1 repository · arXiv:2404.16859
-
Structure-aware Fine-tuning for Code Pre-trained Models 11 Apr 2024 · 0 repositories · arXiv:2404.07471
-
Token Space: A Category Theory Framework for AI Computations 11 Apr 2024 · 0 repositories · arXiv:2404.11624
-
ViM-UNet: Vision Mamba for Biomedical Segmentation 11 Apr 2024 · 1 repository · arXiv:2404.07705
-
Control-DAG: Constrained Decoding for Non-Autoregressive Directed Acyclic T5 using Weighted Finite State Automata 10 Apr 2024 · 1 repository · arXiv:2404.06854
-
Dynamic Generation of Personalities with Large Language Models 10 Apr 2024 · 1 repository · arXiv:2404.07084
-
Emotion-cause pair extraction method based on multi-granularity information and multi-module interaction 10 Apr 2024 · 0 repositories · arXiv:2404.06812
-
Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention 10 Apr 2024 · 5 repositories · arXiv:2404.07143Syntology 15 ran (of which 4 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 2 violated, 5 with no contract checked; 8 where Syntology's instrument failed) · 1 unverified (of 16 harvested samples) · 6 pointer-only (licence)
-
Llama-VITS: Enhancing TTS Synthesis with Semantic Awareness 10 Apr 2024 · 1 repository · arXiv:2404.06714Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
NFARec: A Negative Feedback-Aware Recommender Model 10 Apr 2024 · 1 repository · arXiv:2404.06900
-
Not All Contexts Are Equal: Teaching LLMs Credibility-aware Generation 10 Apr 2024 · 1 repository · arXiv:2404.06809
-
Simpler becomes Harder: Do LLMs Exhibit a Coherent Behavior on Simplified Corpora? 10 Apr 2024 · 1 repository · arXiv:2404.06838
-
Superposition Prompting: Improving and Accelerating Retrieval-Augmented Generation 10 Apr 2024 · 1 repository · arXiv:2404.06910
-
Heuristic-enhanced Candidates Selection strategy for GPTs tackle Few-Shot Aspect-Based Sentiment Analysis 9 Apr 2024 · 1 repository · arXiv:2404.06063
-
Characterizing Multimodal Long-form Summarization: A Case Study on Financial Reports 9 Apr 2024 · 0 repositories · arXiv:2404.06162
-
Comparing Two Model Designs for Clinical Note Generation; Is an LLM a Useful Evaluator of Consistency? 9 Apr 2024 · 0 repositories · arXiv:2404.06503
-
Exploring the Potential of Large Foundation Models for Open-Vocabulary HOI Detection 9 Apr 2024 · 1 repository · arXiv:2404.06194Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Generative Pre-Trained Transformer for Symbolic Regression Base In-Context Reinforcement Learning 9 Apr 2024 · 0 repositories · arXiv:2404.06330
-
InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD 9 Apr 2024 · 2 repositories · arXiv:2404.06512
-
LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders 9 Apr 2024 · 1 repository · arXiv:2404.05961
-
LLMs' Reading Comprehension Is Affected by Parametric Knowledge and Struggles with Hypothetical Statements 9 Apr 2024 · 0 repositories · arXiv:2404.06283
-
Efficient Concertormer for Image Deblurring and Beyond 9 Apr 2024 · 0 repositories · arXiv:2404.06135
-
PGTNet: A Process Graph Transformer Network for Remaining Time Prediction of Business Process Instances 9 Apr 2024 · 1 repository · arXiv:2404.06267
-
Sandwich attack: Multi-language Mixture Adaptive Attack on LLMs 9 Apr 2024 · 0 repositories · arXiv:2404.07242
-
scRDiT: Generating single-cell RNA-seq data by diffusion transformers and accelerating sampling 9 Apr 2024 · 1 repository · arXiv:2404.06153
-
WebCode2M: A Real-World Dataset for Code Generation from Webpage Designs 9 Apr 2024 · 0 repositories · arXiv:2404.06369
-
Comprehensive Study on German Language Models for Clinical and Biomedical Text Understanding 8 Apr 2024 · 0 repositories · arXiv:2404.05694
-
Guiding Large Language Models to Generate Computer-Parsable Content 8 Apr 2024 · 0 repositories · arXiv:2404.05499
-
Decision Transformers for Wireless Communications: A New Paradigm of Resource Management 8 Apr 2024 · 0 repositories · arXiv:2404.05199
-
Deep Optics for Video Snapshot Compressive Imaging 8 Apr 2024 · 1 repository · arXiv:2404.05274Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
Enhancing Lip Reading with Multi-Scale Video and Multi-Encoder 8 Apr 2024 · 0 repositories · arXiv:2404.05466
-
Evaluating Interventional Reasoning Capabilities of Large Language Models 8 Apr 2024 · 0 repositories · arXiv:2404.05545
-
Evaluation of an LLM in Identifying Logical Fallacies: A Call for Rigor When Adopting LLMs in HCI Research 8 Apr 2024 · 0 repositories · arXiv:2404.05213
-
Fighting crime with Transformers: Empirical analysis of address parsing methods in payment data 8 Apr 2024 · 1 repository · arXiv:2404.05632
-
HSViT: Horizontally Scalable Vision Transformer 8 Apr 2024 · 1 repository · arXiv:2404.05196
-
LLM Reasoners: New Evaluation, Library, and Analysis of Step-by-Step Reasoning with Large Language Models 8 Apr 2024 · 1 repository · arXiv:2404.05221Syntology 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
LTNER: Large Language Model Tagging for Named Entity Recognition with Contextualized Entity Marking 8 Apr 2024 · 0 repositories · arXiv:2404.05624
-
MedExpQA: Multilingual Benchmarking of Large Language Models for Medical Question Answering 8 Apr 2024 · 0 repositories · arXiv:2404.05590
-
MLP Can Be A Good Transformer Learner 8 Apr 2024 · 1 repository · arXiv:2404.05657Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Multi-head Attention-based Deep Multiple Instance Learning 8 Apr 2024 · 1 repository · arXiv:2404.05362
-
PetKaz at SemEval-2024 Task 3: Advancing Emotion Classification with an LLM for Emotion-Cause Pair Extraction in Conversations 8 Apr 2024 · 1 repository · arXiv:2404.05502
-
Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws 8 Apr 2024 · 0 repositories · arXiv:2404.05405
-
Relation Extraction Using Large Language Models: A Case Study on Acupuncture Point Locations 8 Apr 2024 · 0 repositories · arXiv:2404.05415
-
Semantic Stealth: Adversarial Text Attacks on NLP Using Several Methods 8 Apr 2024 · 0 repositories · arXiv:2404.05159
-
Use of a Structured Knowledge Base Enhances Metadata Curation by Large Language Models 8 Apr 2024 · 1 repository · arXiv:2404.05893
-
Xiwu: A Basis Flexible and Learnable LLM for High Energy Physics 8 Apr 2024 · 1 repository · arXiv:2404.08001
-
A Multi-Level Framework for Accelerating Training Transformer Models 7 Apr 2024 · 1 repository · arXiv:2404.07999Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
MM-MATH: Advancing Multimodal Math Evaluation with Process Evaluation and Fine-grained Classification 7 Apr 2024 · 1 repository · arXiv:2404.05091
-
Contextual Chart Generation for Cyber Deception 7 Apr 2024 · 0 repositories · arXiv:2404.04854
-
CSA-Trans: Code Structure Aware Transformer for AST 7 Apr 2024 · 1 repository · arXiv:2404.05767Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Data Bias According to Bipol: Men are Naturally Right and It is the Role of Women to Follow Their Lead 7 Apr 2024 · 1 repository · arXiv:2404.04838
-
Dual-Scale Transformer for Large-Scale Single-Pixel Imaging 7 Apr 2024 · 1 repository · arXiv:2404.05001Syntology official (archive's flag): 6 ran · 6 ran (of which 3 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
GvT: A Graph-based Vision Transformer with Talking-Heads Utilizing Sparsity, Trained from Scratch on Small Datasets 7 Apr 2024 · 0 repositories · arXiv:2404.04924
-
Hyperbolic Learning with Synthetic Captions for Open-World Detection 7 Apr 2024 · 0 repositories · arXiv:2404.05016
-
Initial Exploration of Zero-Shot Privacy Utility Tradeoffs in Tabular Data Using GPT-4 7 Apr 2024 · 0 repositories · arXiv:2404.05047
-
Joint Reconstruction of 3D Human and Object via Contact-Based Refinement Transformer 7 Apr 2024 · 1 repository · arXiv:2404.04819Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
LHU-Net: A Light Hybrid U-Net for Cost-Efficient, High-Performance Volumetric Medical Image Segmentation 7 Apr 2024 · 1 repository · arXiv:2404.05102Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
PagPassGPT: Pattern Guided Password Guessing via Generative Pretrained Transformer 7 Apr 2024 · 1 repository · arXiv:2404.04886
-
Rethinking Diffusion Model for Multi-Contrast MRI Super-Resolution 7 Apr 2024 · 1 repository · arXiv:2404.04785Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 7 where Syntology's instrument failed) · 0 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
VMambaMorph: a Multi-Modality Deformable Image Registration Framework based on Visual State Space Model with Cross-Scan Module 7 Apr 2024 · 1 repository · arXiv:2404.05105
-
A Morphology-Based Investigation of Positional Encodings 6 Apr 2024 · 0 repositories · arXiv:2404.04530
-
Cluster-based Video Summarization with Temporal Context Awareness 6 Apr 2024 · 1 repository · arXiv:2404.04511
-
Convolutional Neural Network Transformer (CNNT) for Fluorescence Microscopy image Denoising with Improved Generalization and Fast Adaptation 6 Apr 2024 · 0 repositories · arXiv:2404.04726
-
Empowering Image Recovery_ A Multi-Attention Approach 6 Apr 2024 · 0 repositories · arXiv:2404.04617
-
Decomposition-based Unsupervised Domain Adaptation for Remote Sensing Image Semantic Segmentation 6 Apr 2024 · 1 repository · arXiv:2404.04531Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples)
-
IITK at SemEval-2024 Task 2: Exploring the Capabilities of LLMs for Safe Biomedical Natural Language Inference for Clinical Trials 6 Apr 2024 · 1 repository · arXiv:2404.04510
-
Large Language Model (LLM) AI text generation detection based on transformer deep learning algorithm 6 Apr 2024 · 0 repositories · arXiv:2405.06652
-
MACM: Utilizing a Multi-Agent System for Condition Mining in Solving Complex Mathematical Problems 6 Apr 2024 · 1 repository · arXiv:2404.04735Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Mixed-Query Transformer: A Unified Image Segmentation Architecture 6 Apr 2024 · 0 repositories · arXiv:2404.04469
-
Q-PEFT: Query-dependent Parameter Efficient Fine-tuning for Text Reranking with Large Language Models 6 Apr 2024 · 0 repositories · arXiv:2404.04522
-
RecGPT: Generative Personalized Prompts for Sequential Recommendation via ChatGPT Training Paradigm 6 Apr 2024 · 0 repositories · arXiv:2404.08675
-
Bayesian Additive Regression Networks 5 Apr 2024 · 1 repository · arXiv:2404.04425
-
Cleared for Takeoff? Compositional & Conditional Reasoning may be the Achilles Heel to (Flight-Booking) Language Agents 5 Apr 2024 · 0 repositories · arXiv:2404.04237
-
Deciphering Political Entity Sentiment in News with Large Language Models: Zero-Shot and Few-Shot Strategies 5 Apr 2024 · 1 repository · arXiv:2404.04361
-
Effects of Different Prompts on the Quality of GPT-4 Responses to Dementia Care Questions 5 Apr 2024 · 0 repositories · arXiv:2404.08674
-
Context-Aware Aerial Object Detection: Leveraging Inter-Object and Background Relationships 5 Apr 2024 · 0 repositories · arXiv:2404.04140
-
Investigating the Robustness of Modelling Decisions for Few-Shot Cross-Topic Stance Detection: A Preregistered Study 5 Apr 2024 · 1 repository · arXiv:2404.03987
-
Learning Correlation Structures for Vision Transformers 5 Apr 2024 · 0 repositories · arXiv:2404.03924
-
PhysPT: Physics-aware Pretrained Transformer for Estimating Human Dynamics from Monocular Videos 5 Apr 2024 · 0 repositories · arXiv:2404.04430
-
player2vec: A Language Modeling Approach to Understand Player Behavior in Games 5 Apr 2024 · 0 repositories · arXiv:2404.04234
-
Re-pseudonymization Strategies for Smart Meter Data Are Not Robust to Deep Learning Profiling Attacks 5 Apr 2024 · 0 repositories · arXiv:2404.03948
-
Scope Ambiguities in Large Language Models 5 Apr 2024 · 1 repository · arXiv:2404.04332
-
The NES Video-Music Database: A Dataset of Symbolic Video Game Music Paired with Gameplay Videos 5 Apr 2024 · 1 repository · arXiv:2404.04420
-
Uformer: A UNet-Transformer fused robust end-to-end deep learning framework for real-time denoising of lung sounds 5 Apr 2024 · 0 repositories · arXiv:2404.04365
-
A Comprehensive Survey on Self-Supervised Learning for Recommendation 4 Apr 2024 · 1 repository · arXiv:2404.03354
-
A Directional Diffusion Graph Transformer for Recommendation 4 Apr 2024 · 0 repositories · arXiv:2404.03326
-
AutoWebGLM: A Large Language Model-based Web Navigating Agent 4 Apr 2024 · 1 repository · arXiv:2404.03648Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
BanglaAutoKG: Automatic Bangla Knowledge Graph Construction with Semantic Neural Graph Filtering 4 Apr 2024 · 1 repository · arXiv:2404.03528Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Capabilities of Large Language Models in Control Engineering: A Benchmark Study on GPT-4, Claude 3 Opus, and Gemini 1.0 Ultra 4 Apr 2024 · 0 repositories · arXiv:2404.03647
-
CBR-RAG: Case-Based Reasoning for Retrieval Augmented Generation in LLMs for Legal Question Answering 4 Apr 2024 · 1 repository · arXiv:2404.04302
-
CONFLARE: CONFormal LArge language model REtrieval 4 Apr 2024 · 1 repository · arXiv:2404.04287
-
Conversational Disease Diagnosis via External Planner-Controlled Large Language Models 4 Apr 2024 · 1 repository · arXiv:2404.04292
-
Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences 4 Apr 2024 · 0 repositories · arXiv:2404.03715
-
Do Large Language Models Rank Fairly? An Empirical Study on the Fairness of LLMs as Rankers 4 Apr 2024 · 0 repositories · arXiv:2404.03192
-
Evaluating LLMs at Detecting Errors in LLM Responses 4 Apr 2024 · 1 repository · arXiv:2404.03602Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
NLP at UC Santa Cruz at SemEval-2024 Task 5: Legal Answer Validation using Few-Shot Multi-Choice QA 4 Apr 2024 · 1 repository · arXiv:2404.03150
-
On the Surprising Efficacy of Distillation as an Alternative to Pre-Training Small Models 4 Apr 2024 · 1 repository · arXiv:2404.03263Syntology official (archive's flag): 4 ran · 4 ran (of which 3 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
On the Theoretical Expressive Power and the Design Space of Higher-Order Graph Transformers 4 Apr 2024 · 1 repository · arXiv:2404.03380Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples) · 5 pointer-only (licence)
-
OpenNeRF: Open Set 3D Neural Scene Segmentation with Pixel-Wise Features and Rendered Novel Views 4 Apr 2024 · 0 repositories · arXiv:2404.03650