Methods › General › Attention Modules › Multi-Head Attention › Papers, page 80
Multi-Head Attention
Papers archive 2025-07-28
archive papers tagged: 24,855 · with a code link: 11,214 · where Syntology ran a sample: 3,454 (2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,454 of 24,855 tagged: 2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument)
Page 80 of 249: papers 7,901 to 8,000 of 24,855, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Revealing Trends in Datasets from the 2022 ACL and EMNLP Conferences 31 Mar 2024 · 0 repositories · arXiv:2404.08666
-
RQ-RAG: Learning to Refine Queries for Retrieval Augmented Generation 31 Mar 2024 · 1 repository · arXiv:2404.00610
-
Training-Free Semantic Segmentation via LLM-Supervision 31 Mar 2024 · 0 repositories · arXiv:2404.00701
-
Transformer based Pluralistic Image Completion with Reduced Information Loss 31 Mar 2024 · 1 repository · arXiv:2404.00513
-
A Comprehensive Study on NLP Data Augmentation for Hate Speech Detection: Legacy Methods, BERT, and LLMs 30 Mar 2024 · 0 repositories · arXiv:2404.00303
-
A Novel Feature Map Enhancement Technique Integrating Residual CNN and Transformer for Alzheimer Diseases Diagnosis 30 Mar 2024 · 0 repositories · arXiv:2405.12986
-
Can LLMs Master Math? Investigating Large Language Models on Math Stack Exchange 30 Mar 2024 · 1 repository · arXiv:2404.00344Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 4 harvested samples)
-
Dependability Evaluation of Stable Diffusion with Soft Errors on the Model Parameters 30 Mar 2024 · 0 repositories · arXiv:2404.00352
-
Edinburgh Clinical NLP at SemEval-2024 Task 2: Fine-tune your model unless you have access to GPT-4 30 Mar 2024 · 1 repository · arXiv:2404.00484
-
Injecting New Knowledge into Large Language Models via Supervised Fine-Tuning 30 Mar 2024 · 0 repositories · arXiv:2404.00213
-
Leveraging Pre-trained and Transformer-derived Embeddings from EHRs to Characterize Heterogeneity Across Alzheimer's Disease and Related Dementias 30 Mar 2024 · 0 repositories · arXiv:2404.00464
-
Small Language Models Learn Enhanced Reasoning Skills from Medical Textbooks 30 Mar 2024 · 0 repositories · arXiv:2404.00376
-
Spread Your Wings: A Radial Strip Transformer for Image Deblurring 30 Mar 2024 · 0 repositories · arXiv:2404.00358
-
A Parallel Attention Network for Cattle Face Recognition 29 Mar 2024 · 0 repositories · arXiv:2403.19980
-
AgileFormer: Spatially Agile Transformer UNet for Medical Image Segmentation 29 Mar 2024 · 1 repository · arXiv:2404.00122
-
ChatGPT v.s. Media Bias: A Comparative Study of GPT-3.5 and Fine-tuned Language Models 29 Mar 2024 · 0 repositories · arXiv:2403.20158
-
Classifying Conspiratorial Narratives At Scale: False Alarms and Erroneous Connections 29 Mar 2024 · 1 repository · arXiv:2404.00141
-
DataAgent: Evaluating Large Language Models' Ability to Answer Zero-Shot, Natural Language Queries 29 Mar 2024 · 0 repositories · arXiv:2404.00188
-
Decision Mamba: Reinforcement Learning via Sequence Modeling with Selective State Spaces 29 Mar 2024 · 1 repository · arXiv:2403.19925Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
DiJiang: Efficient Large Language Models through Compact Kernelization 29 Mar 2024 · 1 repository · arXiv:2403.19928Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Enhancing the General Agent Capabilities of Low-Parameter LLMs through Tuning and Multi-Branch Reasoning 29 Mar 2024 · 1 repository · arXiv:2403.19962Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Explainable Deep Learning: A Visual Analytics Approach with Transition Matrices 29 Mar 2024 · 1 repository
-
Latxa: An Open Language Model and Evaluation Suite for Basque 29 Mar 2024 · 1 repository · arXiv:2403.20266
-
LayerNorm: A key component in parameter-efficient fine-tuning 29 Mar 2024 · 0 repositories · arXiv:2403.20284
-
Localising the Seizure Onset Zone from Single-Pulse Electrical Stimulation Responses with a CNN Transformer 29 Mar 2024 · 1 repository · arXiv:2403.20324
-
MANGO: A Benchmark for Evaluating Mapping and Navigation Abilities of Large Language Models 29 Mar 2024 · 1 repository · arXiv:2403.19913Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
On-the-fly Definition Augmentation of LLMs for Biomedical NER 29 Mar 2024 · 1 repository · arXiv:2404.00152Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples)
-
ReALM: Reference Resolution As Language Modeling 29 Mar 2024 · 0 repositories · arXiv:2403.20329
-
SceneTracker: Long-term Scene Flow Estimation Network 29 Mar 2024 · 1 repository · arXiv:2403.19924Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Shallow Cross-Encoders for Low-Latency Retrieval 29 Mar 2024 · 1 repository · arXiv:2403.20222
-
TDANet: A Novel Temporal Denoise Convolutional Neural Network With Attention for Fault Diagnosis 29 Mar 2024 · 0 repositories · arXiv:2403.19943
-
A Novel Stochastic Transformer-based Approach for Post-Traumatic Stress Disorder Detection using Audio Recording of Clinical Interviews 28 Mar 2024 · 0 repositories · arXiv:2403.19441
-
A Review of Multi-Modal Large Language and Vision Models 28 Mar 2024 · 0 repositories · arXiv:2404.01322
-
AAPMT: AGI Assessment Through Prompt and Metric Transformer 28 Mar 2024 · 1 repository · arXiv:2403.19101
-
AlloyBERT: Alloy Property Prediction with Large Language Models 28 Mar 2024 · 0 repositories · arXiv:2403.19783
-
Are Large Language Models Good at Utility Judgments? 28 Mar 2024 · 1 repository · arXiv:2403.19216Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Checkpoint Merging via Bayesian Optimization in LLM Pretraining 28 Mar 2024 · 0 repositories · arXiv:2403.19390
-
Code Comparison Tuning for Code Large Language Models 28 Mar 2024 · 0 repositories · arXiv:2403.19121
-
DenseNets Reloaded: Paradigm Shift Beyond ResNets and ViTs 28 Mar 2024 · 3 repositories · arXiv:2403.19588Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 3 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Enhancing Efficiency in Vision Transformer Networks: Design Techniques and Insights 28 Mar 2024 · 0 repositories · arXiv:2403.19882
-
FACTOID: FACtual enTailment fOr hallucInation Detection 28 Mar 2024 · 0 repositories · arXiv:2403.19113
-
Generating Multi-Aspect Queries for Conversational Search 28 Mar 2024 · 0 repositories · arXiv:2403.19302
-
Genetic Quantization-Aware Approximation for Non-Linear Operations in Transformers 28 Mar 2024 · 1 repository · arXiv:2403.19591
-
Intelligent Classification and Personalized Recommendation of E-commerce Products Based on Machine Learning 28 Mar 2024 · 0 repositories · arXiv:2403.19345
-
Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models 28 Mar 2024 · 1 repository · arXiv:2403.19521Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 14 harvested samples) · 14 pointer-only (licence)
-
Jamba: A Hybrid Transformer-Mamba Language Model 28 Mar 2024 · 3 repositories · arXiv:2403.19887
-
Just-DNA-Seq, open-source personal genomics platform: longevity science for everyone 28 Mar 2024 · 0 repositories · arXiv:2403.19087
-
Keypoint Action Tokens Enable In-Context Imitation Learning in Robotics 28 Mar 2024 · 0 repositories · arXiv:2403.19578
-
MATEval: A Multi-Agent Discussion Framework for Advancing Open-Ended Text Evaluation 28 Mar 2024 · 1 repository · arXiv:2403.19305
-
Mitigating Misleading Chain-of-Thought Reasoning with Selective Filtering 28 Mar 2024 · 1 repository · arXiv:2403.19167
-
Patch Spatio-Temporal Relation Prediction for Video Anomaly Detection 28 Mar 2024 · 0 repositories · arXiv:2403.19111
-
Risk prediction of pathological gambling on social media 28 Mar 2024 · 0 repositories · arXiv:2403.19358
-
Siamese Vision Transformers are Scalable Audio-visual Learners 28 Mar 2024 · 1 repository · arXiv:2403.19638Syntology official (archive's flag): 15 ran · 15 ran (of which 3 constructed an object rather than computing a result; 15 with no instrument failure: 0 honoured, 0 violated, 15 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 19 harvested samples) · 19 pointer-only (licence)
-
Single-Shared Network with Prior-Inspired Loss for Parameter-Efficient Multi-Modal Imaging Skin Lesion Classification 28 Mar 2024 · 0 repositories · arXiv:2403.19203
-
A Novel Corpus of Annotated Medical Imaging Reports and Information Extraction Results Using BERT-based Language Models 27 Mar 2024 · 1 repository · arXiv:2403.18975
-
A Survey on Large Language Models from Concept to Implementation 27 Mar 2024 · 0 repositories · arXiv:2403.18969
-
AcTED: Automatic Acquisition of Typical Event Duration for Semi-supervised Temporal Commonsense QA 27 Mar 2024 · 0 repositories · arXiv:2403.18504
-
Attention-aware semantic relevance predicting Chinese sentence reading 27 Mar 2024 · 0 repositories · arXiv:2403.18542
-
BioMedLM: A 2.7B Parameter Language Model Trained On Biomedical Text 27 Mar 2024 · 1 repository · arXiv:2403.18421Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
BLADE: Enhancing Black-box Large Language Models with Small Domain-Specific Models 27 Mar 2024 · 0 repositories · arXiv:2403.18365
-
Boosting Conversational Question Answering with Fine-Grained Retrieval-Augmentation and Self-Check 27 Mar 2024 · 0 repositories · arXiv:2403.18243
-
CPR: Retrieval Augmented Generation for Copyright Protection 27 Mar 2024 · 0 repositories · arXiv:2403.18920
-
Cross-domain Fiber Cluster Shape Analysis for Language Performance Cognitive Score Prediction 27 Mar 2024 · 0 repositories · arXiv:2403.19001
-
ECoDepth: Effective Conditioning of Diffusion Models for Monocular Depth Estimation 27 Mar 2024 · 1 repository · arXiv:2403.18807Syntology official (archive's flag): 11 ran · 11 ran (of which 1 constructed an object rather than computing a result; 6 with no instrument failure: 2 honoured, 2 violated, 2 with no contract checked; 5 where Syntology's instrument failed) · 4 unverified (of 15 harvested samples) · 15 pointer-only (licence)
-
Evaluating Large Language Models for Health-Related Text Classification Tasks with Public Social Media Data 27 Mar 2024 · 0 repositories · arXiv:2403.19031
-
Faster Convergence for Transformer Fine-tuning with Line Search Methods 27 Mar 2024 · 1 repository · arXiv:2403.18506
-
Cross-System Categorization of Abnormal Traces in Microservice-Based Systems via Meta-Learning 27 Mar 2024 · 0 repositories · arXiv:2403.18998
-
Fourier or Wavelet bases as counterpart self-attention in spikformer for efficient visual classification 27 Mar 2024 · 0 repositories · arXiv:2403.18228
-
Fusion approaches for emotion recognition from speech using acoustic and text-based features 27 Mar 2024 · 0 repositories · arXiv:2403.18635
-
Illicit object detection in X-ray images using Vision Transformers 27 Mar 2024 · 0 repositories · arXiv:2403.19043
-
Lift3D: Zero-Shot Lifting of Any 2D Vision Model to 3D 27 Mar 2024 · 0 repositories · arXiv:2403.18922
-
LLMs in HCI Data Work: Bridging the Gap Between Information Retrieval and Responsible Research Practices 27 Mar 2024 · 0 repositories · arXiv:2403.18173
-
Long-form factuality in large language models 27 Mar 2024 · 3 repositories · arXiv:2403.18802Syntology official (archive's flag): 4 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 4 pointer-only (licence)
-
mALBERT: Is a Compact Multilingual BERT Model Still Worth It? 27 Mar 2024 · 0 repositories · arXiv:2403.18338
-
Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models 27 Mar 2024 · 2 repositories · arXiv:2403.18814Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 1 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples) · 1 pointer-only (licence)
-
Multi-Layer Dense Attention Decoder for Polyp Segmentation 27 Mar 2024 · 1 repository · arXiv:2403.18180
-
ParCo: Part-Coordinating Text-to-Motion Synthesis 27 Mar 2024 · 1 repository · arXiv:2403.18512
-
RankMamba: Benchmarking Mamba's Document Ranking Performance in the Era of Transformers 27 Mar 2024 · 1 repository · arXiv:2403.18276
-
Reshaping Free-Text Radiology Notes Into Structured Reports With Generative Transformers 27 Mar 2024 · 1 repository · arXiv:2403.18938
-
SemRoDe: Macro Adversarial Training to Learn Representations That are Robust to Word-Level Attacks 27 Mar 2024 · 1 repository · arXiv:2403.18423Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples)
-
ViTAR: Vision Transformer with Any Resolution 27 Mar 2024 · 0 repositories · arXiv:2403.18361
-
Vulnerability Detection with Code Language Models: How Far Are We? 27 Mar 2024 · 1 repository · arXiv:2403.18624Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples)
-
Accuracy enhancement method for speech emotion recognition from spectrogram using temporal frequency correlation and positional information learning through knowledge transfer 26 Mar 2024 · 1 repository · arXiv:2403.17327
-
ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV Caching 26 Mar 2024 · 0 repositories · arXiv:2403.17312
-
Are Compressed Language Models Less Subgroup Robust? 26 Mar 2024 · 1 repository · arXiv:2403.17811
-
Automated Report Generation for Lung Cytological Images Using a CNN Vision Classifier and Multiple-Transformer Text Decoders: Preliminary Study 26 Mar 2024 · 0 repositories · arXiv:2403.18151
-
CCDSReFormer: Traffic Flow Prediction with a Criss-Crossed Dual-Stream Enhanced Rectified Transformer Model 26 Mar 2024 · 0 repositories · arXiv:2403.17753
-
Constructions Are So Difficult That Even Large Language Models Get Them Right for the Wrong Reasons 26 Mar 2024 · 1 repository · arXiv:2403.17760
-
Decoding Probing: Revealing Internal Linguistic Structures in Neural Language Models using Minimal Pairs 26 Mar 2024 · 0 repositories · arXiv:2403.17299
-
Disambiguate Entity Matching using Large Language Models through Relation Discovery 26 Mar 2024 · 0 repositories · arXiv:2403.17344
-
Don't Trust: Verify -- Grounding LLM Quantitative Reasoning with Autoformalization 26 Mar 2024 · 1 repository · arXiv:2403.18120Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 8 where Syntology's instrument failed) · 2 unverified (of 12 harvested samples)
-
EgoPoseFormer: A Simple Baseline for Stereo Egocentric 3D Human Pose Estimation 26 Mar 2024 · 1 repository · arXiv:2403.18080
-
ELLEN: Extremely Lightly Supervised Learning For Efficient Named Entity Recognition 26 Mar 2024 · 1 repository · arXiv:2403.17385
-
Enhancing Legal Document Retrieval: A Multi-Phase Approach with Large Language Models 26 Mar 2024 · 0 repositories · arXiv:2403.18093
-
Evaluating the Efficacy of Prompt-Engineered Large Multimodal Models Versus Fine-Tuned Vision Transformers in Image-Based Security Applications 26 Mar 2024 · 0 repositories · arXiv:2403.17787
-
Fingerprinting web servers through Transformer-encoded HTTP response headers 26 Mar 2024 · 1 repository · arXiv:2404.00056
-
Hierarchical Light Transformer Ensembles for Multimodal Trajectory Forecasting 26 Mar 2024 · 1 repository · arXiv:2403.17678
-
Hierarchical Multi-label Classification for Fine-level Event Extraction from Aviation Accident Reports 26 Mar 2024 · 0 repositories · arXiv:2403.17914
-
Integrative Graph-Transformer Framework for Histopathology Whole Slide Image Representation and Classification 26 Mar 2024 · 0 repositories · arXiv:2403.18134
-
InternLM2 Technical Report 26 Mar 2024 · 3 repositories · arXiv:2403.17297Syntology 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)