Methods › General › Attention Modules › Multi-Head Attention › Papers, page 73
Multi-Head Attention
Papers archive 2025-07-28
archive papers tagged: 24,855 · with a code link: 11,214 · where Syntology ran a sample: 3,454 (2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,454 of 24,855 tagged: 2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument)
Page 73 of 249: papers 7,201 to 7,300 of 24,855, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
RETTA: Retrieval-Enhanced Test-Time Adaptation for Zero-Shot Video Captioning 11 May 2024 · 0 repositories · arXiv:2405.07046
-
RoTHP: Rotary Position Embedding-based Transformer Hawkes Process 11 May 2024 · 0 repositories · arXiv:2405.06985
-
Super-Resolving Blurry Images with Events 11 May 2024 · 0 repositories · arXiv:2405.06918
-
TacoERE: Cluster-aware Compression for Event Relation Extraction 11 May 2024 · 0 repositories · arXiv:2405.06890
-
A Lightweight Sparse Focus Transformer for Remote Sensing Image Change Captioning 10 May 2024 · 1 repository · arXiv:2405.06598
-
A Survey on RAG Meeting LLMs: Towards Retrieval-Augmented Large Language Models 10 May 2024 · 0 repositories · arXiv:2405.06211
-
An Assessment of Model-On-Model Deception 10 May 2024 · 0 repositories · arXiv:2405.12999
-
CANAL -- Cyber Activity News Alerting Language Model: Empirical Approach vs. Expensive LLM 10 May 2024 · 0 repositories · arXiv:2405.06772
-
Characterizing the Accuracy -- Efficiency Trade-off of Low-rank Decomposition in Language Models 10 May 2024 · 0 repositories · arXiv:2405.06626
-
ChatGPTest: opportunities and cautionary tales of utilizing AI for questionnaire pretesting 10 May 2024 · 0 repositories · arXiv:2405.06329
-
Dual-Task Vision Transformer for Rapid and Accurate Intracerebral Hemorrhage CT Image Classification 10 May 2024 · 0 repositories · arXiv:2405.06814
-
Large Language Model in Financial Regulatory Interpretation 10 May 2024 · 0 repositories · arXiv:2405.06808
-
Mesh Denoising Transformer 10 May 2024 · 0 repositories · arXiv:2405.06536
-
Multimodal LLMs Struggle with Basic Visual Network Analysis: a VNA Benchmark 10 May 2024 · 1 repository · arXiv:2405.06634
-
A Mixture of Experts Approach to 3D Human Motion Prediction 9 May 2024 · 1 repository · arXiv:2405.06088
-
Artificial Intelligence as the New Hacker: Developing Agents for Offensive Security 9 May 2024 · 0 repositories · arXiv:2406.07561
-
Bidirectional Progressive Transformer for Interaction Intention Anticipation 9 May 2024 · 0 repositories · arXiv:2405.05552
-
Can large language models understand uncommon meanings of common words? 9 May 2024 · 0 repositories · arXiv:2405.05741
-
Digital Diagnostics: The Potential Of Large Language Models In Recognizing Symptoms Of Common Illnesses 9 May 2024 · 0 repositories · arXiv:2405.06712
-
Ditto: Quantization-aware Secure Inference of Transformers upon MPC 9 May 2024 · 1 repository · arXiv:2405.05525
-
HMT: Hierarchical Memory Transformer for Long Context Language Processing 9 May 2024 · 1 repository · arXiv:2405.06067Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 5 harvested samples)
-
Iris: An AI-Driven Virtual Tutor For Computer Science Education 9 May 2024 · 0 repositories · arXiv:2405.08008
-
Letter to the Editor: What are the legal and ethical considerations of submitting radiology reports to ChatGPT? 9 May 2024 · 0 repositories · arXiv:2405.05647
-
Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers 9 May 2024 · 2 repositories · arXiv:2405.05945
-
People cannot distinguish GPT-4 from a human in a Turing test 9 May 2024 · 0 repositories · arXiv:2405.08007
-
Reddit-Impacts: A Named Entity Recognition Dataset for Analyzing Clinical and Social Effects of Substance Use Derived from Social Media 9 May 2024 · 0 repositories · arXiv:2405.06145
-
Self-Supervised Learning of Time Series Representation via Diffusion Process and Imputation-Interpolation-Forecasting Mask 9 May 2024 · 2 repositories · arXiv:2405.05959
-
Similarity Guided Multimodal Fusion Transformer for Semantic Location Prediction in Social Media 9 May 2024 · 0 repositories · arXiv:2405.05760
-
Smurfs: Leveraging Multiple Proficiency Agents with Context-Efficiency for Tool Planning 9 May 2024 · 1 repository · arXiv:2405.05955Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Unveiling the Competitive Dynamics: A Comparative Evaluation of American and Chinese LLMs 9 May 2024 · 0 repositories · arXiv:2405.06713
-
Vision-Language Modeling with Regularized Spatial Transformer Networks for All Weather Crosswind Landing of Aircraft 9 May 2024 · 0 repositories · arXiv:2405.05574
-
VM-DDPM: Vision Mamba Diffusion for Medical Image Synthesis 9 May 2024 · 0 repositories · arXiv:2405.05667
-
ACORN: Aspect-wise Commonsense Reasoning Explanation Evaluation 8 May 2024 · 1 repository · arXiv:2405.04818
-
AirGapAgent: Protecting Privacy-Conscious Conversational Agents 8 May 2024 · 0 repositories · arXiv:2405.05175
-
CARE-SD: Classifier-based analysis for recognizing and eliminating stigmatizing and doubt marker labels in electronic health records: model development and validation 8 May 2024 · 1 repository · arXiv:2405.05204
-
Evaluating Students' Open-ended Written Responses with LLMs: Using the RAG Framework for GPT-3.5, GPT-4, Claude-3, and Mistral-Large 8 May 2024 · 0 repositories · arXiv:2405.05444
-
Few-Shot Class Incremental Learning via Robust Transformer Approach 8 May 2024 · 1 repository · arXiv:2405.05984
-
HMANet: Hybrid Multi-Axis Aggregation Network for Image Super-Resolution 8 May 2024 · 1 repository · arXiv:2405.05001
-
LLMs Can Patch Up Missing Relevance Judgments in Evaluation 8 May 2024 · 0 repositories · arXiv:2405.04727
-
Multi-scale Bottleneck Transformer for Weakly Supervised Multimodal Violence Detection 8 May 2024 · 1 repository · arXiv:2405.05130
-
Open Source Language Models Can Provide Feedback: Evaluating LLMs' Ability to Help Students Using GPT-4-As-A-Judge 8 May 2024 · 1 repository · arXiv:2405.05253
-
Seeds of Stereotypes: A Large-Scale Textual Analysis of Race and Gender Associations with Diseases in Online Sources 8 May 2024 · 0 repositories · arXiv:2405.05049
-
Transformer Architecture for NetsDB 8 May 2024 · 1 repository · arXiv:2405.04807
-
Utilizing Large Language Models to Generate Synthetic Data to Increase the Performance of BERT-Based Neural Networks 8 May 2024 · 0 repositories · arXiv:2405.06695
-
You Only Cache Once: Decoder-Decoder Architectures for Language Models 8 May 2024 · 1 repository · arXiv:2405.05254
-
An Advanced Features Extraction Module for Remote Sensing Image Super-Resolution 7 May 2024 · 0 repositories · arXiv:2405.04595
-
An LLM-Tool Compiler for Fused Parallel Function Calling 7 May 2024 · 0 repositories · arXiv:2405.17438
-
D-TrAttUnet: Toward Hybrid CNN-Transformer Architecture for Generic and Subtle Segmentation in Medical Images 7 May 2024 · 1 repository · arXiv:2405.04169
-
Enhancing the Efficiency and Accuracy of Underlying Asset Reviews in Structured Finance: The Application of Multi-agent Framework 7 May 2024 · 1 repository · arXiv:2405.04294
-
Enriched BERT Embeddings for Scholarly Publication Classification 7 May 2024 · 1 repository · arXiv:2405.04136
-
ERATTA: Extreme RAG for Table To Answers with Large Language Models 7 May 2024 · 0 repositories · arXiv:2405.03963
-
Evaluating Text Summaries Generated by Large Language Models Using OpenAI's GPT 7 May 2024 · 0 repositories · arXiv:2405.04053
-
Folded Context Condensation in Path Integral Formalism for Infinite Context Transformers 7 May 2024 · 0 repositories · arXiv:2405.04620
-
GPT-Enabled Cybersecurity Training: A Tailored Approach for Effective Awareness 7 May 2024 · 0 repositories · arXiv:2405.04138
-
HAFFormer: A Hierarchical Attention-Free Framework for Alzheimer's Disease Detection From Spontaneous Speech 7 May 2024 · 0 repositories · arXiv:2405.03952
-
How does GPT-2 Predict Acronyms? Extracting and Understanding a Circuit via Mechanistic Interpretability 7 May 2024 · 1 repository · arXiv:2405.04156Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Learning Linear Block Error Correction Codes 7 May 2024 · 1 repository · arXiv:2405.04050
-
Long Context Alignment with Short Instructions and Synthesized Positions 7 May 2024 · 0 repositories · arXiv:2405.03939
-
Masked Graph Transformer for Large-Scale Recommendation 7 May 2024 · 0 repositories · arXiv:2405.04028
-
MEDVOC: Vocabulary Adaptation for Fine-tuning Pre-trained Language Models on Medical Text Summarization 7 May 2024 · 1 repository · arXiv:2405.04163Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 1 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts 7 May 2024 · 1 repository · arXiv:2405.04520
-
POV Learning: Individual Alignment of Multimodal Models using Human Perception 7 May 2024 · 0 repositories · arXiv:2405.04443
-
Predictive Modeling with Temporal Graphical Representation on Electronic Health Records 7 May 2024 · 1 repository · arXiv:2405.03943Syntology official (archive's flag): 10 ran · 10 ran (of which 3 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Remote Diffusion 7 May 2024 · 0 repositories · arXiv:2405.04717
-
Revisiting Character-level Adversarial Attacks for Language Models 7 May 2024 · 1 repository · arXiv:2405.04346Syntology official (archive's flag): 23 ran · 23 ran (of which 3 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 10 where Syntology's instrument failed) · 8 unverified (of 31 harvested samples)
-
Robust Implementation of Retrieval-Augmented Generation on Edge-based Computing-in-Memory Architectures 7 May 2024 · 0 repositories · arXiv:2405.04700
-
S3Former: Self-supervised High-resolution Transformer for Solar PV Profiling 7 May 2024 · 0 repositories · arXiv:2405.04489
-
Structured Click Control in Transformer-based Interactive Segmentation 7 May 2024 · 1 repository · arXiv:2405.04009Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 0 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
SUTRA: Scalable Multilingual Language Model Architecture 7 May 2024 · 0 repositories · arXiv:2405.06694
-
The Silicon Ceiling: Auditing GPT's Race and Gender Biases in Hiring 7 May 2024 · 0 repositories · arXiv:2405.04412
-
Utilizing GPT to Enhance Text Summarization: A Strategy to Minimize Hallucinations 7 May 2024 · 0 repositories · arXiv:2405.04039
-
Vision Mamba: A Comprehensive Survey and Taxonomy 7 May 2024 · 1 repository · arXiv:2405.04404
-
xLSTM: Extended Long Short-Term Memory 7 May 2024 · 5 repositories · arXiv:2405.04517Syntology community repositories only · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples)
-
AlphaMath Almost Zero: Process Supervision without Process 6 May 2024 · 1 repository · arXiv:2405.03553
-
Anchored Answers: Unravelling Positional Bias in GPT-2's Multiple-Choice Questions 6 May 2024 · 1 repository · arXiv:2405.03205
-
Characterizing the Dilemma of Performance and Index Size in Billion-Scale Vector Search and Breaking It with Second-Tier Memory 6 May 2024 · 0 repositories · arXiv:2405.03267
-
Class-relevant Patch Embedding Selection for Few-Shot Image Classification 6 May 2024 · 0 repositories · arXiv:2405.03722
-
Compressing Long Context for Enhancing RAG with AMR-based Concept Distillation 6 May 2024 · 0 repositories · arXiv:2405.03085
-
CRA5: Extreme Compression of ERA5 for Portable Global Climate and Weather Research via an Efficient Variational Transformer 6 May 2024 · 1 repository · arXiv:2405.03376
-
Swin transformers are robust to distribution and concept drift in endoscopy-based longitudinal rectal cancer assessment 6 May 2024 · 0 repositories · arXiv:2405.03762
-
Detecting Android Malware: From Neural Embeddings to Hands-On Validation with BERTroid 6 May 2024 · 0 repositories · arXiv:2405.03620
-
Detecting Anti-Semitic Hate Speech using Transformer-based Large Language Models 6 May 2024 · 0 repositories · arXiv:2405.03794
-
Dual Relation Mining Network for Zero-Shot Learning 6 May 2024 · 0 repositories · arXiv:2405.03613
-
Enhancing DETRs Variants through Improved Content Query and Similar Query Aggregation 6 May 2024 · 0 repositories · arXiv:2405.03318
-
ERAGent: Enhancing Retrieval-Augmented Language Models with Improved Accuracy, Efficiency, and Personalization 6 May 2024 · 1 repository · arXiv:2405.06683
-
Exploring the Frontiers of Softmax: Provable Optimization, Applications in Diffusion Model, and Beyond 6 May 2024 · 0 repositories · arXiv:2405.03251
-
GREEN: Generative Radiology Report Evaluation and Error Notation 6 May 2024 · 0 repositories · arXiv:2405.03595
-
Hire Me or Not? Examining Language Model's Behavior with Occupation Attributes 6 May 2024 · 1 repository · arXiv:2405.06687Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
Intra-task Mutual Attention based Vision Transformer for Few-Shot Learning 6 May 2024 · 0 repositories · arXiv:2405.03109
-
Investigating Personalized Driving Behaviors in Dilemma Zones: Analysis and Prediction of Stop-or-Go Decisions 6 May 2024 · 0 repositories · arXiv:2405.03873
-
Large Language Models Reveal Information Operation Goals, Tactics, and Narrative Frames 6 May 2024 · 1 repository · arXiv:2405.03688
-
MAmmoTH2: Scaling Instructions from the Web 6 May 2024 · 0 repositories · arXiv:2405.03548
-
Modality Prompts for Arbitrary Modality Salient Object Detection 6 May 2024 · 0 repositories · arXiv:2405.03351
-
ReCycle: Fast and Efficient Long Time Series Forecasting with Residual Cyclic Transformers 6 May 2024 · 1 repository · arXiv:2405.03429
-
Salient Object Detection From Arbitrary Modalities 6 May 2024 · 1 repository · arXiv:2405.03352
-
SocialFormer: Social Interaction Modeling with Edge-enhanced Heterogeneous Graph Transformers for Trajectory Prediction 6 May 2024 · 0 repositories · arXiv:2405.03809
-
Transformer-based RGB-T Tracking with Channel and Spatial Feature Fusion 6 May 2024 · 1 repository · arXiv:2405.03177
-
Transformer models as an efficient replacement for statistical test suites to evaluate the quality of random numbers 6 May 2024 · 0 repositories · arXiv:2405.03904
-
Can Large Language Models Make the Grade? An Empirical Study Evaluating LLMs Ability to Mark Short Answer Questions in K-12 Education 5 May 2024 · 0 repositories · arXiv:2405.02985
-
E-TSL: A Continuous Educational Turkish Sign Language Dataset with Baseline Methods 5 May 2024 · 0 repositories · arXiv:2405.02984