Methods › General › Attention Modules › Multi-Head Attention › Papers, page 117
Multi-Head Attention
Papers archive 2025-07-28
archive papers tagged: 24,855 · with a code link: 11,214 · where Syntology ran a sample: 3,454 (2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,454 of 24,855 tagged: 2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument)
Page 117 of 249: papers 11,601 to 11,700 of 24,855, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
DebCSE: Rethinking Unsupervised Contrastive Sentence Embedding Learning in the Debiasing Perspective 14 Sep 2023 · 0 repositories · arXiv:2309.07396
-
Echotune: A Modular Extractor Leveraging the Variable-Length Nature of Speech in ASR Tasks 14 Sep 2023 · 0 repositories · arXiv:2309.07765
-
EnCodecMAE: Leveraging neural codecs for universal audio representation learning 14 Sep 2023 · 2 repositories · arXiv:2309.07391Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Folding Attention: Memory and Power Optimization for On-Device Transformer-based Streaming Speech Recognition 14 Sep 2023 · 0 repositories · arXiv:2309.07988
-
HIGT: Hierarchical Interaction Graph-Transformer for Whole Slide Image Analysis 14 Sep 2023 · 1 repository · arXiv:2309.07400Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Language Embedded Radiance Fields for Zero-Shot Task-Oriented Grasping 14 Sep 2023 · 0 repositories · arXiv:2309.07970
-
Learning Quasi-Static 3D Models of Markerless Deformable Linear Objects for Bimanual Robotic Manipulation 14 Sep 2023 · 1 repository · arXiv:2309.07609
-
Text Classification of Cancer Clinical Trial Eligibility Criteria 14 Sep 2023 · 0 repositories · arXiv:2309.07812
-
Two Timin': Repairing Smart Contracts With A Two-Layered Approach 14 Sep 2023 · 0 repositories · arXiv:2309.07841
-
Virchow: A Million-Slide Digital Pathology Foundation Model 14 Sep 2023 · 1 repository · arXiv:2309.07778Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Aggregating Nearest Sharp Features via Hybrid Transformers for Video Deblurring 13 Sep 2023 · 1 repository · arXiv:2309.07054
-
Benchmarking Procedural Language Understanding for Low-Resource Languages: A Case Study on Turkish 13 Sep 2023 · 1 repository · arXiv:2309.06698
-
CCSPNet-Joint: Efficient Joint Training Method for Traffic Sign Detection Under Extreme Conditions 13 Sep 2023 · 1 repository · arXiv:2309.06902
-
Enhancing Keyphrase Generation by BART Finetuning with Splitting and Shuffling 13 Sep 2023 · 0 repositories · arXiv:2309.06726
-
Generative AI 13 Sep 2023 · 0 repositories · arXiv:2309.07930
-
In-Contextual Gender Bias Suppression for Large Language Models 13 Sep 2023 · 1 repository · arXiv:2309.07251
-
Large Language Models Can Infer Psychological Dispositions of Social Media Users 13 Sep 2023 · 0 repositories · arXiv:2309.08631
-
Neural network-based coronary dominance classification of RCA angiograms 13 Sep 2023 · 0 repositories · arXiv:2309.06958
-
RAIN: Your Language Models Can Align Themselves without Finetuning 13 Sep 2023 · 1 repository · arXiv:2309.07124Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
SafetyBench: Evaluating the Safety of Large Language Models 13 Sep 2023 · 1 repository · arXiv:2309.07045Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
ShaDocFormer: A Shadow-Attentive Threshold Detector With Cascaded Fusion Refiner for Document Shadow Removal 13 Sep 2023 · 1 repository · arXiv:2309.06670
-
Sudden Drops in the Loss: Syntax Acquisition, Phase Transitions, and Simplicity Bias in MLMs 13 Sep 2023 · 1 repository · arXiv:2309.07311Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Traveling Words: A Geometric Interpretation of Transformers 13 Sep 2023 · 1 repository · arXiv:2309.07315
-
Balanced and Explainable Social Media Analysis for Public Health with Large Language Models 12 Sep 2023 · 1 repository · arXiv:2309.05951
-
Neural Network Layer Matrix Decomposition reveals Latent Manifold Encoding and Memory Capacity 12 Sep 2023 · 0 repositories · arXiv:2309.05968
-
Circuit Breaking: Removing Model Behaviors with Targeted Ablation 12 Sep 2023 · 1 repository · arXiv:2309.05973Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
RAP-Gen: Retrieval-Augmented Patch Generation with CodeT5 for Automatic Program Repair 12 Sep 2023 · 0 repositories · arXiv:2309.06057
-
BHASA: A Holistic Southeast Asian Linguistic and Cultural Evaluation Suite for Large Language Models 12 Sep 2023 · 3 repositories · arXiv:2309.06085
-
Characterizing Latent Perspectives of Media Houses Towards Public Figures 12 Sep 2023 · 0 repositories · arXiv:2309.06112
-
Long-term drought prediction using deep neural networks based on geospatial weather data 12 Sep 2023 · 1 repository · arXiv:2309.06212
-
A 3M-Hybrid Model for the Restoration of Unique Giant Murals: A Case Study on the Murals of Yongle Palace 12 Sep 2023 · 0 repositories · arXiv:2309.06194
-
ACT: Empowering Decision Transformer with Dynamic Programming via Advantage Conditioning 12 Sep 2023 · 1 repository · arXiv:2309.05915Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
PRESTI: Predicting Repayment Effort of Self-Admitted Technical Debt Using Textual Information 12 Sep 2023 · 0 repositories · arXiv:2309.06020
-
Comparing Llama-2 and GPT-3 LLMs for HPC kernels generation 12 Sep 2023 · 0 repositories · arXiv:2309.07103
-
Exploring Large Language Models for Ontology Alignment 12 Sep 2023 · 1 repository · arXiv:2309.07172
-
Exploring the Benefits of Differentially Private Pre-training and Parameter-Efficient Fine-tuning for Table Transformers 12 Sep 2023 · 1 repository · arXiv:2309.06526
-
Feature Aggregation Network for Building Extraction from High-resolution Remote Sensing Images 12 Sep 2023 · 0 repositories · arXiv:2309.06017
-
FLDNet: A Foreground-Aware Network for Polyp Segmentation Leveraging Long-Distance Dependencies 12 Sep 2023 · 0 repositories · arXiv:2309.05987
-
Hierarchical Multi-Task Learning Framework for Session-based Recommendations 12 Sep 2023 · 0 repositories · arXiv:2309.06533
-
Breaking through the learning plateaus of in-context learning in Transformer 12 Sep 2023 · 0 repositories · arXiv:2309.06054
-
IBAFormer: Intra-batch Attention Transformer for Domain Generalized Semantic Segmentation 12 Sep 2023 · 0 repositories · arXiv:2309.06282
-
Jersey Number Recognition using Keyframe Identification from Low-Resolution Broadcast Videos 12 Sep 2023 · 0 repositories · arXiv:2309.06285
-
Leveraging Large Language Models and Weak Supervision for Social Media data annotation: an evaluation using COVID-19 self-reported vaccination tweets 12 Sep 2023 · 0 repositories · arXiv:2309.06503
-
Overview of Memotion 3: Sentiment and Emotion Analysis of Codemixed Hinglish Memes 12 Sep 2023 · 0 repositories · arXiv:2309.06517
-
Strategic Behavior of Large Language Models: Game Structure vs. Contextual Framing 12 Sep 2023 · 0 repositories · arXiv:2309.05898
-
The Moral Machine Experiment on Large Language Models 12 Sep 2023 · 1 repository · arXiv:2309.05958
-
Unveiling the potential of large language models in generating semantic and cross-language clones 12 Sep 2023 · 0 repositories · arXiv:2309.06424
-
An Empirical Study of NetOps Capability of Pre-Trained Large Language Models 11 Sep 2023 · 0 repositories · arXiv:2309.05557
-
Applying BioBERT to Extract Germline Gene-Disease Associations for Building a Knowledge Graph from the Biomedical Literature 11 Sep 2023 · 1 repository · arXiv:2309.13061
-
Black-Box Analysis: GPTs Across Time in Legal Textual Entailment Task 11 Sep 2023 · 0 repositories · arXiv:2309.05501
-
Circle Feature Graphormer: Can Circle Features Stimulate Graph Transformer? 11 Sep 2023 · 1 repository · arXiv:2309.06574
-
Toward a Deeper Understanding: RetNet Viewed through Convolution 11 Sep 2023 · 1 repository · arXiv:2309.05375
-
CrisisTransformers: Pre-trained language models and sentence encoders for crisis-related social media texts 11 Sep 2023 · 0 repositories · arXiv:2309.05494
-
Detecting Natural Language Biases with Prompt-based Learning 11 Sep 2023 · 0 repositories · arXiv:2309.05227
-
HAT: Hybrid Attention Transformer for Image Restoration 11 Sep 2023 · 2 repositories · arXiv:2309.05239Syntology community repositories only · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Large Language Model for Science: A Study on P vs. NP 11 Sep 2023 · 1 repository · arXiv:2309.05689
-
Long-Range Transformer Architectures for Document Understanding 11 Sep 2023 · 1 repository · arXiv:2309.05503
-
Memory Injections: Correcting Multi-Hop Reasoning Failures during Inference in Transformer-Based Language Models 11 Sep 2023 · 1 repository · arXiv:2309.05605Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples)
-
Restoring Snow-Degraded Single Images With Wavelet in Vision Transformer 11 Sep 2023 · 2 repositories
-
SparseSwin: Swin Transformer with Sparse Transformer Block 11 Sep 2023 · 1 repository · arXiv:2309.05224
-
Zero-shot Learning with Minimum Instruction to Extract Social Determinants and Family History from Clinical Notes using GPT Model 11 Sep 2023 · 0 repositories · arXiv:2309.05475
-
DeViT: Decomposing Vision Transformers for Collaborative Inference in Edge Devices 10 Sep 2023 · 0 repositories · arXiv:2309.05015
-
Implementing Learning Principles with a Personal AI Tutor: A Case Study 10 Sep 2023 · 0 repositories · arXiv:2309.13060
-
Learning Personalized User Preference from Cold Start in Multi-turn Conversations 10 Sep 2023 · 0 repositories · arXiv:2309.05127
-
Neural-Hidden-CRF: A Robust Weakly-Supervised Sequence Labeler 10 Sep 2023 · 1 repository · arXiv:2309.05086
-
RGAT: A Deeper Look into Syntactic Dependency Information for Coreference Resolution 10 Sep 2023 · 0 repositories · arXiv:2309.04977
-
Unified Contrastive Fusion Transformer for Multimodal Human Action Recognition 10 Sep 2023 · 0 repositories · arXiv:2309.05032
-
DeNoising-MOT: Towards Multiple Object Tracking with Severe Occlusions 9 Sep 2023 · 0 repositories · arXiv:2309.04682
-
Efficient Finetuning Large Language Models For Vietnamese Chatbot 9 Sep 2023 · 0 repositories · arXiv:2309.04646
-
Few-Shot Medical Image Segmentation via a Region-enhanced Prototypical Transformer 9 Sep 2023 · 1 repository · arXiv:2309.04825
-
How to Evaluate Semantic Communications for Images with ViTScore Metric? 9 Sep 2023 · 0 repositories · arXiv:2309.04891
-
Latent Spatiotemporal Adaptation for Generalized Face Forgery Video Detection 9 Sep 2023 · 0 repositories · arXiv:2309.04795
-
Transformer-Based Deep Learning Detector for Dual-Mode Index Modulation 3D-OFDM 9 Sep 2023 · 0 repositories · arXiv:2309.04764
-
Can NLP Models 'Identify', 'Distinguish', and 'Justify' Questions that Don't have a Definitive Answer? 8 Sep 2023 · 0 repositories · arXiv:2309.04635
-
CNN Injected Transformer for Image Exposure Correction 8 Sep 2023 · 1 repository · arXiv:2309.04366
-
Context-Aware Prompt Tuning for Vision-Language Model with Dual-Alignment 8 Sep 2023 · 0 repositories · arXiv:2309.04158
-
Curve Your Attention: Mixed-Curvature Transformers for Graph Representation Learning 8 Sep 2023 · 0 repositories · arXiv:2309.04082
-
Encoding Multi-Domain Scientific Papers by Ensembling Multiple CLS Tokens 8 Sep 2023 · 1 repository · arXiv:2309.04333Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
FIMO: A Challenge Formal Dataset for Automated Theorem Proving 8 Sep 2023 · 1 repository · arXiv:2309.04295
-
From Sparse to Dense: GPT-4 Summarization with Chain of Density Prompting 8 Sep 2023 · 0 repositories · arXiv:2309.04269
-
Fuzzy Fingerprinting Transformer Language-Models for Emotion Recognition in Conversations 8 Sep 2023 · 0 repositories · arXiv:2309.04292
-
Language Prompt for Autonomous Driving 8 Sep 2023 · 1 repository · arXiv:2309.04379Syntology official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 1 pointer-only (licence)
-
Leveraging Pretrained Image-text Models for Improving Audio-Visual Learning 8 Sep 2023 · 0 repositories · arXiv:2309.04628
-
NESTLE: a No-Code Tool for Statistical Analysis of Legal Corpus 8 Sep 2023 · 1 repository · arXiv:2309.04146
-
UQ at #SMM4H 2023: ALEX for Public Health Analysis with Social Media 8 Sep 2023 · 1 repository · arXiv:2309.04213
-
Adapting Self-Supervised Representations to Multi-Domain Setups 7 Sep 2023 · 0 repositories · arXiv:2309.03999
-
Enhancing Pipeline-Based Conversational Agents with Large Language Models 7 Sep 2023 · 0 repositories · arXiv:2309.03748
-
Evaluating ChatGPT as a Recommender System: A Rigorous Approach 7 Sep 2023 · 1 repository · arXiv:2309.03613
-
Supervised Learning and Large Language Model Benchmarks on Mental Health Datasets: Cognitive Distortions and Suicidal Risks in Chinese Social Media 7 Sep 2023 · 2 repositories · arXiv:2309.03564
-
Evaluation of large language models for discovery of gene set function 7 Sep 2023 · 1 repository · arXiv:2309.04019
-
FLM-101B: An Open LLM and How to Train It with $100K Budget 7 Sep 2023 · 0 repositories · arXiv:2309.03852
-
MS-UNet-v2: Adaptive Denoising Method and Training Strategy for Medical Image Segmentation with Small Training Data 7 Sep 2023 · 0 repositories · arXiv:2309.03686
-
MMSFormer: Multimodal Transformer for Material and Semantic Segmentation 7 Sep 2023 · 1 repository · arXiv:2309.04001
-
ProPainter: Improving Propagation and Transformer for Video Inpainting 7 Sep 2023 · 3 repositories · arXiv:2309.03897Syntology official (archive's flag): 10 ran · 20 ran (of which 10 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 8 where Syntology's instrument failed) · 14 unverified (of 34 harvested samples) · 16 pointer-only (licence)
-
S-Adapter: Generalizing Vision Transformer for Face Anti-Spoofing with Statistical Tokens 7 Sep 2023 · 3 repositories · arXiv:2309.04038
-
Zero-Shot Audio Captioning via Audibility Guidance 7 Sep 2023 · 0 repositories · arXiv:2309.03884
-
Certifying LLM Safety against Adversarial Prompting 6 Sep 2023 · 1 repository · arXiv:2309.02705
-
Character Queries: A Transformer-based Approach to On-Line Handwritten Character Segmentation 6 Sep 2023 · 1 repository · arXiv:2309.03072
-
GPT Can Solve Mathematical Problems Without a Calculator 6 Sep 2023 · 1 repository · arXiv:2309.03241Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
HAE-RAE Bench: Evaluation of Korean Knowledge in Language Models 6 Sep 2023 · 1 repository · arXiv:2309.02706