Methods › General › Attention Modules › Multi-Head Attention › Papers, page 135
Multi-Head Attention
Papers archive 2025-07-28
archive papers tagged: 24,855 · with a code link: 11,214 · where Syntology ran a sample: 3,454 (2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,454 of 24,855 tagged: 2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument)
Page 135 of 249: papers 13,401 to 13,500 of 24,855, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
What Makes for Good Visual Tokenizers for Large Language Models? 20 May 2023 · 1 repository · arXiv:2305.12223Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
A Sequence-to-Sequence Approach for Arabic Pronoun Resolution 19 May 2023 · 0 repositories · arXiv:2305.11529
-
AutoTrial: Prompting Language Models for Clinical Trial Design 19 May 2023 · 0 repositories · arXiv:2305.11366
-
Brain Captioning: Decoding human brain activity into images and text 19 May 2023 · 1 repository · arXiv:2305.11560
-
CCT-Code: Cross-Consistency Training for Multilingual Clone Detection and Code Search 19 May 2023 · 0 repositories · arXiv:2305.11626
-
Clinical Camel: An Open Expert-Level Medical Language Model with Dialogue-Based Knowledge Encoding 19 May 2023 · 2 repositories · arXiv:2305.12031Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Examining Inter-Consistency of Large Language Models Collaboration: An In-depth Analysis via Debate 19 May 2023 · 1 repository · arXiv:2305.11595Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Do Models Really Learn to Follow Instructions? An Empirical Study of Instruction Tuning 19 May 2023 · 0 repositories · arXiv:2305.11383
-
Efficient Mixed Transformer for Single Image Super-Resolution 19 May 2023 · 2 repositories · arXiv:2305.11403
-
Enhancing Short-Term Wind Speed Forecasting using Graph Attention and Frequency-Enhanced Mechanisms 19 May 2023 · 0 repositories · arXiv:2305.11526
-
Exploring the Upper Limits of Text-Based Collaborative Filtering Using Large Language Models: Discoveries and Insights 19 May 2023 · 0 repositories · arXiv:2305.11700
-
Eye-SpatialNet: Spatial Information Extraction from Ophthalmology Notes 19 May 2023 · 0 repositories · arXiv:2305.11948
-
Federated Foundation Models: Privacy-Preserving and Collaborative Learning for Large Models 19 May 2023 · 0 repositories · arXiv:2305.11414
-
Graph Propagation Transformer for Graph Representation Learning 19 May 2023 · 1 repository · arXiv:2305.11424Syntology official (archive's flag): 4 ran · 4 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified; every one of the 4 samples that ran constructed an object rather than computing a result (of 8 harvested samples) · 8 pointer-only (licence)
-
How Does Generative Retrieval Scale to Millions of Passages? 19 May 2023 · 0 repositories · arXiv:2305.11841
-
Language-Universal Phonetic Representation in Multilingual Speech Pretraining for Low-Resource Speech Recognition 19 May 2023 · 0 repositories · arXiv:2305.11569
-
Multimodal Web Navigation with Instruction-Finetuned Foundation Models 19 May 2023 · 0 repositories · arXiv:2305.11854
-
OL-Transformer: A Fast and Universal Surrogate Simulator for Optical Multilayer Thin Film Structures 19 May 2023 · 0 repositories · arXiv:2305.11984
-
PointGPT: Auto-regressively Generative Pre-training from Point Clouds 19 May 2023 · 1 repository · arXiv:2305.11487Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples) · 4 pointer-only (licence)
-
Reciprocal Attention Mixing Transformer for Lightweight Image Restoration 19 May 2023 · 1 repository · arXiv:2305.11474
-
Recycle-and-Distill: Universal Compression Strategy for Transformer-based Speech SSL Models with Attention Map Reusing and Masking Distillation 19 May 2023 · 1 repository · arXiv:2305.11685
-
Reducing Sequence Length by Predicting Edit Operations with Large Language Models 19 May 2023 · 0 repositories · arXiv:2305.11862
-
Scaling laws for language encoding models in fMRI 19 May 2023 · 1 repository · arXiv:2305.11863
-
SeeGULL: A Stereotype Benchmark with Broad Geo-Cultural Coverage Leveraging Generative Models 19 May 2023 · 1 repository · arXiv:2305.11840
-
Self-Agreement: A Framework for Fine-tuning Language Models to Find Agreement among Diverse Opinions 19 May 2023 · 0 repositories · arXiv:2305.11460
-
Self-QA: Unsupervised Knowledge Guided Language Model Alignment 19 May 2023 · 1 repository · arXiv:2305.11952
-
Hint of Thought prompting: an explainable and zero-shot approach to reasoning tasks with LLMs 19 May 2023 · 0 repositories · arXiv:2305.11461
-
Surgical-VQLA: Transformer with Gated Vision-Language Embedding for Visual Question Localized-Answering in Robotic Surgery 19 May 2023 · 2 repositories · arXiv:2305.11692Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples) · 4 pointer-only (licence)
-
A Comparative Study on E-Branchformer vs Conformer in Speech Recognition, Translation, and Understanding Tasks 18 May 2023 · 2 repositories · arXiv:2305.11073
-
Ahead-of-Time P-Tuning 18 May 2023 · 0 repositories · arXiv:2305.10835
-
AIwriting: Relations Between Image Generation and Digital Writing 18 May 2023 · 0 repositories · arXiv:2305.10834
-
Annotation-free Audio-Visual Segmentation 18 May 2023 · 0 repositories · arXiv:2305.11019
-
BioAug: Conditional Generation based Data Augmentation for Low-Resource Biomedical NER 18 May 2023 · 1 repository · arXiv:2305.10647
-
Boost Vision Transformer with GPU-Friendly Sparsity and Quantization 18 May 2023 · 0 repositories · arXiv:2305.10727
-
Comparing Machines and Children: Using Developmental Psychology Experiments to Assess the Strengths and Weaknesses of LaMDA Responses 18 May 2023 · 0 repositories · arXiv:2305.11243
-
Coordinated Transformer with Position & Sample-aware Central Loss for Anatomical Landmark Detection 18 May 2023 · 0 repositories · arXiv:2305.11338
-
Deep Learning Methods for Extracting Metaphorical Names of Flowers and Plants 18 May 2023 · 0 repositories · arXiv:2305.10833
-
Ditto: A Simple and Efficient Approach to Improve Sentence Embeddings 18 May 2023 · 1 repository · arXiv:2305.10786
-
Emergent Representations of Program Semantics in Language Models Trained on Programs 18 May 2023 · 1 repository · arXiv:2305.11169Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 12 where Syntology's instrument failed) · 1 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
FunASR: A Fundamental End-to-End Speech Recognition Toolkit 18 May 2023 · 1 repository · arXiv:2305.11013
-
Generalized Multiple Intent Conditioned Slot Filling 18 May 2023 · 0 repositories · arXiv:2305.11023
-
Generalized Planning in PDDL Domains with Pretrained Large Language Models 18 May 2023 · 1 repository · arXiv:2305.11014Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Large Language Models can be Guided to Evade AI-Generated Text Detection 18 May 2023 · 1 repository · arXiv:2305.10847
-
Less is More! A slim architecture for optimal language translation 18 May 2023 · 0 repositories · arXiv:2305.10991
-
LIMA: Less Is More for Alignment 18 May 2023 · 5 repositories · arXiv:2305.11206
-
mLongT5: A Multilingual and Efficient Text-To-Text Transformer for Longer Sequences 18 May 2023 · 1 repository · arXiv:2305.11129
-
MolXPT: Wrapping Molecules with Text for Generative Pre-training 18 May 2023 · 1 repository · arXiv:2305.10688
-
Brain Imaging-to-Graph Generation using Adversarial Hierarchical Diffusion Models for MCI Causality Analysis 18 May 2023 · 0 repositories · arXiv:2305.10754
-
PDP: Parameter-free Differentiable Pruning is All You Need 18 May 2023 · 0 repositories · arXiv:2305.11203
-
Selecting Learnable Training Samples is All DETRs Need in Crowded Pedestrian Detection 18 May 2023 · 0 repositories · arXiv:2305.10801
-
Support for Stock Trend Prediction Using Transformers and Sentiment Analysis 18 May 2023 · 0 repositories · arXiv:2305.14368
-
TextDiffuser: Diffusion Models as Text Painters 18 May 2023 · 0 repositories · arXiv:2305.10855
-
Trading Syntax Trees for Wordpieces: Target-oriented Opinion Words Extraction with Wordpieces and Aspect Enhancement 18 May 2023 · 0 repositories · arXiv:2305.11034
-
Vaxformer: Antigenicity-controlled Transformer for Vaccine Design Against SARS-CoV-2 18 May 2023 · 1 repository · arXiv:2305.11194
-
A quantitative study of NLP approaches to question difficulty estimation 17 May 2023 · 1 repository · arXiv:2305.10236
-
A survey of the Vision Transformers and their CNN-Transformer based Variants 17 May 2023 · 0 repositories · arXiv:2305.09880
-
AD-KD: Attribution-Driven Knowledge Distillation for Language Model Compression 17 May 2023 · 1 repository · arXiv:2305.10010Syntology official (archive's flag): 3 ran · 3 ran (of which 1 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
CageViT: Convolutional Activation Guided Efficient Vision Transformer 17 May 2023 · 0 repositories · arXiv:2305.09924
-
CoEdIT: Text Editing by Task-Specific Instruction Tuning 17 May 2023 · 1 repository · arXiv:2305.09857
-
CostFormer:Cost Transformer for Cost Aggregation in Multi-view Stereo 17 May 2023 · 0 repositories · arXiv:2305.10320
-
EENED: End-to-End Neural Epilepsy Detection based on Convolutional Transformer 17 May 2023 · 0 repositories · arXiv:2305.10502
-
EfficientSCI: Densely Connected Network with Space-time Factorization for Large-scale Video Snapshot Compressive Imaging 17 May 2023 · 1 repository · arXiv:2305.10006Syntology official (archive's flag): 4 ran · 4 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; every one of the 4 samples that ran constructed an object rather than computing a result (of 5 harvested samples) · 5 pointer-only (licence)
-
Explaining black box text modules in natural language with language models 17 May 2023 · 2 repositories · arXiv:2305.09863Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
From chocolate bunny to chocolate crocodile: Do Language Models Understand Noun Compounds? 17 May 2023 · 0 repositories · arXiv:2305.10568
-
G-Adapter: Towards Structure-Aware Parameter-Efficient Transfer Learning for Graph Transformer Networks 17 May 2023 · 0 repositories · arXiv:2305.10329
-
Improving Speaker Verification with Self-Pretrained Transformer Models 17 May 2023 · 0 repositories · arXiv:2305.10517
-
Instruction Tuned Models are Quick Learners 17 May 2023 · 1 repository · arXiv:2306.05539Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Interactive Learning of Hierarchical Tasks from Dialog with GPT 17 May 2023 · 0 repositories · arXiv:2305.10349
-
Knowledge Graph Completion Models are Few-shot Learners: An Empirical Study of Relation Labeling in E-commerce with LLMs 17 May 2023 · 0 repositories · arXiv:2305.09858
-
Large-Scale Text Analysis Using Generative Language Models: A Case Study in Discovering Public Value Expressions in AI Patents 17 May 2023 · 0 repositories · arXiv:2305.10383
-
M3KE: A Massive Multi-Level Multi-Subject Knowledge Evaluation Benchmark for Chinese Large Language Models 17 May 2023 · 1 repository · arXiv:2305.10263
-
Large Language Models Leverage External Knowledge to Extend Clinical Insight Beyond Language Boundaries 17 May 2023 · 0 repositories · arXiv:2305.10163
-
Rethinking Data Augmentation for Tabular Data in Deep Learning 17 May 2023 · 1 repository · arXiv:2305.10308Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
SAM for Poultry Science 17 May 2023 · 0 repositories · arXiv:2305.10254
-
Short-Term Electricity Load Forecasting Using the Temporal Fusion Transformer: Effect of Grid Hierarchies and Data Sources 17 May 2023 · 0 repositories · arXiv:2305.10559
-
Smaller Language Models are Better Black-box Machine-Generated Text Detectors 17 May 2023 · 1 repository · arXiv:2305.09859
-
Solving Cosine Similarity Underestimation between High Frequency Words by L2 Norm Discounting 17 May 2023 · 1 repository · arXiv:2305.10610
-
Token-wise Decomposition of Autoregressive Language Model Hidden States for Analyzing Model Predictions 17 May 2023 · 1 repository · arXiv:2305.10614Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 12 harvested samples) · 7 pointer-only (licence)
-
Tree of Thoughts: Deliberate Problem Solving with Large Language Models 17 May 2023 · 6 repositories · arXiv:2305.10601Syntology official (archive's flag): 2 ran · 19 ran (of which 6 constructed an object rather than computing a result; 18 with no instrument failure: 0 honoured, 0 violated, 18 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (of 24 harvested samples) · 1 pointer-only (licence)
-
When Gradient Descent Meets Derivative-Free Optimization: A Match Made in Black-Box Scenario 17 May 2023 · 0 repositories · arXiv:2305.10013
-
A Preliminary Analysis on the Code Generation Capabilities of GPT-3.5 and Bard AI Models for Java Functions 16 May 2023 · 0 repositories · arXiv:2305.09402
-
CWTM: Leveraging Contextualized Word Embeddings from BERT for Neural Topic Modeling 16 May 2023 · 1 repository · arXiv:2305.09329
-
Blind Image Quality Assessment via Transformer Predicted Error Map and Perceptual Quality Token 16 May 2023 · 1 repository · arXiv:2305.09353
-
CB-HVTNet: A channel-boosted hybrid vision transformer network for lymphocyte assessment in histopathological images 16 May 2023 · 0 repositories · arXiv:2305.09211
-
Cooperation Is All You Need 16 May 2023 · 0 repositories · arXiv:2305.10449
-
Exploring the Impact of Layer Normalization for Zero-shot Neural Machine Translation 16 May 2023 · 0 repositories · arXiv:2305.09312
-
Generative Table Pre-training Empowers Models for Tabular Prediction 16 May 2023 · 1 repository · arXiv:2305.09696
-
GIFT: Graph-Induced Fine-Tuning for Multi-Party Conversation Understanding 16 May 2023 · 1 repository · arXiv:2305.09360
-
Life of PII -- A PII Obfuscation Transformer 16 May 2023 · 0 repositories · arXiv:2305.09550
-
Measuring Dimensions of Self-Presentation in Twitter Bios and their Links to Misinformation Sharing 16 May 2023 · 1 repository · arXiv:2305.09548
-
PanelNet: Understanding 360 Indoor Environment via Panel Representation 16 May 2023 · 0 repositories · arXiv:2305.09078
-
Weight-Inherited Distillation for Task-Agnostic BERT Compression 16 May 2023 · 1 repository · arXiv:2305.09098
-
AdamR at SemEval-2023 Task 10: Solving the Class Imbalance Problem in Sexism Detection with Ensemble Learning 15 May 2023 · 0 repositories · arXiv:2305.08636
-
AutoRecon: Automated 3D Object Discovery and Reconstruction 15 May 2023 · 0 repositories · arXiv:2305.08810
-
C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models 15 May 2023 · 1 repository · arXiv:2305.08322Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Continual Multimodal Knowledge Graph Construction 15 May 2023 · 1 repository · arXiv:2305.08698
-
Coreference-aware Double-channel Attention Network for Multi-party Dialogue Reading Comprehension 15 May 2023 · 1 repository · arXiv:2305.08348
-
Document Understanding Dataset and Evaluation (DUDE) 15 May 2023 · 1 repository · arXiv:2305.08455Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
GeoMAE: Masked Geometric Target Prediction for Self-supervised Point Cloud Pre-Training 15 May 2023 · 1 repository · arXiv:2305.08808
-
Keras GPT Copilot: Integrating the Power of Large Language Models in Deep Learning Model Development 15 May 2023 · 1 repository