Methods › General › Attention Modules › Multi-Head Attention › Papers, page 129
Multi-Head Attention
Papers archive 2025-07-28
archive papers tagged: 24,855 · with a code link: 11,214 · where Syntology ran a sample: 3,454 (2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,454 of 24,855 tagged: 2,916 with a run with no instrument failure, 538 where every run was a failure of Syntology's instrument)
Page 129 of 249: papers 12,801 to 12,900 of 24,855, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Thrilled by Your Progress! Large Language Models (GPT-4) No Longer Struggle to Pass Assessments in Higher Education Programming Courses 15 Jun 2023 · 0 repositories · arXiv:2306.10073
-
Ensembled Prediction Intervals for Causal Outcomes Under Hidden Confounding 15 Jun 2023 · 0 repositories · arXiv:2306.09520
-
DiffAug: A Diffuse-and-Denoise Augmentation for Training Robust Classifiers 15 Jun 2023 · 0 repositories · arXiv:2306.09192
-
ViP: A Differentially Private Foundation Model for Computer Vision 15 Jun 2023 · 1 repository · arXiv:2306.08842Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
A semantically enhanced dual encoder for aspect sentiment triplet extraction 14 Jun 2023 · 1 repository · arXiv:2306.08373
-
Assessing the Effectiveness of GPT-3 in Detecting False Political Statements: A Case Study on the LIAR Dataset 14 Jun 2023 · 1 repository · arXiv:2306.08190
-
Building a Corpus for Biomedical Relation Extraction of Species Mentions 14 Jun 2023 · 0 repositories · arXiv:2306.08403
-
Language models are not naysayers: An analysis of language models on negation benchmarks 14 Jun 2023 · 1 repository · arXiv:2306.08189
-
M^2UNet: MetaFormer Multi-scale Upsampling Network for Polyp Segmentation 14 Jun 2023 · 0 repositories · arXiv:2306.08600
-
MCR-Data2vec 2.0: Improving Self-supervised Speech Pre-training via Model-level Consistency Regularization 14 Jun 2023 · 0 repositories · arXiv:2306.08463
-
MUBen: Benchmarking the Uncertainty of Molecular Representation Models 14 Jun 2023 · 2 repositories · arXiv:2306.10060Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Multimodal Optimal Transport-based Co-Attention Transformer with Global Structure Consistency for Survival Prediction 14 Jun 2023 · 3 repositories · arXiv:2306.08330Syntology official: harvested, nothing ran · 10 ran (of which 6 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 3 where Syntology's instrument failed) · 8 unverified (of 18 harvested samples) · 18 pointer-only (licence)
-
Research on Named Entity Recognition in Improved transformer with R-Drop structure 14 Jun 2023 · 0 repositories · arXiv:2306.08315
-
The Expressive Leaky Memory Neuron: an Efficient and Expressive Phenomenological Neuron Model Can Solve Long-Horizon Tasks 14 Jun 2023 · 1 repository · arXiv:2306.16922Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 2 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
Towards AGI in Computer Vision: Lessons Learned from GPT and Large Language Models 14 Jun 2023 · 0 repositories · arXiv:2306.08641
-
TSMixer: Lightweight MLP-Mixer Model for Multivariate Time Series Forecasting 14 Jun 2023 · 1 repository · arXiv:2306.09364Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples)
-
Unraveling the ARC Puzzle: Mimicking Human Solutions with Object-Centric Decision Transformer 14 Jun 2023 · 0 repositories · arXiv:2306.08204
-
When to Use Efficient Self Attention? Profiling Text, Speech and Image Transformer Variants 14 Jun 2023 · 1 repository · arXiv:2306.08667
-
World-to-Words: Grounded Open Vocabulary Acquisition through Fast Mapping in Vision-Language Models 14 Jun 2023 · 1 repository · arXiv:2306.08685Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples)
-
arXiVeri: Automatic table verification with GPT 13 Jun 2023 · 1 repository · arXiv:2306.07968
-
Can ChatGPT Enable ITS? The Case of Mixed Traffic Control via Reinforcement Learning 13 Jun 2023 · 1 repository · arXiv:2306.08094
-
Enhancing Social Network Hate Detection Using Back Translation and GPT-3 Augmentations During Training and Test-Time 13 Jun 2023 · 1 repository
-
FLamE: Few-shot Learning from Natural Language Explanations 13 Jun 2023 · 0 repositories · arXiv:2306.08042
-
GEmo-CLAP: Gender-Attribute-Enhanced Contrastive Language-Audio Pretraining for Accurate Speech Emotion Recognition 13 Jun 2023 · 0 repositories · arXiv:2306.07848
-
h2oGPT: Democratizing Large Language Models 13 Jun 2023 · 2 repositories · arXiv:2306.08161
-
Human-Like Intuitive Behavior and Reasoning Biases Emerged in Language Models -- and Disappeared in GPT-4 13 Jun 2023 · 0 repositories · arXiv:2306.07622
-
Improving Zero-Shot Detection of Low Prevalence Chest Pathologies using Domain Pre-trained Language Models 13 Jun 2023 · 1 repository · arXiv:2306.08000
-
MolCAP: Molecular Chemical reActivity pretraining and prompted-finetuning enhanced molecular representation learning 13 Jun 2023 · 0 repositories · arXiv:2306.09187
-
Monolingual and Cross-Lingual Knowledge Transfer for Topic Classification 13 Jun 2023 · 0 repositories · arXiv:2306.07797
-
Reviving Shift Equivariance in Vision Transformers 13 Jun 2023 · 0 repositories · arXiv:2306.07470
-
Semi-supervised learning made simple with self-supervised clustering 13 Jun 2023 · 1 repository · arXiv:2306.07483Syntology official (archive's flag): 2 ran · 2 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Discrete Graph Auto-Encoder 13 Jun 2023 · 0 repositories · arXiv:2306.07735
-
XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models 13 Jun 2023 · 1 repository · arXiv:2306.07971
-
A Survey of Vision-Language Pre-training from the Lens of Multimodal Machine Translation 12 Jun 2023 · 0 repositories · arXiv:2306.07198
-
AerialFormer: Multi-resolution Transformer for Aerial Image Segmentation 12 Jun 2023 · 1 repository · arXiv:2306.06842
-
CD-CTFM: A Lightweight CNN-Transformer Network for Remote Sensing Cloud Detection Fusing Multiscale Features 12 Jun 2023 · 0 repositories · arXiv:2306.07186
-
Enhancing COVID-19 Diagnosis through Vision Transformer-Based Analysis of Chest X-ray Images 12 Jun 2023 · 0 repositories · arXiv:2306.06914
-
Exploring Attention Mechanisms for Multimodal Emotion Recognition in an Emergency Call Center Corpus 12 Jun 2023 · 0 repositories · arXiv:2306.07115
-
Imbalanced Multi-label Classification for Business-related Text with Moderately Large Label Spaces 12 Jun 2023 · 0 repositories · arXiv:2306.07046
-
Large language models and (non-)linguistic recursion 12 Jun 2023 · 0 repositories · arXiv:2306.07195
-
Large Language Models as Tax Attorneys: A Case Study in Legal Capabilities Emergence 12 Jun 2023 · 0 repositories · arXiv:2306.07075
-
Learning Multilingual Sentence Representations with Cross-lingual Consistency Regularization 12 Jun 2023 · 1 repository · arXiv:2306.06919
-
Learning to Mask and Permute Visual Tokens for Vision Transformer Pre-Training 12 Jun 2023 · 1 repository · arXiv:2306.07346
-
Leveraging Skill-to-Skill Supervision for Knowledge Tracing 12 Jun 2023 · 0 repositories · arXiv:2306.06841
-
Linear Classifier: An Often-Forgotten Baseline for Text Classification 12 Jun 2023 · 1 repository · arXiv:2306.07111Syntology official: no sample here; runs from other or unrecorded repositories · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples)
-
Lost in Translation: Large Language Models in Non-English Content Analysis 12 Jun 2023 · 0 repositories · arXiv:2306.07377
-
MaskedFusion360: Reconstruct LiDAR Data by Querying Camera Features 12 Jun 2023 · 1 repository · arXiv:2306.07087
-
Mitigating Transformer Overconfidence via Lipschitz Regularization 12 Jun 2023 · 1 repository · arXiv:2306.06849
-
Multimodal Audio-textual Architecture for Robust Spoken Language Understanding 12 Jun 2023 · 0 repositories · arXiv:2306.06819
-
A Graph Transformer-Driven Approach for Network Robustness Learning 12 Jun 2023 · 0 repositories · arXiv:2306.06913
-
NPVForensics: Jointing Non-critical Phonemes and Visemes for Deepfake Detection 12 Jun 2023 · 0 repositories · arXiv:2306.06885
-
On the N-gram Approximation of Pre-trained Language Models 12 Jun 2023 · 0 repositories · arXiv:2306.06892
-
Prompt-based Extraction of Social Determinants of Health Using Few-shot Learning 12 Jun 2023 · 0 repositories · arXiv:2306.07170
-
Recursion of Thought: A Divide-and-Conquer Approach to Multi-Context Reasoning with Language Models 12 Jun 2023 · 1 repository · arXiv:2306.06891
-
The BEA 2023 Shared Task on Generating AI Teacher Responses in Educational Dialogues 12 Jun 2023 · 0 repositories · arXiv:2306.06941
-
UniPoll: A Unified Social Media Poll Generation Framework via Multi-Objective Optimization 12 Jun 2023 · 1 repository · arXiv:2306.06851
-
Waffling around for Performance: Visual Classification with Random Words and Broad Concepts 12 Jun 2023 · 2 repositories · arXiv:2306.07282Syntology official (archive's flag): 4 ran · 5 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
CoTran: An LLM-based Code Translator using Reinforcement Learning with Feedback from Compiler and Symbolic Execution 11 Jun 2023 · 1 repository · arXiv:2306.06755
-
E(2)-Equivariant Vision Transformer 11 Jun 2023 · 1 repository · arXiv:2306.06722Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
EaSyGuide : ESG Issue Identification Framework leveraging Abilities of Generative Large Language Models 11 Jun 2023 · 1 repository · arXiv:2306.06662
-
Inductive reasoning in humans and large language models 11 Jun 2023 · 1 repository · arXiv:2306.06548
-
Local-to-global Perspectives on Graph Neural Networks 11 Jun 2023 · 1 repository · arXiv:2306.06547
-
RoBERTweet: A BERT Language Model for Romanian Tweets 11 Jun 2023 · 0 repositories · arXiv:2306.06598
-
PVPUFormer: Probabilistic Visual Prompt Unified Transformer for Interactive Image Segmentation 11 Jun 2023 · 2 repositories · arXiv:2306.06656
-
Enhancing Low Resource NER Using Assisting Language And Transfer Learning 10 Jun 2023 · 0 repositories · arXiv:2306.06477
-
Medical Data Augmentation via ChatGPT: A Case Study on Medication Identification and Medication Event Classification 10 Jun 2023 · 0 repositories · arXiv:2306.07297
-
Multi-modal Pre-training for Medical Vision-language Understanding and Generation: An Empirical Study with A New Benchmark 10 Jun 2023 · 1 repository · arXiv:2306.06494
-
Shuffled Autoregression For Motion Interpolation 10 Jun 2023 · 0 repositories · arXiv:2306.06367
-
Vista-Morph: Unsupervised Image Registration of Visible-Thermal Facial Pairs 10 Jun 2023 · 0 repositories · arXiv:2306.06505
-
What Can an Accent Identifier Learn? Probing Phonetic and Prosodic Information in a Wav2vec2-based Accent Identification Model 10 Jun 2023 · 0 repositories · arXiv:2306.06524
-
14 Examples of How LLMs Can Transform Materials Science and Chemistry: A Reflection on a Large Language Model Hackathon 9 Jun 2023 · 2 repositories · arXiv:2306.06283Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples)
-
A Gated Attention Transformer for Multi-Person Pose Tracking 9 Jun 2023 · 0 repositories · arXiv:2306.05807
-
A Unified Generative Approach to Product Attribute-Value Identification 9 Jun 2023 · 0 repositories · arXiv:2306.05605
-
Boosting Fast and High-Quality Speech Synthesis with Linear Diffusion 9 Jun 2023 · 0 repositories · arXiv:2306.05708
-
COVER: A Heuristic Greedy Adversarial Attack on Prompt-based Learning in Language Models 9 Jun 2023 · 0 repositories · arXiv:2306.05659
-
Customizing General-Purpose Foundation Models for Medical Report Generation 9 Jun 2023 · 0 repositories · arXiv:2306.05642
-
End-to-End Neural Network Compression via ℓ₁/ℓ₂ Regularized Latency Surrogates 9 Jun 2023 · 0 repositories · arXiv:2306.05785
-
Everybody Compose: Deep Beats To Music 9 Jun 2023 · 1 repository · arXiv:2306.06284
-
Exploring the Responses of Large Language Models to Beginner Programmers' Help Requests 9 Jun 2023 · 0 repositories · arXiv:2306.05715
-
GPT-Calls: Enhancing Call Segmentation and Tagging by Generating Synthetic Conversations via Large Language Models 9 Jun 2023 · 0 repositories · arXiv:2306.07941
-
Illumination Controllable Dehazing Network based on Unsupervised Retinex Embedding 9 Jun 2023 · 1 repository · arXiv:2306.05675
-
Implementing BERT and fine-tuned RobertA to detect AI generated news by ChatGPT 9 Jun 2023 · 0 repositories · arXiv:2306.07401
-
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena 9 Jun 2023 · 11 repositories · arXiv:2306.05685Syntology community repositories only · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 12 harvested samples)
-
Language Models Can Learn Exceptions to Syntactic Rules 9 Jun 2023 · 1 repository · arXiv:2306.05969
-
Lightweight Monocular Depth Estimation via Token-Sharing Transformer 9 Jun 2023 · 0 repositories · arXiv:2306.05682
-
ModeT: Learning Deformable Image Registration via Motion Decomposition Transformer 9 Jun 2023 · 1 repository · arXiv:2306.05688
-
Morphosyntactic probing of multilingual BERT models 9 Jun 2023 · 1 repository · arXiv:2306.06205
-
PoET: A generative model of protein families as sequences-of-sequences 9 Jun 2023 · 1 repository · arXiv:2306.06156Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Prodigy: An Expeditiously Adaptive Parameter-Free Learner 9 Jun 2023 · 1 repository · arXiv:2306.06101
-
RankFormer: Listwise Learning-to-Rank Using Listwide Labels 9 Jun 2023 · 1 repository · arXiv:2306.05808
-
Reconstructing Human Expressiveness in Piano Performances with a Transformer Network 9 Jun 2023 · 1 repository · arXiv:2306.06040
-
Reliability Check: An Analysis of GPT-3's Response to Sensitive Topics and Prompt Wording 9 Jun 2023 · 2 repositories · arXiv:2306.06199
-
Understanding Telecom Language Through Large Language Models 9 Jun 2023 · 0 repositories · arXiv:2306.07933
-
Virtual Node Tuning for Few-shot Node Classification 9 Jun 2023 · 0 repositories · arXiv:2306.06063
-
Augmenting Hessians with Inter-Layer Dependencies for Mixed-Precision Post-Training Quantization 8 Jun 2023 · 0 repositories · arXiv:2306.04879
-
Bias Against 93 Stigmatized Groups in Masked Language Models and Downstream Sentiment Classification Tasks 8 Jun 2023 · 1 repository · arXiv:2306.05550Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Complexity-aware Large Scale Origin-Destination Network Generation via Diffusion Model 8 Jun 2023 · 0 repositories · arXiv:2306.04873
-
Connectional-Style-Guided Contextual Representation Learning for Brain Disease Diagnosis 8 Jun 2023 · 0 repositories · arXiv:2306.05297
-
DACCORD : un jeu de données pour la Détection Automatique d'énonCés COntRaDictoires en français 8 Jun 2023 · 2 repositories
-
Deep Learning Method for Cell-Wise Object Tracking, Velocity Estimation and Projection of Sensor Data over Time 8 Jun 2023 · 0 repositories · arXiv:2306.06126