Methods › General › Attention Mechanisms › Attention › Papers, page 237
Attention Is All You Need
Attention
Papers archive 2025-07-28
archive papers tagged: 31,583 · with a code link: 13,473 · where Syntology ran a sample: 3,998 (3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,998 of 31,583 tagged: 3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument)
Page 237 of 316: papers 23,601 to 23,700 of 31,583, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Modeling Image Composition for Complex Scene Generation 2 Jun 2022 · 1 repository · arXiv:2206.00923Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Optimizing Relevance Maps of Vision Transformers Improves Robustness 2 Jun 2022 · 1 repository · arXiv:2206.01161
-
Transforming medical imaging with Transformers? A comparative review of key properties, current progresses, and future perspectives 2 Jun 2022 · 0 repositories · arXiv:2206.01136
-
VL-BEiT: Generative Vision-Language Pretraining 2 Jun 2022 · 0 repositories · arXiv:2206.01127
-
A comparative study between vision transformers and CNNs in digital pathology 1 Jun 2022 · 0 repositories · arXiv:2206.00389
-
Assessing Group-level Gender Bias in Professional Evaluations: The Case of Medical Student End-of-Shift Feedback 1 Jun 2022 · 0 repositories · arXiv:2206.00234
-
BERT-Sort: A Zero-shot MLM Semantic Encoder on Ordinal Features for AutoML 1 Jun 2022 · 1 repository
-
CLIP4IDC: CLIP for Image Difference Captioning 1 Jun 2022 · 1 repository · arXiv:2206.00629
-
Cross-domain Detection Transformer based on Spatial-aware and Semantic-aware Token Alignment 1 Jun 2022 · 0 repositories · arXiv:2206.00222
-
Efficient Self-supervised Vision Pretraining with Local Masked Reconstruction 1 Jun 2022 · 1 repository · arXiv:2206.00790
-
Entity Resolution with Hierarchical Graph Attention Networks 1 Jun 2022 · 1 repository
-
Floorplan Restoration by Structure Hallucinating Transformer Cascades 1 Jun 2022 · 0 repositories · arXiv:2206.00645
-
Learning Sequential Contexts using Transformer for 3D Hand Pose Estimation 1 Jun 2022 · 0 repositories · arXiv:2206.00171
-
Order-sensitive Shapley Values for Evaluating Conceptual Soundness of NLP Models 1 Jun 2022 · 0 repositories · arXiv:2206.00192
-
Romantic-Computing 1 Jun 2022 · 0 repositories · arXiv:2206.11864
-
Stargazer: A transformer-based driver action detection system for intelligent transportation 1 Jun 2022 · 1 repository
-
The Fully Convolutional Transformer for Medical Image Segmentation 1 Jun 2022 · 2 repositories · arXiv:2206.00566
-
A Unified Framework for Emotion Identification and Generation in Dialogues 31 May 2022 · 0 repositories · arXiv:2205.15513
-
CodeAttack: Code-Based Adversarial Attacks for Pre-trained Programming Language Models 31 May 2022 · 2 repositories · arXiv:2206.00052
-
Decomposing NeRF for Editing via Feature Field Distillation 31 May 2022 · 1 repository · arXiv:2205.15585Syntology 10 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 1 honoured, 0 violated, 8 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples) · 1 pointer-only (licence)
-
GateNLP-UShef at SemEval-2022 Task 8: Entity-Enriched Siamese Transformer for Multilingual News Article Similarity 31 May 2022 · 1 repository · arXiv:2205.15812
-
Inferring 3D change detection from bitemporal optical images 31 May 2022 · 0 repositories · arXiv:2205.15903
-
Joint Spatial-Temporal and Appearance Modeling with Transformer for Multiple Object Tracking 31 May 2022 · 1 repository · arXiv:2205.15495
-
Knowledge Graph - Deep Learning: A Case Study in Question Answering in Aviation Safety Domain 31 May 2022 · 0 repositories · arXiv:2205.15952
-
HierarchyNet: Learning to Summarize Source Code with Heterogeneous Representations 31 May 2022 · 0 repositories · arXiv:2205.15479
-
Multilingual Transformers for Product Matching -- Experiments and a New Benchmark in Polish 31 May 2022 · 0 repositories · arXiv:2205.15712
-
Surface Analysis with Vision Transformers 31 May 2022 · 1 repository · arXiv:2205.15836Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
You Can't Count on Luck: Why Decision Transformers and RvS Fail in Stochastic Environments 31 May 2022 · 0 repositories · arXiv:2205.15967
-
Attention Flows for General Transformers 30 May 2022 · 1 repository · arXiv:2205.15389
-
Automatic Short Math Answer Grading via In-context Meta-learning 30 May 2022 · 1 repository · arXiv:2205.15219
-
Billions of Parameters Are Worth More Than In-domain Training Data: A case study in the Legal Case Entailment Task 30 May 2022 · 1 repository · arXiv:2205.15172
-
Can Transformer be Too Compositional? Analysing Idiom Processing in Neural Machine Translation 30 May 2022 · 1 repository · arXiv:2205.15301
-
E2S2: Encoding-Enhanced Sequence-to-Sequence Pretraining for Language Understanding and Generation 30 May 2022 · 1 repository · arXiv:2205.14912
-
Few-Shot Diffusion Models 30 May 2022 · 1 repository · arXiv:2205.15463
-
HiViT: Hierarchical Vision Transformer Meets Masked Image Modeling 30 May 2022 · 1 repository · arXiv:2205.14949
-
You Only Need 90K Parameters to Adapt Light: A Light Weight Transformer for Image Enhancement and Exposure Correction 30 May 2022 · 1 repository · arXiv:2205.14871Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Multi-Agent Reinforcement Learning is a Sequence Modeling Problem 30 May 2022 · 1 repository · arXiv:2205.14953
-
Multi-Game Decision Transformers 30 May 2022 · 1 repository · arXiv:2205.15241
-
Prompting ELECTRA: Few-Shot Learning with Discriminative Pre-Trained Models 30 May 2022 · 1 repository · arXiv:2205.15223
-
Temporal Latent Bottleneck: Synthesis of Fast and Slow Processing Mechanisms in Sequence Learning 30 May 2022 · 2 repositories · arXiv:2205.14794Syntology 0 ran · 1 unverified (of 1 harvested sample)
-
Transformer with Tree-order Encoding for Neural Program Generation 30 May 2022 · 1 repository · arXiv:2206.13354
-
Zero-Shot and Few-Shot Learning for Lung Cancer Multi-Label Classification using Vision Transformer 30 May 2022 · 0 repositories · arXiv:2205.15290
-
COVID-19 Literature Mining and Retrieval using Text Mining Approaches 29 May 2022 · 0 repositories · arXiv:2205.14781
-
CPED: A Large-Scale Chinese Personalized and Emotional Dialogue Dataset for Conversational AI 29 May 2022 · 1 repository · arXiv:2205.14727
-
EfficientViT: Multi-Scale Linear Attention for High-Resolution Dense Prediction 29 May 2022 · 6 repositories · arXiv:2205.14756Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Learning Locality and Isotropy in Dialogue Modeling 29 May 2022 · 1 repository · arXiv:2205.14583Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Micro-Expression Recognition Based on Attribute Information Embedding and Cross-modal Contrastive Learning 29 May 2022 · 0 repositories · arXiv:2205.14643
-
SFE-AI at SemEval-2022 Task 11: Low-Resource Named Entity Recognition using Large Pre-trained Language Models 29 May 2022 · 0 repositories · arXiv:2205.14660
-
TransforMAP: Transformer for Memory Access Prediction 29 May 2022 · 0 repositories · arXiv:2205.14778
-
Urdu News Article Recommendation Model using Natural Language Processing Techniques 29 May 2022 · 0 repositories · arXiv:2206.11862
-
Approximate Conditional Coverage & Calibration via Neural Model Approximations 28 May 2022 · 0 repositories · arXiv:2205.14310
-
Learning Non-Autoregressive Models from Search for Unsupervised Sentence Summarization 28 May 2022 · 1 repository · arXiv:2205.14521
-
MDMLP: Image Classification from Scratch on Small Datasets with MLP 28 May 2022 · 2 repositories · arXiv:2205.14477
-
Multi-Task Learning with Multi-Query Transformer for Dense Prediction 28 May 2022 · 1 repository · arXiv:2205.14354
-
Non-stationary Transformers: Exploring the Stationarity in Time Series Forecasting 28 May 2022 · 2 repositories · arXiv:2205.14415
-
Teaching Models to Express Their Uncertainty in Words 28 May 2022 · 1 repository · arXiv:2205.14334Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 2 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Variational Transformer: A Framework Beyond the Trade-off between Accuracy and Diversity for Image Captioning 28 May 2022 · 1 repository · arXiv:2205.14458
-
WT-MVSNet: Window-based Transformers for Multi-view Stereo 28 May 2022 · 0 repositories · arXiv:2205.14319
-
Architecture-Agnostic Masked Image Modeling -- From ViT back to CNN 27 May 2022 · 3 repositories · arXiv:2205.13943
-
FedFormer: Contextual Federation with Attention in Reinforcement Learning 27 May 2022 · 1 repository · arXiv:2205.13697Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 2 pointer-only (licence)
-
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness 27 May 2022 · 13 repositories · arXiv:2205.14135Syntology official: no sample here; runs from other or unrecorded repositories · 24 ran (of which 2 constructed an object rather than computing a result; 18 with no instrument failure: 0 honoured, 1 violated, 17 with no contract checked; 6 where Syntology's instrument failed) · 6 unverified (of 30 harvested samples) · 1 pointer-only (licence)
-
Future Transformer for Long-term Action Anticipation 27 May 2022 · 1 repository · arXiv:2205.14022Syntology 5 ran (of which 4 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 1 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
GIT: A Generative Image-to-text Transformer for Vision and Language 27 May 2022 · 1 repository · arXiv:2205.14100Syntology official (archive's flag): 14 ran · 14 ran (of which 0 constructed an object rather than computing a result; 14 with no instrument failure: 0 honoured, 0 violated, 14 with no contract checked; 0 where Syntology's instrument failed) · 7 unverified (of 21 harvested samples)
-
Momentum Stiefel Optimizer, with Applications to Suitably-Orthogonal Attention, and Optimal Transport 27 May 2022 · 1 repository · arXiv:2205.14173Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 1 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples)
-
Multimodal Masked Autoencoders Learn Transferable Representations 27 May 2022 · 3 repositories · arXiv:2205.14204Syntology official (archive's flag): 3 ran · 12 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 12 harvested samples)
-
kNN-Prompt: Nearest Neighbor Zero-Shot Inference 27 May 2022 · 1 repository · arXiv:2205.13792
-
NLU for Game-based Learning in Real: Initial Evaluations 27 May 2022 · 0 repositories · arXiv:2205.13754
-
Patching Leaks in the Charformer for Efficient Character-Level Generation 27 May 2022 · 1 repository · arXiv:2205.14086
-
Probabilistic Transformer: Modelling Ambiguities and Distributions for RNA Folding and Molecule Design 27 May 2022 · 1 repository · arXiv:2205.13927
-
Transformers from an Optimization Perspective 27 May 2022 · 1 repository · arXiv:2205.13891Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples)
-
TURJUMAN: A Public Toolkit for Neural Arabic Machine Translation 27 May 2022 · 1 repository · arXiv:2206.03933
-
Understanding Long Programming Languages with Structure-Aware Sparse Attention 27 May 2022 · 1 repository · arXiv:2205.13730
-
What Dense Graph Do You Need for Self-Attention? 27 May 2022 · 1 repository · arXiv:2205.14014
-
AdaptFormer: Adapting Vision Transformers for Scalable Visual Recognition 26 May 2022 · 2 repositories · arXiv:2205.13535Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Are Transformers Effective for Time Series Forecasting? 26 May 2022 · 10 repositories · arXiv:2205.13504Syntology official: no sample here; runs from other or unrecorded repositories · 13 ran (of which 4 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified (of 19 harvested samples) · 13 pointer-only (licence)
-
Benchmarking of Deep Learning models on 2D Laminar Flow behind Cylinder 26 May 2022 · 0 repositories · arXiv:2205.13485
-
Clinical Dialogue Transcription Error Correction using Seq2Seq Models 26 May 2022 · 0 repositories · arXiv:2205.13572
-
DT-SV: A Transformer-based Time-domain Approach for Speaker Verification 26 May 2022 · 0 repositories · arXiv:2205.13249
-
Do we really need temporal convolutions in action segmentation? 26 May 2022 · 1 repository · arXiv:2205.13425
-
Embed to Control Partially Observed Systems: Representation Learning with Provable Sample Efficiency 26 May 2022 · 0 repositories · arXiv:2205.13476
-
Federated Split BERT for Heterogeneous Text Classification 26 May 2022 · 0 repositories · arXiv:2205.13299
-
Green Hierarchical Vision Transformer for Masked Image Modeling 26 May 2022 · 1 repository · arXiv:2205.13515Syntology official (archive's flag): 4 ran · 4 ran (of which 1 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Learning Dialogue Representations from Consecutive Utterances 26 May 2022 · 1 repository · arXiv:2205.13568
-
Leveraging Dependency Grammar for Fine-Grained Offensive Language Detection using Graph Convolutional Networks 26 May 2022 · 1 repository · arXiv:2205.13164
-
MixMAE: Mixed and Masked Autoencoder for Efficient Pretraining of Hierarchical Vision Transformers 26 May 2022 · 1 repository · arXiv:2205.13137
-
SemAffiNet: Semantic-Affine Transformation for Point Cloud Segmentation 26 May 2022 · 1 repository · arXiv:2205.13490Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 1 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 11 harvested samples) · 2 pointer-only (licence)
-
The Document Vectors Using Cosine Similarity Revisited 26 May 2022 · 1 repository · arXiv:2205.13357
-
Towards Learning Universal Hyperparameter Optimizers with Transformers 26 May 2022 · 1 repository · arXiv:2205.13320
-
Training and Inference on Any-Order Autoregressive Models the Right Way 26 May 2022 · 1 repository · arXiv:2205.13554
-
Transformer for Partial Differential Equations' Operator Learning 26 May 2022 · 1 repository · arXiv:2205.13671Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 1 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 13 harvested samples)
-
VIDI: A Video Dataset of Incidents 26 May 2022 · 1 repository · arXiv:2205.13277
-
Your Transformer May Not be as Powerful as You Expect 26 May 2022 · 1 repository · arXiv:2205.13401Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
BiT: Robustly Binarized Multi-distilled Transformer 25 May 2022 · 3 repositories · arXiv:2205.13016Syntology official (archive's flag): 2 ran · 3 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Conditional set generation using Seq2seq models 25 May 2022 · 0 repositories · arXiv:2205.12485
-
Eye-gaze-guided Vision Transformer for Rectifying Shortcut Learning 25 May 2022 · 0 repositories · arXiv:2205.12466
-
Factorizing Content and Budget Decisions in Abstractive Summarization of Long Documents 25 May 2022 · 1 repository · arXiv:2205.12486
-
Inception Transformer 25 May 2022 · 4 repositories · arXiv:2205.12956Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Large Language Models are Few-Shot Clinical Information Extractors 25 May 2022 · 0 repositories · arXiv:2205.12689
-
Lifelong Learning Natural Language Processing Approach for Multilingual Data Classification 25 May 2022 · 0 repositories · arXiv:2206.11867
-
MoCoViT: Mobile Convolutional Vision Transformer 25 May 2022 · 1 repository · arXiv:2205.12635