Methods › General › Attention Mechanisms › Attention › Papers, page 247
Attention Is All You Need
Attention
Papers archive 2025-07-28
archive papers tagged: 31,583 · with a code link: 13,473 · where Syntology ran a sample: 3,998 (3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,998 of 31,583 tagged: 3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument)
Page 247 of 316: papers 24,601 to 24,700 of 31,583, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Mixture-of-Experts with Expert Choice Routing 18 Feb 2022 · 0 repositories · arXiv:2202.09368
-
Task Specific Attention is one more thing you need for object detection 18 Feb 2022 · 1 repository · arXiv:2202.09048
-
Unleashing the Power of Transformer for Graphs 18 Feb 2022 · 0 repositories · arXiv:2202.10581
-
Multi-Scale Hybrid Vision Transformer for Learning Gastric Histology: AI-Based Decision Support System for Gastric Cancer Treatment 17 Feb 2022 · 0 repositories · arXiv:2202.08510
-
ST-MoE: Designing Stable and Transferable Sparse Expert Models 17 Feb 2022 · 3 repositories · arXiv:2202.08906Syntology community repositories only · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 3 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Graph Masked Autoencoders with Transformers 17 Feb 2022 · 1 repository · arXiv:2202.08391
-
Improving English to Sinhala Neural Machine Translation using Part-of-Speech Tag 17 Feb 2022 · 0 repositories · arXiv:2202.08882
-
Revisiting Over-smoothing in BERT from the Perspective of Graph 17 Feb 2022 · 0 repositories · arXiv:2202.08625
-
SGPT: GPT Sentence Embeddings for Semantic Search 17 Feb 2022 · 1 repository · arXiv:2202.08904Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Transformer for Graphs: An Overview from Architecture Perspective 17 Feb 2022 · 1 repository · arXiv:2202.08455Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
TraSeTR: Track-to-Segment Transformer with Contrastive Query for Instance-level Instrument Segmentation in Robotic Surgery 17 Feb 2022 · 0 repositories · arXiv:2202.08453
-
When BERT Meets Quantum Temporal Convolution Learning for Text Classification in Heterogeneous Computing 17 Feb 2022 · 0 repositories · arXiv:2203.03550
-
A Survey of Pretraining on Graphs: Taxonomy, Methods, and Applications 16 Feb 2022 · 3 repositories · arXiv:2202.07893
-
ActionFormer: Localizing Moments of Actions with Transformers 16 Feb 2022 · 1 repository · arXiv:2202.07925
-
EdgeFormer: A Parameter-Efficient Transformer for On-Device Seq2seq Generation 16 Feb 2022 · 1 repository · arXiv:2202.07959Syntology official: no sample here; runs from other or unrecorded repositories · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
No One Left Behind: Inclusive Federated Learning over Heterogeneous Devices 16 Feb 2022 · 0 repositories · arXiv:2202.08036
-
Probing Pretrained Models of Source Code 16 Feb 2022 · 1 repository · arXiv:2202.08975
-
The NLP Task Effectiveness of Long-Range Transformers 16 Feb 2022 · 0 repositories · arXiv:2202.07856
-
A Survey on Dynamic Neural Networks for Natural Language Processing 15 Feb 2022 · 0 repositories · arXiv:2202.07101
-
A Survey on Model Compression and Acceleration for Pretrained Language Models 15 Feb 2022 · 0 repositories · arXiv:2202.07105
-
BLUE at Memotion 2.0 2022: You have my Image, my Text and my Transformer 15 Feb 2022 · 0 repositories · arXiv:2202.07543
-
Defending against Reconstruction Attacks with Rényi Differential Privacy 15 Feb 2022 · 0 repositories · arXiv:2202.07623
-
One Configuration to Rule Them All? Towards Hyperparameter Transfer in Topic Models using Multi-Objective Bayesian Optimization 15 Feb 2022 · 1 repository · arXiv:2202.07631
-
Personalized Prompt Learning for Explainable Recommendation 15 Feb 2022 · 1 repository · arXiv:2202.07371
-
Predicting on the Edge: Identifying Where a Larger Model Does Better 15 Feb 2022 · 0 repositories · arXiv:2202.07652
-
Tomayto, Tomahto. Beyond Token-level Answer Equivalence for Question Answering Evaluation 15 Feb 2022 · 1 repository · arXiv:2202.07654
-
Toxic Comments Hunter : Score Severity of Toxic Comments 15 Feb 2022 · 0 repositories · arXiv:2203.03548
-
Transformers in Time Series: A Survey 15 Feb 2022 · 11 repositories · arXiv:2202.07125
-
ViNTER: Image Narrative Generation with Emotion-Arc-Aware Transformer 15 Feb 2022 · 0 repositories · arXiv:2202.07305
-
XAI for Transformers: Better Explanations through Conservative Propagation 15 Feb 2022 · 1 repository · arXiv:2202.07304Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 12 harvested samples)
-
CodeFill: Multi-token Code Completion by Jointly Learning from Structure and Naming Sequences 14 Feb 2022 · 1 repository · arXiv:2202.06689
-
Geometric Transformer for Fast and Robust Point Cloud Registration 14 Feb 2022 · 2 repositories · arXiv:2202.06688Syntology community repositories only · 4 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified; every one of the 4 samples that ran constructed an object rather than computing a result (of 9 harvested samples) · 9 pointer-only (licence)
-
Handcrafted Histological Transformer (H2T): Unsupervised Representation of Whole Slide Images 14 Feb 2022 · 1 repository · arXiv:2202.07001Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Mixing and Shifting: Exploiting Global and Local Dependencies in Vision MLPs 14 Feb 2022 · 2 repositories · arXiv:2202.06510Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 1 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples) · 6 pointer-only (licence)
-
Punctuation restoration in Swedish through fine-tuned KB-BERT 14 Feb 2022 · 0 repositories · arXiv:2202.06769
-
QA4QG: Using Question Answering to Constrain Multi-Hop Question Generation 14 Feb 2022 · 1 repository · arXiv:2202.06538
-
Sequence-to-Sequence Resources for Catalan 14 Feb 2022 · 1 repository · arXiv:2202.06871
-
Source Code Summarization with Structural Relative Position Guided Transformer 14 Feb 2022 · 1 repository · arXiv:2202.06521
-
Transformer Memory as a Differentiable Search Index 14 Feb 2022 · 1 repository · arXiv:2202.06991
-
UserBERT: Modeling Long- and Short-Term User Preferences via Self-Supervision 14 Feb 2022 · 0 repositories · arXiv:2202.07605
-
What Do They Capture? -- A Structural Analysis of Pre-Trained Language Models for Source Code 14 Feb 2022 · 1 repository · arXiv:2202.06840
-
Assessment of contextualised representations in detecting outcome phrases in clinical trials 13 Feb 2022 · 0 repositories · arXiv:2203.03547
-
BViT: Broad Attention based Vision Transformer 13 Feb 2022 · 1 repository · arXiv:2202.06268
-
ET-BERT: A Contextualized Datagram Representation with Pre-training Transformers for Encrypted Traffic Classification 13 Feb 2022 · 1 repository · arXiv:2202.06335
-
LighTN: Light-weight Transformer Network for Performance-overhead Tradeoff in Point Cloud Downsampling 13 Feb 2022 · 0 repositories · arXiv:2202.06263
-
LMN at SemEval-2022 Task 11: A Transformer-based System for English Named Entity Recognition 13 Feb 2022 · 0 repositories · arXiv:2203.03546
-
A multi-task semi-supervised framework for Text2Graph & Graph2Text 12 Feb 2022 · 1 repository · arXiv:2202.06041
-
Automatic Issue Classifier: A Transfer Learning Framework for Classifying Issue Reports 12 Feb 2022 · 1 repository · arXiv:2202.06149
-
Benchmark Assessment for DeepSpeed Optimization Library 12 Feb 2022 · 0 repositories · arXiv:2202.12831
-
Maximizing Communication Efficiency for Large-scale Training via 0/1 Adam 12 Feb 2022 · 1 repository · arXiv:2202.06009
-
Multi-direction and Multi-scale Pyramid in Transformer for Video-based Pedestrian Retrieval 12 Feb 2022 · 1 repository · arXiv:2202.06014
-
Multi-direction and Multi-scale Pyramid in Transformer for Video-based Pedestrian Retrieval 12 Feb 2022 · 1 repository
-
HaT5: Hate Language Identification using Text-to-Text Transfer Transformer 11 Feb 2022 · 0 repositories · arXiv:2202.05690
-
Including Facial Expressions in Contextual Embeddings for Sign Language Generation 11 Feb 2022 · 0 repositories · arXiv:2202.05383
-
White-Box Attacks on Hate-speech BERT Classifiers in German with Explicit and Implicit Character Level Defense 11 Feb 2022 · 1 repository · arXiv:2202.05778
-
A Multi-task Learning Framework for Product Ranking with BERT 10 Feb 2022 · 0 repositories · arXiv:2202.05317
-
AA-TransUNet: Attention Augmented TransUNet For Nowcasting Tasks 10 Feb 2022 · 1 repository · arXiv:2202.04996
-
Slovene SuperGLUE Benchmark: Translation and Evaluation 10 Feb 2022 · 0 repositories · arXiv:2202.04994
-
Can Open Domain Question Answering Systems Answer Visual Knowledge Questions? 9 Feb 2022 · 0 repositories · arXiv:2202.04306
-
Social Media as an Instant Source of Feedback on Water Quality 9 Feb 2022 · 0 repositories · arXiv:2202.04462
-
pNLP-Mixer: an Efficient all-MLP Architecture for Language 9 Feb 2022 · 1 repository · arXiv:2202.04350
-
The Volcspeech system for the ICASSP 2022 multi-channel multi-party meeting transcription challenge 9 Feb 2022 · 0 repositories · arXiv:2202.04261
-
CALM: Contrastive Aligned Audio-Language Multirate and Multimodal Representations 8 Feb 2022 · 0 repositories · arXiv:2202.03587
-
Do Language Models Learn Position-Role Mappings? 8 Feb 2022 · 0 repositories · arXiv:2202.03611
-
Efficacy of Transformer Networks for Classification of Raw EEG Data 8 Feb 2022 · 0 repositories · arXiv:2202.05170
-
HistBERT: A Pre-trained Language Model for Diachronic Lexical Semantic Analysis 8 Feb 2022 · 1 repository · arXiv:2202.03612
-
Logical Reasoning for Task Oriented Dialogue Systems 8 Feb 2022 · 0 repositories · arXiv:2202.04161
-
Particle Transformer for Jet Tagging 8 Feb 2022 · 1 repository · arXiv:2202.03772Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Semantic features of object concepts generated with GPT-3 8 Feb 2022 · 1 repository · arXiv:2202.03753
-
What are the best systems? New perspectives on NLP Benchmarking 8 Feb 2022 · 1 repository · arXiv:2202.03799
-
Cedille: A large autoregressive French language model 7 Feb 2022 · 1 repository · arXiv:2202.03371
-
data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language 7 Feb 2022 · 12 repositories · arXiv:2202.03555Syntology community repositories only · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples)
-
Efficient Adapter Transfer of Self-Supervised Speech Models for Automatic Speech Recognition 7 Feb 2022 · 1 repository · arXiv:2202.03218Syntology 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 4 harvested samples)
-
Recent Trends in 2D Object Detection and Applications in Video Event Recognition 7 Feb 2022 · 0 repositories · arXiv:2202.03206
-
Structure-Aware Transformer for Graph Representation Learning 7 Feb 2022 · 3 repositories · arXiv:2202.03036Syntology official (archive's flag): 2 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 3 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples) · 1 pointer-only (licence)
-
OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework 7 Feb 2022 · 4 repositories · arXiv:2202.03052Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
User Satisfaction Estimation with Sequential Dialogue Act Modeling in Goal-oriented Conversational Systems 7 Feb 2022 · 1 repository · arXiv:2202.02912
-
Evaluating natural language processing models with generalization metrics that do not need access to any training or testing data 6 Feb 2022 · 1 repository · arXiv:2202.02842
-
No Parameters Left Behind: Sensitivity Guided Adaptive Learning Rate for Training Large Transformer Models 6 Feb 2022 · 1 repository · arXiv:2202.02664Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Exploring Self-Attention Mechanisms for Speech Separation 6 Feb 2022 · 1 repository · arXiv:2202.02884
-
Classification on Sentence Embeddings for Legal Assistance 5 Feb 2022 · 0 repositories · arXiv:2202.02639
-
Ethics, Rules of Engagement, and AI: Neural Narrative Mapping Using Large Transformer Language Models 5 Feb 2022 · 0 repositories · arXiv:2202.02647
-
A Benchmark Corpus for the Detection of Automatically Generated Text in Academic Publications 4 Feb 2022 · 1 repository · arXiv:2202.02013
-
Image-to-Image MLP-mixer for Image Reconstruction 4 Feb 2022 · 1 repository · arXiv:2202.02018
-
Pre-Trained Neural Language Models for Automatic Mobile App User Feedback Answer Generation 4 Feb 2022 · 0 repositories · arXiv:2202.02294
-
StonkBERT: Can Language Models Predict Medium-Run Stock Price Movements? 4 Feb 2022 · 0 repositories · arXiv:2202.02268
-
Supervised Contrastive Learning for Product Matching 4 Feb 2022 · 1 repository · arXiv:2202.02098
-
Temporal Attention for Language Models 4 Feb 2022 · 1 repository · arXiv:2202.02093
-
The 6-Ds of Creating AI-Enabled Systems 4 Feb 2022 · 0 repositories · arXiv:2202.03172
-
TransFollower: Long-Sequence Car-Following Trajectory Prediction through Transformer 4 Feb 2022 · 0 repositories · arXiv:2202.03183
-
Towards 3D Scene Reconstruction from Locally Scale-Aligned Monocular Video Depth 3 Feb 2022 · 0 repositories · arXiv:2202.01470
-
Brain Cancer Survival Prediction on Treatment-na ive MRI using Deep Anchor Attention Learning with Vision Transformer 3 Feb 2022 · 0 repositories · arXiv:2202.01857
-
ETSformer: Exponential Smoothing Transformers for Time-series Forecasting 3 Feb 2022 · 3 repositories · arXiv:2202.01381Syntology official: no sample here; runs from other or unrecorded repositories · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 2 pointer-only (licence)
-
ASR-Aware End-to-end Neural Diarization 2 Feb 2022 · 0 repositories · arXiv:2202.01286
-
Exploring Transformer Backbones for Heterogeneous Treatment Effect Estimation 2 Feb 2022 · 1 repository · arXiv:2202.01336Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Co-training Improves Prompt-based Learning for Large Language Models 2 Feb 2022 · 1 repository · arXiv:2202.00828Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 10 harvested samples)
-
Error Correction in ASR using Sequence-to-Sequence Models 2 Feb 2022 · 0 repositories · arXiv:2202.01157
-
GatorTron: A Large Clinical Language Model to Unlock Patient Information from Unstructured Electronic Health Records 2 Feb 2022 · 0 repositories · arXiv:2203.03540
-
L3Cube-MahaCorpus and MahaBERT: Marathi Monolingual Corpus, Marathi BERT Language Models, and Resources 2 Feb 2022 · 1 repository · arXiv:2202.01159
-
RescoreBERT: Discriminative Speech Recognition Rescoring with BERT 2 Feb 2022 · 0 repositories · arXiv:2202.01094