Methods › General › Attention Mechanisms › Attention › Papers, page 276
Attention Is All You Need
Attention
Papers archive 2025-07-28
archive papers tagged: 31,583 · with a code link: 13,473 · where Syntology ran a sample: 3,998 (3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,998 of 31,583 tagged: 3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument)
Page 276 of 316: papers 27,501 to 27,600 of 31,583, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Are Multilingual BERT models robust? A Case Study on Adversarial Attacks for Multilingual Question Answering 15 Apr 2021 · 0 repositories · arXiv:2104.07646
-
BERT based Transformers lead the way in Extraction of Health Information from Social Media 15 Apr 2021 · 1 repository · arXiv:2104.07367
-
Cross-domain Speech Recognition with Unsupervised Character-level Distribution Matching 15 Apr 2021 · 1 repository · arXiv:2104.07491
-
Robust Optimization for Multilingual Translation with Imbalanced Data 15 Apr 2021 · 0 repositories · arXiv:2104.07639
-
Does BERT Pretrained on Clinical Notes Reveal Sensitive Data? 15 Apr 2021 · 4 repositories · arXiv:2104.07762Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Emotion Dynamics Modeling via BERT 15 Apr 2021 · 0 repositories · arXiv:2104.07252
-
ExplaGraphs: An Explanation Graph Generation Task for Structured Commonsense Reasoning 15 Apr 2021 · 1 repository · arXiv:2104.07644Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples)
-
How to Train BERT with an Academic Budget 15 Apr 2021 · 4 repositories · arXiv:2104.07705
-
NT5?! Training T5 to Perform Numerical Reasoning 15 Apr 2021 · 1 repository · arXiv:2104.07307
-
Points as Queries: Weakly Semi-supervised Object Detection by Points 15 Apr 2021 · 1 repository · arXiv:2104.07434Syntology 0 ran · 2 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Natural Language Understanding with Privacy-Preserving BERT 15 Apr 2021 · 0 repositories · arXiv:2104.07504
-
Rethinking Text Line Recognition Models 15 Apr 2021 · 0 repositories · arXiv:2104.07787
-
Self-supervised Video Object Segmentation by Motion Grouping 15 Apr 2021 · 0 repositories · arXiv:2104.07658
-
Shoulder Implant X-Ray Manufacturer Classification: Exploring with Vision Transformer 15 Apr 2021 · 1 repository · arXiv:2104.07667
-
SINA-BERT: A pre-trained Language Model for Analysis of Medical Texts in Persian 15 Apr 2021 · 0 repositories · arXiv:2104.07613
-
Syntax-Aware Graph-to-Graph Transformer for Semantic Role Labelling 15 Apr 2021 · 0 repositories · arXiv:2104.07704
-
Text Guide: Improving the quality of long text classification by a text selection method based on feature importance 15 Apr 2021 · 1 repository · arXiv:2104.07225
-
TorontoCL at CMCL 2021 Shared Task: RoBERTa with Multi-Stage Fine-Tuning for Eye-Tracking Prediction 15 Apr 2021 · 1 repository · arXiv:2104.07244
-
Ultra-High Dimensional Sparse Representations with Binarization for Efficient Text Retrieval 15 Apr 2021 · 0 repositories · arXiv:2104.07198
-
UIT-E10dot3 at SemEval-2021 Task 5: Toxic Spans Detection with Named Entity Recognition and Question-Answering Approaches 15 Apr 2021 · 0 repositories · arXiv:2104.07376
-
Vision Transformer using Low-level Chest X-ray Feature Corpus for COVID-19 Diagnosis and Severity Quantification 15 Apr 2021 · 0 repositories · arXiv:2104.07235
-
An Interpretability Illusion for BERT 14 Apr 2021 · 0 repositories · arXiv:2104.07143
-
An Introduction of mini-AlphaStar 14 Apr 2021 · 1 repository · arXiv:2104.06890
-
Decoupled Spatial-Temporal Transformer for Video Inpainting 14 Apr 2021 · 1 repository · arXiv:2104.06637
-
Demystifying BERT: Implications for Accelerator Design 14 Apr 2021 · 0 repositories · arXiv:2104.08335
-
Enhancing Word-Level Semantic Representation via Dependency Structure for Expressive Text-to-Speech Synthesis 14 Apr 2021 · 0 repositories · arXiv:2104.06835
-
Disentangling Representations of Text by Masking Transformers 14 Apr 2021 · 0 repositories · arXiv:2104.07155
-
Knowledge-driven Answer Generation for Conversational Search 14 Apr 2021 · 0 repositories · arXiv:2104.06892
-
NAREOR: The Narrative Reordering Problem 14 Apr 2021 · 1 repository · arXiv:2104.06669
-
Non-autoregressive sequence-to-sequence voice conversion 14 Apr 2021 · 0 repositories · arXiv:2104.06793
-
On the Robustness of Intent Classification and Slot Labeling in Goal-oriented Dialog Systems to Real-world Noise 14 Apr 2021 · 1 repository · arXiv:2104.07149
-
Sparse Attention with Linear Units 14 Apr 2021 · 3 repositories · arXiv:2104.07012Syntology official: harvested, nothing ran · 3 ran (of which 1 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples)
-
Static Embeddings as Efficient Knowledge Bases? 14 Apr 2021 · 1 repository · arXiv:2104.07094
-
TSDAE: Using Transformer-based Sequential Denoising Auto-Encoder for Unsupervised Sentence Embedding Learning 14 Apr 2021 · 6 repositories · arXiv:2104.06979Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
TWEAC: Transformer with Extendable QA Agent Classifiers 14 Apr 2021 · 1 repository · arXiv:2104.07081
-
VTGAN: Semi-supervised Retinal Image Synthesis and Disease Prediction using Vision Transformers 14 Apr 2021 · 2 repositories · arXiv:2104.06757
-
1-bit LAMB: Communication Efficient Large-Scale Large-Batch Training with LAMB's Convergence Speed 13 Apr 2021 · 1 repository · arXiv:2104.06069
-
Can a Transformer Pass the Wug Test? Tuning Copying Bias in Neural Morphological Inflection Models 13 Apr 2021 · 0 repositories · arXiv:2104.06483
-
Discourse Probing of Pretrained Language Models 13 Apr 2021 · 1 repository · arXiv:2104.05882
-
Large-Scale Contextualised Language Modelling for Norwegian 13 Apr 2021 · 2 repositories · arXiv:2104.06546Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples)
-
Mediators in Determining what Processing BERT Performs First 13 Apr 2021 · 1 repository · arXiv:2104.06400
-
MS2: Multi-Document Summarization of Medical Studies 13 Apr 2021 · 2 repositories · arXiv:2104.06486Syntology official (archive's flag): 3 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
QA-GNN: Reasoning with Language Models and Knowledge Graphs for Question Answering 13 Apr 2021 · 6 repositories · arXiv:2104.06378Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 1 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 17 unverified (of 25 harvested samples) · 5 pointer-only (licence)
-
Semantic maps and metrics for science Semantic maps and metrics for science using deep transformer encoders 13 Apr 2021 · 0 repositories · arXiv:2104.05928
-
Transformer-based Methods for Recognizing Ultra Fine-grained Entities (RUFES) 13 Apr 2021 · 0 repositories · arXiv:2104.06048
-
Understanding Transformers for Bot Detection in Twitter 13 Apr 2021 · 1 repository · arXiv:2104.06182
-
UPB at SemEval-2021 Task 7: Adversarial Multi-Task Learning for Detecting and Rating Humor and Offense 13 Apr 2021 · 0 repositories · arXiv:2104.06063
-
ViT-V-Net: Vision Transformer for Unsupervised Volumetric Medical Image Registration 13 Apr 2021 · 1 repository · arXiv:2104.06468Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 2 honoured, 0 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples)
-
BERT based freedom to operate patent analysis 12 Apr 2021 · 0 repositories · arXiv:2105.00817
-
Cloth Interactive Transformer for Virtual Try-On 12 Apr 2021 · 1 repository · arXiv:2104.05519
-
Family of Origin and Family of Choice: Massively Parallel Lexiconized Iterative Pretraining for Severely Low Resource Machine Translation 12 Apr 2021 · 0 repositories · arXiv:2104.05848
-
Fighting the COVID-19 Infodemic with a Holistic BERT Ensemble 12 Apr 2021 · 1 repository · arXiv:2104.05745
-
Fine-Tuning Transformers for Identifying Self-Reporting Potential Cases and Symptoms of COVID-19 in Tweets 12 Apr 2021 · 1 repository · arXiv:2104.05501
-
Learning dynamic and hierarchical traffic spatiotemporal features with Transformer 12 Apr 2021 · 0 repositories · arXiv:2104.05163
-
Learning to Remove: Towards Isotropic Pre-trained BERT Embedding 12 Apr 2021 · 1 repository · arXiv:2104.05274
-
Learning to Synthesize Data for Semantic Parsing 12 Apr 2021 · 1 repository · arXiv:2104.05827
-
Multilingual Language Models Predict Human Reading Behavior 12 Apr 2021 · 1 repository · arXiv:2104.05433
-
On Representation Learning for Scientific News Articles Using Heterogeneous Knowledge Graphs 12 Apr 2021 · 0 repositories · arXiv:2104.05866
-
Paragraph-level Simplification of Medical Texts 12 Apr 2021 · 1 repository · arXiv:2104.05767
-
Updater-Extractor Architecture for Inductive World State Representations 12 Apr 2021 · 0 repositories · arXiv:2104.05500
-
WHOSe Heritage: Classification of UNESCO World Heritage "Outstanding Universal Value" Documents with Soft Labels 12 Apr 2021 · 1 repository · arXiv:2104.05547
-
Does syntax matter? A strong baseline for Aspect-based Sentiment Analysis with RoBERTa 11 Apr 2021 · 1 repository · arXiv:2104.04986
-
Fine-tuning Encoders for Improved Monolingual and Zero-shot Polylingual Neural Topic Modeling 11 Apr 2021 · 1 repository · arXiv:2104.05064
-
Innovative Bert-based Reranking Language Models for Speech Recognition 11 Apr 2021 · 0 repositories · arXiv:2104.04950
-
UniDrop: A Simple yet Effective Technique to Improve Transformer without Extra Cost 11 Apr 2021 · 0 repositories · arXiv:2104.04946
-
Adapting Language Models for Zero-shot Learning by Meta-tuning on Dataset and Prompt Collections 10 Apr 2021 · 1 repository · arXiv:2104.04670
-
MIPT-NSU-UTMN at SemEval-2021 Task 5: Ensembling Learning with Pre-trained Language Models for Toxic Spans Detection 10 Apr 2021 · 1 repository · arXiv:2104.04739
-
Non-autoregressive Transformer-based End-to-end ASR using BERT 10 Apr 2021 · 0 repositories · arXiv:2104.04805
-
ZS-BERT: Towards Zero-Shot Relation Extraction with Attribute Representation Learning 10 Apr 2021 · 1 repository · arXiv:2104.04697
-
Deep Transformer Networks for Time Series Classification: The NPP Safety Case 9 Apr 2021 · 0 repositories · arXiv:2104.05448
-
KI-BERT: Infusing Knowledge Context for Better Language and Domain Understanding 9 Apr 2021 · 0 repositories · arXiv:2104.08145
-
The Road to Know-Where: An Object-and-Room Informed Sequential BERT for Indoor Vision-Language Navigation 9 Apr 2021 · 1 repository · arXiv:2104.04167
-
Knowledge-Aware Graph-Enhanced GPT-2 for Dialogue State Tracking 9 Apr 2021 · 1 repository · arXiv:2104.04466
-
Text2Chart: A Multi-Staged Chart Generator from Natural Language Text 9 Apr 2021 · 1 repository · arXiv:2104.04584
-
Transformers: "The End of History" for NLP? 9 Apr 2021 · 0 repositories · arXiv:2105.00813
-
Layer Reduction: Accelerating Conformer-Based Self-Supervised Model via Layer Consistency 8 Apr 2021 · 0 repositories · arXiv:2105.00812
-
Lone Pine at SemEval-2021 Task 5: Fine-Grained Detection of Hate Speech Using BERToxic 8 Apr 2021 · 1 repository · arXiv:2104.03506
-
Probing BERT in Hyperbolic Spaces 8 Apr 2021 · 1 repository · arXiv:2104.03869Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified; the one sample that ran constructed an object rather than computing a result (of 3 harvested samples) · 3 pointer-only (licence)
-
Revisiting Simple Neural Probabilistic Language Models 8 Apr 2021 · 1 repository · arXiv:2104.03474Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Uppsala NLP at SemEval-2021 Task 2: Multilingual Language Models for Fine-tuning and Feature Extraction in Word-in-Context Disambiguation 8 Apr 2021 · 0 repositories · arXiv:2104.03767
-
Better Neural Machine Translation by Extracting Linguistic Information from BERT 7 Apr 2021 · 1 repository · arXiv:2104.02831
-
Combining Pre-trained Word Embeddings and Linguistic Features for Sequential Metaphor Identification 7 Apr 2021 · 0 repositories · arXiv:2104.03285
-
Everything's Talkin': Pareidolia Face Reenactment 7 Apr 2021 · 1 repository · arXiv:2104.03061
-
Facial Attribute Transformers for Precise and Robust Makeup Transfer 7 Apr 2021 · 0 repositories · arXiv:2104.02894
-
Interpreting A Pre-trained Model Is A Key For Model Architecture Optimization: A Case Study On Wav2Vec 2.0 7 Apr 2021 · 0 repositories · arXiv:2104.02851
-
Interpreting Verbal Metaphors by Paraphrasing 7 Apr 2021 · 0 repositories · arXiv:2104.03391
-
LI-Net: Large-Pose Identity-Preserving Face Reenactment Network 7 Apr 2021 · 0 repositories · arXiv:2104.02850
-
Seeing Out of tHe bOx: End-to-End Pre-training for Vision-Language Representation Learning 7 Apr 2021 · 3 repositories · arXiv:2104.03135
-
Speak or Chat with Me: End-to-End Spoken Language Understanding System with Flexible Inputs 7 Apr 2021 · 1 repository · arXiv:2104.05752
-
An Empirical Evaluation of Word Embedding Models for Subjectivity Analysis Tasks 6 Apr 2021 · 1 repository
-
Attention Head Masking for Inference Time Content Selection in Abstractive Summarization 6 Apr 2021 · 0 repositories · arXiv:2104.02205
-
CodeTrans: Towards Cracking the Language of Silicon's Code Through Self-Supervised Deep Learning and High Performance Computing 6 Apr 2021 · 1 repository · arXiv:2104.02443
-
Efficient transfer learning for NLP with ELECTRA 6 Apr 2021 · 1 repository · arXiv:2104.02756
-
Fourier Image Transformer 6 Apr 2021 · 1 repository · arXiv:2104.02555
-
HBert + BiasCorp -- Fighting Racism on the Web 6 Apr 2021 · 0 repositories · arXiv:2104.02242
-
LT-LM: a novel non-autoregressive language model for single-shot lattice rescoring 6 Apr 2021 · 1 repository · arXiv:2104.02526
-
MuSLCAT: Multi-Scale Multi-Level Convolutional Attention Transformer for Discriminative Music Modeling on Raw Waveforms 6 Apr 2021 · 0 repositories · arXiv:2104.02309
-
ODE Transformer: An Ordinary Differential Equation-Inspired Model for Neural Machine Translation 6 Apr 2021 · 0 repositories · arXiv:2104.02308
-
Variable selection with missing data in both covariates and outcomes: Imputation and machine learning 6 Apr 2021 · 1 repository · arXiv:2104.02769
-
Variational Transformer Networks for Layout Generation 6 Apr 2021 · 0 repositories · arXiv:2104.02416