Methods › General › Attention Mechanisms › Attention › Papers, page 261
Attention Is All You Need
Attention
Papers archive 2025-07-28
archive papers tagged: 31,583 · with a code link: 13,473 · where Syntology ran a sample: 3,998 (3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,998 of 31,583 tagged: 3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument)
Page 261 of 316: papers 26,001 to 26,100 of 31,583, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Non-Autoregressive Models are Better Multilingual Translators 29 Sep 2021 · 0 repositories
-
Offline Pre-trained Multi-Agent Decision Transformer 29 Sep 2021 · 0 repositories
-
Offline Reinforcement Learning for Large Scale Language Action Spaces 29 Sep 2021 · 0 repositories
-
Pretraining for Language Conditioned Imitation with Transformers 29 Sep 2021 · 0 repositories
-
Privacy-preserving Task-Agnostic Vision Transformer for Image Processing 29 Sep 2021 · 1 repository
-
Rank4Class: Examining Multiclass Classification through the Lens of Learning to Rank 29 Sep 2021 · 0 repositories
-
Robot Intent Recognition Method Based on State Grid Business Office 29 Sep 2021 · 0 repositories
-
ScaLA: Speeding-Up Fine-tuning of Pre-trained Transformer Networks via Efficient and Scalable Adversarial Perturbation 29 Sep 2021 · 0 repositories
-
Scale Efficiently: Insights from Pretraining and Finetuning Transformers 29 Sep 2021 · 0 repositories
-
Scaling the Depth of Vision Transformers via the Fourier Domain Analysis 29 Sep 2021 · 0 repositories
-
Semi-supervised Offline Reinforcement Learning with Pre-trained Decision Transformers 29 Sep 2021 · 0 repositories
-
SeqPATE: Differentially Private Text Generation via Knowledge Distillation 29 Sep 2021 · 0 repositories
-
SiT: Simulation Transformer for Particle-based Physics Simulation 29 Sep 2021 · 0 repositories
-
Spanning Tree-based Graph Generation for Molecules 29 Sep 2021 · 0 repositories
-
Sparse Attention with Learning to Hash 29 Sep 2021 · 0 repositories
-
Specialized Transformers: Faster, Smaller and more Accurate NLP Models 29 Sep 2021 · 0 repositories
-
Subdimensional Expansion Using Attention-Based Learning For Multi-Agent Path Finding 29 Sep 2021 · 1 repository · arXiv:2109.14695
-
Successive POI Recommendation via Brain-inspired Spatiotemporal Aware Representation 29 Sep 2021 · 0 repositories
-
Temporal Action Localization with Global Segmentation Mask Transformers 29 Sep 2021 · 0 repositories
-
Test Time Robustification of Deep Models via Adaptation and Augmentation 29 Sep 2021 · 0 repositories
-
Topic Aware Neural Language Model: Domain Adaptation of Unconditional Text Generation Models 29 Sep 2021 · 0 repositories
-
Training sequence labeling models using prior knowledge 29 Sep 2021 · 0 repositories
-
Transliteration: A Simple Technique For Improving Multilingual Language Modeling 29 Sep 2021 · 0 repositories
-
TransSlowDown: Efficiency Attacks on Neural Machine Translation Systems 29 Sep 2021 · 0 repositories
-
TransTCN: An Attention-based TCN Framework for Sequential Modeling 29 Sep 2021 · 0 repositories
-
Tuformer: Data-Driven Design of Expressive Transformer by Tucker Tensor Representation 29 Sep 2021 · 0 repositories
-
UFO-ViT: High Performance Linear Vision Transformer without Softmax 29 Sep 2021 · 1 repository · arXiv:2109.14382
-
Understanding the Role of Self Attention for Efficient Speech Recognition 29 Sep 2021 · 0 repositories
-
Video Forgery Detection Using Multiple Cues on Fusion of EfficientNet and Swin Transformer 29 Sep 2021 · 0 repositories
-
VUT: Versatile UI Transformer for Multimodal Multi-Task User Interface Modeling 29 Sep 2021 · 0 repositories
-
Fine-tuning Vision Transformers for the Prediction of State Variables in Ising Models 28 Sep 2021 · 0 repositories · arXiv:2109.13925
-
How Different Text-preprocessing Techniques Using The BERT Model Affect The Gender Profiling of Authors 28 Sep 2021 · 0 repositories · arXiv:2109.13890
-
Nana-HDR: A Non-attentive Non-autoregressive Hybrid Model for TTS 28 Sep 2021 · 0 repositories · arXiv:2109.13673
-
RAFT: A Real-World Few-Shot Text Classification Benchmark 28 Sep 2021 · 1 repository · arXiv:2109.14076
-
Single-dataset Experts for Multi-dataset Question Answering 28 Sep 2021 · 1 repository · arXiv:2109.13880Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 1 violated, 8 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 11 harvested samples) · 3 pointer-only (licence)
-
What to Prioritize? Natural Language Processing for the Development of a Modern Bug Tracking Solution in Hardware Development 28 Sep 2021 · 0 repositories · arXiv:2109.13825
-
Effective Use of Graph Convolution Network and Contextual Sub-Tree forCommodity News Event Extraction 27 Sep 2021 · 1 repository · arXiv:2109.12781
-
Fast-MD: Fast Multi-Decoder End-to-End Speech Translation with Non-Autoregressive Hidden Intermediates 27 Sep 2021 · 1 repository · arXiv:2109.12804
-
Improving Stack Overflow question title generation with copying enhanced CodeBERT model and bi-modal information 27 Sep 2021 · 1 repository · arXiv:2109.13073
-
Integrated Training for Sequence-to-Sequence Models Using Non-Autoregressive Transformer 27 Sep 2021 · 0 repositories · arXiv:2109.12950
-
PASS: An ImageNet replacement for self-supervised pretraining without humans 27 Sep 2021 · 1 repository · arXiv:2109.13228Syntology official: no sample here; runs from other or unrecorded repositories · 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 0 violated, 9 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 14 harvested samples) · 9 pointer-only (licence)
-
Patterns of Lexical Ambiguity in Contextualised Language Models 27 Sep 2021 · 0 repositories · arXiv:2109.13032
-
Sparse Spatial Transformers for Few-Shot Learning 27 Sep 2021 · 1 repository · arXiv:2109.12932Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples)
-
TURINGBENCH: A Benchmark Environment for Turing Test in the Age of Neural Text Generation 27 Sep 2021 · 3 repositories · arXiv:2109.13296Syntology 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
Understanding and Overcoming the Challenges of Efficient Transformer Quantization 27 Sep 2021 · 1 repository · arXiv:2109.12948Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Improving Question Answering Performance Using Knowledge Distillation and Active Learning 26 Sep 2021 · 1 repository · arXiv:2109.12662Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Multi-Transformer: A New Neural Network-Based Architecture for Forecasting S&P Volatility 26 Sep 2021 · 1 repository · arXiv:2109.12621
-
On the Prunability of Attention Heads in Multilingual BERT 26 Sep 2021 · 0 repositories · arXiv:2109.12683
-
Parallel Refinements for Lexically Constrained Text Generation with BART 26 Sep 2021 · 1 repository · arXiv:2109.12487
-
Vision Transformer Hashing for Image Retrieval 26 Sep 2021 · 1 repository · arXiv:2109.12564
-
ViT Cane: Visual Assistant for the Visually Impaired 26 Sep 2021 · 0 repositories · arXiv:2109.13857
-
BiTr-Unet: a CNN-Transformer Combined Network for MRI Brain Tumor Segmentation 25 Sep 2021 · 1 repository · arXiv:2109.12271
-
Finetuning Transformer Models to Build ASAG System 25 Sep 2021 · 0 repositories · arXiv:2109.12300
-
Learning to Selectively Learn for Weakly-supervised Paraphrase Generation 25 Sep 2021 · 0 repositories · arXiv:2109.12457
-
TEMGNet: Deep Transformer-based Decoding of Upperlimb sEMG for Hand Gestures Recognition 25 Sep 2021 · 0 repositories · arXiv:2109.12379
-
AES Systems Are Both Overstable And Oversensitive: Explaining Why And Proposing Defenses 24 Sep 2021 · 0 repositories · arXiv:2109.11728
-
DACT-BERT: Differentiable Adaptive Computation Time for an Efficient BERT Inference 24 Sep 2021 · 0 repositories · arXiv:2109.11745
-
Dense Contrastive Visual-Linguistic Pretraining 24 Sep 2021 · 0 repositories · arXiv:2109.11778
-
Identification of Enzymatic Active Sites with Unsupervised Language Modeling 24 Sep 2021 · 0 repositories
-
Lacking the embedding of a word? Look it up into a traditional dictionary 24 Sep 2021 · 0 repositories · arXiv:2109.11763
-
Leveraging Pretrained Models for Automatic Summarization of Doctor-Patient Conversations 24 Sep 2021 · 1 repository · arXiv:2109.12174
-
Localizing Infinity-shaped fishes: Sketch-guided object localization in the wild 24 Sep 2021 · 0 repositories · arXiv:2109.11874
-
Long-Range Transformers for Dynamic Spatiotemporal Forecasting 24 Sep 2021 · 2 repositories · arXiv:2109.12218Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
Robustness and Sensitivity of BERT Models Predicting Alzheimer's Disease from Text 24 Sep 2021 · 0 repositories · arXiv:2109.11888
-
Transformers Generalize Linearly 24 Sep 2021 · 1 repository · arXiv:2109.12036
-
Breaking BERT: Understanding its Vulnerabilities for Named Entity Recognition through Adversarial Attack 23 Sep 2021 · 1 repository · arXiv:2109.11308
-
Dependency Structure for News Document Summarization 23 Sep 2021 · 0 repositories · arXiv:2109.11199
-
OH-Former: Omni-Relational High-Order Transformer for Person Re-Identification 23 Sep 2021 · 0 repositories · arXiv:2109.11159
-
Putting Words in BERT's Mouth: Navigating Contextualized Vector Spaces with Pseudowords 23 Sep 2021 · 1 repository · arXiv:2109.11491
-
The Volctrans GLAT System: Non-autoregressive Translation Meets WMT21 23 Sep 2021 · 0 repositories · arXiv:2109.11247
-
Alzheimers Dementia Detection using Acoustic & Linguistic features and Pre-Trained BERT 22 Sep 2021 · 0 repositories · arXiv:2109.11010
-
Controlled Evaluation of Grammatical Knowledge in Mandarin Chinese Language Models 22 Sep 2021 · 1 repository · arXiv:2109.11058
-
DialogueBERT: A Self-Supervised Learning based Dialogue Pre-training Encoder 22 Sep 2021 · 0 repositories · arXiv:2109.10480
-
Hierarchical Multimodal Transformer to Summarize Videos 22 Sep 2021 · 0 repositories · arXiv:2109.10559
-
Improving 360 Monocular Depth Estimation via Non-local Dense Prediction Transformer and Joint Supervised and Self-supervised Learning 22 Sep 2021 · 1 repository · arXiv:2109.10563Syntology 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
KD-VLP: Improving End-to-End Vision-and-Language Pretraining with Object Knowledge Distillation 22 Sep 2021 · 1 repository · arXiv:2109.10504
-
Language Models as Recommender Systems: Evaluations and Limitations 22 Sep 2021 · 0 repositories
-
Predicting Efficiency/Effectiveness Trade-offs for Dense vs. Sparse Retrieval Strategy Selection 22 Sep 2021 · 0 repositories · arXiv:2109.10739
-
Recursively Summarizing Books with Human Feedback 22 Sep 2021 · 0 repositories · arXiv:2109.10862
-
Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers 22 Sep 2021 · 3 repositories · arXiv:2109.10686Syntology community repositories only · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 4 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 12 harvested samples) · 1 pointer-only (licence)
-
T6D-Direct: Transformers for Multi-Object 6D Pose Direct Regression 22 Sep 2021 · 0 repositories · arXiv:2109.10948
-
The NiuTrans Machine Translation Systems for WMT21 22 Sep 2021 · 0 repositories · arXiv:2109.10485
-
Unsupervised Contextualized Document Representation 22 Sep 2021 · 1 repository · arXiv:2109.10509
-
A Comprehensive Review on Summarizing Financial News Using Deep Learning 21 Sep 2021 · 0 repositories · arXiv:2109.10118
-
BERTweetFR : Domain Adaptation of Pre-Trained Language Models for French Tweets 21 Sep 2021 · 0 repositories · arXiv:2109.10234
-
DS-Net++: Dynamic Weight Slicing for Efficient Inference in CNNs and Transformers 21 Sep 2021 · 1 repository · arXiv:2109.10060
-
InvBERT: Reconstructing Text from Contextualized Word Embeddings by inverting the BERT pipeline 21 Sep 2021 · 0 repositories · arXiv:2109.10104
-
LOTR: Face Landmark Localization Using Localization Transformer 21 Sep 2021 · 0 repositories · arXiv:2109.10057
-
Multi-Task Learning with Sentiment, Emotion, and Target Detection to Recognize Hate Speech and Offensive Language 21 Sep 2021 · 0 repositories · arXiv:2109.10255
-
Representation Learning for Short Text Clustering 21 Sep 2021 · 0 repositories · arXiv:2109.09894
-
TrOCR: Transformer-based Optical Character Recognition with Pre-trained Models 21 Sep 2021 · 8 repositories · arXiv:2109.10282Syntology community repositories only · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
A Plug-and-Play Method for Controlled Text Generation 20 Sep 2021 · 1 repository · arXiv:2109.09707
-
BARTpho: Pre-trained Sequence-to-Sequence Models for Vietnamese 20 Sep 2021 · 3 repositories · arXiv:2109.09701
-
BERT Cannot Align Characters 20 Sep 2021 · 0 repositories · arXiv:2109.09700
-
BERT Has Uncommon Sense: Similarity Ranking for Word Sense BERTology 20 Sep 2021 · 1 repository · arXiv:2109.09780
-
Dyadformer: A Multi-modal Transformer for Long-Range Modeling of Dyadic Interactions 20 Sep 2021 · 0 repositories · arXiv:2109.09487
-
MFEViT: A Robust Lightweight Transformer-based Network for Multimodal 2D+3D Facial Expression Recognition 20 Sep 2021 · 0 repositories · arXiv:2109.13086
-
Model Bias in NLP -- Application to Hate Speech Classification using transfer learning techniques 20 Sep 2021 · 0 repositories · arXiv:2109.09725
-
Well Googled is Half Done: Multimodal Forecasting of New Fashion Product Sales with Image-based Google Trends 20 Sep 2021 · 1 repository · arXiv:2109.09824Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
CLIFF: Contrastive Learning for Improving Faithfulness and Factuality in Abstractive Summarization 19 Sep 2021 · 3 repositories · arXiv:2109.09209