Methods › General › Attention Mechanisms › Attention › Papers, page 280
Attention Is All You Need
Attention
Papers archive 2025-07-28
archive papers tagged: 31,583 · with a code link: 13,473 · where Syntology ran a sample: 3,998 (3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,998 of 31,583 tagged: 3,366 with a run with no instrument failure, 632 where every run was a failure of Syntology's instrument)
Page 280 of 316: papers 27,901 to 28,000 of 31,583, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Translational Equivariance in Kernelizable Attention 15 Feb 2021 · 1 repository · arXiv:2102.07680
-
Within-Document Event Coreference with BERT-Based Contextualized Representations 15 Feb 2021 · 0 repositories · arXiv:2102.09600
-
indicnlp@kgp at DravidianLangTech-EACL2021: Offensive Language Identification in Dravidian Languages 14 Feb 2021 · 1 repository · arXiv:2102.07150
-
indicnlp@ kgp at DravidianLangTech-EACL2021: Offensive Language Identification in Dravidian Languages 14 Feb 2021 · 1 repository
-
Query-by-Example Keyword Spotting system using Multi-head Attention and Softtriple Loss 14 Feb 2021 · 0 repositories · arXiv:2102.07061
-
Characterizing English Variation across Social Media Communities with BERT 12 Feb 2021 · 1 repository · arXiv:2102.06820
-
Dancing along Battery: Enabling Transformer with Run-time Reconfigurability on Mobile Devices 12 Feb 2021 · 0 repositories · arXiv:2102.06336
-
Dynamic Precision Analog Computing for Neural Networks 12 Feb 2021 · 1 repository · arXiv:2102.06365
-
Exploring Classic and Neural Lexical Translation Models for Information Retrieval: Interpretability, Effectiveness, and Efficiency Benefits 12 Feb 2021 · 2 repositories · arXiv:2102.06815
-
Improving Zero-shot Neural Machine Translation on Language-specific Encoders-Decoders 12 Feb 2021 · 0 repositories · arXiv:2102.06578
-
Multiversal views on language models 12 Feb 2021 · 0 repositories · arXiv:2102.06391
-
Optimizing Inference Performance of Transformers on CPUs 12 Feb 2021 · 0 repositories · arXiv:2102.06621
-
Transformer Language Models with LSTM-based Cross-utterance Information Representation 12 Feb 2021 · 1 repository · arXiv:2102.06474
-
Deep Reinforcement Learning for Combinatorial Optimization: Covering Salesman Problems 11 Feb 2021 · 0 repositories · arXiv:2102.05875
-
Proof Artifact Co-training for Theorem Proving with Language Models 11 Feb 2021 · 4 repositories · arXiv:2102.06203Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Text Compression-aided Transformer Encoding 11 Feb 2021 · 0 repositories · arXiv:2102.05951
-
MAIN: Multihead-Attention Imputation Networks 10 Feb 2021 · 0 repositories · arXiv:2102.05428
-
NAST: Non-Autoregressive Spatial-Temporal Transformer for Time Series Forecasting 10 Feb 2021 · 1 repository · arXiv:2102.05624
-
AuGPT: Auxiliary Tasks and Data Augmentation for End-To-End Dialogue with Pre-Trained Language Models 9 Feb 2021 · 1 repository · arXiv:2102.05126
-
Bayesian Transformer Language Models for Speech Recognition 9 Feb 2021 · 0 repositories · arXiv:2102.04754
-
Conversational Query Rewriting with Self-supervised Learning 9 Feb 2021 · 0 repositories · arXiv:2102.04708
-
Joint Intent Detection and Slot Filling with Wheel-Graph Attention Networks 9 Feb 2021 · 0 repositories · arXiv:2102.04610
-
NewsBERT: Distilling Pre-trained Language Model for Intelligent News Application 9 Feb 2021 · 0 repositories · arXiv:2102.04887
-
Point Cloud Transformers applied to Collider Physics 9 Feb 2021 · 1 repository · arXiv:2102.05073
-
Transfer Learning Approach for Arabic Offensive Language Detection System -- BERT-Based Model 9 Feb 2021 · 0 repositories · arXiv:2102.05708
-
A Hybrid Task-Oriented Dialog System with Domain and Task Adaptive Pretraining 8 Feb 2021 · 0 repositories · arXiv:2102.04506
-
Colorization Transformer 8 Feb 2021 · 2 repositories · arXiv:2102.04432
-
Generating Fake Cyber Threat Intelligence Using Transformer-Based Models 8 Feb 2021 · 0 repositories · arXiv:2102.04351
-
Bias Out-of-the-Box: An Empirical Analysis of Intersectional Occupational Biases in Popular Generative Language Models 8 Feb 2021 · 1 repository · arXiv:2102.04130Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
TransReID: Transformer-based Object Re-Identification 8 Feb 2021 · 4 repositories · arXiv:2102.04378
-
TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation 8 Feb 2021 · 22 repositories · arXiv:2102.04306Syntology official: no sample here; runs from other or unrecorded repositories · 6 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Wake Word Detection with Streaming Transformers 8 Feb 2021 · 0 repositories · arXiv:2102.04488
-
Nyströmformer: A Nyström-Based Algorithm for Approximating Self-Attention 7 Feb 2021 · 10 repositories · arXiv:2102.03902Syntology official: no sample here; runs from other or unrecorded repositories · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 1 pointer-only (licence)
-
Spoiler Alert: Using Natural Language Processing to Detect Spoilers in Book Reviews 7 Feb 2021 · 1 repository · arXiv:2102.03882
-
Jointly Improving Language Understanding and Generation with Quality-Weighted Weak Supervision of Automatic Labeling 6 Feb 2021 · 0 repositories · arXiv:2102.03551
-
Neural Data-to-Text Generation with LM-based Text Augmentation 6 Feb 2021 · 0 repositories · arXiv:2102.03556
-
baller2vec: A Multi-Entity Transformer For Multi-Agent Spatiotemporal Modeling 5 Feb 2021 · 1 repository · arXiv:2102.03291Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
PipeTransformer: Automated Elastic Pipelining for Distributed Training of Transformers 5 Feb 2021 · 1 repository · arXiv:2102.03161Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
RpBERT: A Text-image Relation Propagation-based BERT Model for Multimodal NER 5 Feb 2021 · 1 repository · arXiv:2102.02967
-
Understanding Emails and Drafting Responses -- An Approach Using GPT-3 5 Feb 2021 · 0 repositories · arXiv:2102.03062
-
ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision 5 Feb 2021 · 6 repositories · arXiv:2102.03334Syntology official: harvested, nothing ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified; the one sample that ran constructed an object rather than computing a result (of 4 harvested samples) · 1 pointer-only (licence)
-
1-bit Adam: Communication Efficient Large-Scale Training with Adam's Convergence Speed 4 Feb 2021 · 2 repositories · arXiv:2102.02888
-
Adaptive Semiparametric Language Models 4 Feb 2021 · 0 repositories · arXiv:2102.02557
-
Hierarchical Multi-head Attentive Network for Evidence-aware Fake News Detection 4 Feb 2021 · 1 repository · arXiv:2102.02680
-
Understanding the Capabilities, Limitations, and Societal Impact of Large Language Models 4 Feb 2021 · 0 repositories · arXiv:2102.02503
-
Bootstrapping Multilingual AMR with Contextual Word Alignments 3 Feb 2021 · 0 repositories · arXiv:2102.02189
-
HeBERT & HebEMO: a Hebrew BERT Model and a Tool for Polarity Analysis and Emotion Recognition 3 Feb 2021 · 0 repositories · arXiv:2102.01909
-
MUFASA: Multimodal Fusion Architecture Search for Electronic Health Records 3 Feb 2021 · 0 repositories · arXiv:2102.02340
-
Introduction to Neural Transfer Learning with Transformers for Social Science Text Analysis 3 Feb 2021 · 0 repositories · arXiv:2102.02111
-
Mind the Gap: Assessing Temporal Generalization in Neural Language Models 3 Feb 2021 · 1 repository · arXiv:2102.01951
-
Relaxed Transformer Decoders for Direct Action Proposal Generation 3 Feb 2021 · 2 repositories · arXiv:2102.01894Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
Towards Natural and Controllable Cross-Lingual Voice Conversion Based on Neural TTS Model and Phonetic Posteriorgram 3 Feb 2021 · 0 repositories · arXiv:2102.01991
-
AutoFreeze: Automatically Freezing Model Blocks to Accelerate Fine-tuning 2 Feb 2021 · 1 repository · arXiv:2102.01386Syntology official: no sample here; runs from other or unrecorded repositories · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Clickbait Headline Detection in Indonesian News Sites using Multilingual Bidirectional Encoder Representations from Transformers (M-BERT) 2 Feb 2021 · 0 repositories · arXiv:2102.01497
-
Automated Query Reformulation for Efficient Search based on Query Logs From Stack Overflow 1 Feb 2021 · 1 repository · arXiv:2102.00826
-
GTAE: Graph-Transformer based Auto-Encoders for Linguistic-Constrained Text Style Transfer 1 Feb 2021 · 0 repositories · arXiv:2102.00769
-
Improving Distantly-Supervised Relation Extraction through BERT-based Label & Instance Embeddings 1 Feb 2021 · 1 repository · arXiv:2102.01156
-
"Is depression related to cannabis?": A knowledge-infused model for Entity and Relation Extraction with Limited Supervision 1 Feb 2021 · 0 repositories · arXiv:2102.01222
-
Polyphone Disambiguation in Mandarin Chinese with Semi-Supervised Learning 1 Feb 2021 · 0 repositories · arXiv:2102.00621
-
Scaling Federated Learning for Fine-tuning of Large Language Models 1 Feb 2021 · 0 repositories · arXiv:2102.00875
-
SJ_AJ@DravidianLangTech-EACL2021: Task-Adaptive Pre-Training of Multilingual BERT models for Offensive Language Identification 1 Feb 2021 · 1 repository · arXiv:2102.01051
-
Text-to-hashtag Generation using Seq2seq Learning 1 Feb 2021 · 1 repository · arXiv:2102.00904
-
A Runtime-Based Computational Performance Predictor for Deep Neural Network Training 31 Jan 2021 · 1 repository · arXiv:2102.00527
-
[Re] Reproducing Learning to Deceive With Attention-Based Explanations 31 Jan 2021 · 1 repository
-
Short Text Clustering with Transformers 31 Jan 2021 · 0 repositories · arXiv:2102.00541
-
Adversarially learning disentangled speech representations for robust multi-factor voice conversion 30 Jan 2021 · 0 repositories · arXiv:2102.00184
-
EmpathBERT: A BERT-based Framework for Demographic-aware Empathy Prediction 30 Jan 2021 · 0 repositories · arXiv:2102.00272
-
Learning From Human Correction 30 Jan 2021 · 1 repository · arXiv:2102.00225
-
ShufText: A Simple Black Box Approach to Evaluate the Fragility of Text Classification Models 30 Jan 2021 · 0 repositories · arXiv:2102.00238
-
Speech Recognition by Simply Fine-tuning BERT 30 Jan 2021 · 0 repositories · arXiv:2102.00291
-
The characteristic equation of the exceptional Jordan algebra: its eigenvalues, and their possible connection with the mass ratios of quarks and leptons 30 Jan 2021 · 0 repositories
-
Fine-tuning BERT-based models for Plant Health Bulletin Classification 29 Jan 2021 · 1 repository · arXiv:2102.00838
-
Synthesizing Monolingual Data for Neural Machine Translation 29 Jan 2021 · 0 repositories · arXiv:2101.12462
-
Enhancing the Transformer Decoder with Transition-based Syntax 29 Jan 2021 · 1 repository · arXiv:2101.12640
-
A Graph-based Relevance Matching Model for Ad-hoc Retrieval 28 Jan 2021 · 1 repository · arXiv:2101.11873
-
BERTaú: Itaú BERT for digital customer service 28 Jan 2021 · 0 repositories · arXiv:2101.12015
-
Consequence of Enterprise Resource Planning in the Environs of Pedagogical Organization. 28 Jan 2021 · 0 repositories
-
Consequence of Enterprise Resource Planning in the Environs of Pedagogical Organization. 28 Jan 2021 · 0 repositories
-
LSTM-SAKT: LSTM-Encoded SAKT-like Transformer for Knowledge Tracing 28 Jan 2021 · 0 repositories · arXiv:2102.00845
-
Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet 28 Jan 2021 · 13 repositories · arXiv:2101.11986Syntology official (archive's flag): 4 ran · 21 ran (of which 16 constructed an object rather than computing a result; 21 with no instrument failure: 1 honoured, 0 violated, 20 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 26 harvested samples) · 8 pointer-only (licence)
-
An explainable Transformer-based deep learning model for the prediction of incident heart failure 27 Jan 2021 · 0 repositories · arXiv:2101.11359
-
Bottleneck Transformers for Visual Recognition 27 Jan 2021 · 13 repositories · arXiv:2101.11605Syntology 26 ran (of which 9 constructed an object rather than computing a result; 19 with no instrument failure: 1 honoured, 0 violated, 18 with no contract checked; 7 where Syntology's instrument failed) · 23 unverified (of 49 harvested samples) · 8 pointer-only (licence)
-
Exploring multi-task multi-lingual learning of transformer models for hate speech and offensive speech identification in social media 27 Jan 2021 · 1 repository · arXiv:2101.11155
-
KoreALBERT: Pretraining a Lite BERT Model for Korean Language Understanding 27 Jan 2021 · 0 repositories · arXiv:2101.11363
-
On the Evolution of Syntactic Information Encoded by BERT's Contextualized Representations 27 Jan 2021 · 0 repositories · arXiv:2101.11492
-
Spatial-Channel Transformer Network for Trajectory Prediction on the Traffic Scenes 27 Jan 2021 · 0 repositories · arXiv:2101.11472
-
Analyzing Zero-shot Cross-lingual Transfer in Supervised NLP Tasks 26 Jan 2021 · 0 repositories · arXiv:2101.10649
-
Attention Can Reflect Syntactic Structure (If You Let It) 26 Jan 2021 · 0 repositories · arXiv:2101.10927
-
CLiMP: A Benchmark for Chinese Language Model Evaluation 26 Jan 2021 · 0 repositories · arXiv:2101.11131
-
CPTR: Full Transformer Network for Image Captioning 26 Jan 2021 · 0 repositories · arXiv:2101.10804
-
Deep Subjecthood: Higher-Order Grammatical Features in Multilingual BERT 26 Jan 2021 · 1 repository · arXiv:2101.11043
-
Evaluation of BERT and ALBERT Sentence Embedding Performance on Downstream NLP Tasks 26 Jan 2021 · 0 repositories · arXiv:2101.10642
-
First Align, then Predict: Understanding the Cross-Lingual Ability of Multilingual BERT 26 Jan 2021 · 1 repository · arXiv:2101.11109
-
Named Entity Recognition in the Style of Object Detection 26 Jan 2021 · 0 repositories · arXiv:2101.11122
-
Regulatory Compliance through Doc2Doc Information Retrieval: A case study in EU/UK legislation where text similarity has limitations 26 Jan 2021 · 0 repositories · arXiv:2101.10726
-
The Wireless Control Bus: Enabling Efficient Multi-hop Event-Triggered Control with Concurrent Transmissions 26 Jan 2021 · 0 repositories · arXiv:2101.10961
-
A Hybrid Approach to Measure Semantic Relatedness in Biomedical Concepts 25 Jan 2021 · 0 repositories · arXiv:2101.10196
-
EGFI: Drug-Drug Interaction Extraction and Generation with Fusion of Enriched Entity and Sentence Information 25 Jan 2021 · 1 repository · arXiv:2101.09914
-
Randomized Deep Structured Prediction for Discourse-Level Processing 25 Jan 2021 · 0 repositories · arXiv:2101.10435
-
Does Dialog Length matter for Next Response Selection task? An Empirical Study 24 Jan 2021 · 0 repositories · arXiv:2101.09647