Methods › Natural Language Processing › Autoregressive Transformers › Transformer › Papers, page 136
Transformer
Papers archive 2025-07-28
archive papers tagged: 13,999 · with a code link: 6,572 · where Syntology ran a sample: 2,248 (1,919 with a run with no instrument failure, 329 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,248 of 13,999 tagged: 1,919 with a run with no instrument failure, 329 where every run was a failure of Syntology's instrument)
Page 136 of 140: papers 13,501 to 13,600 of 13,999, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Deja-vu: Double Feature Presentation and Iterated Loss in Deep Transformer Networks 23 Oct 2019 · 2 repositories · arXiv:1910.10324
-
TCT: A Cross-supervised Learning Method for Multimodal Sequence Representation 23 Oct 2019 · 0 repositories · arXiv:1911.05186
-
Complex Transformer: A Framework for Modeling Complex-Valued Sequence 22 Oct 2019 · 1 repository · arXiv:1910.10202Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
Exchangeable deep neural networks for set-to-set matching and learning 22 Oct 2019 · 2 repositories · arXiv:1910.09972
-
Depth-Adaptive Transformer 22 Oct 2019 · 0 repositories · arXiv:1910.10073
-
Improving Transformer-based Speech Recognition Using Unsupervised Pre-training 22 Oct 2019 · 1 repository · arXiv:1910.09932
-
Sequence-to-sequence Singing Synthesis Using the Feed-forward Transformer 22 Oct 2019 · 0 repositories · arXiv:1910.09989
-
Transformer-based Acoustic Modeling for Hybrid Speech Recognition 22 Oct 2019 · 0 repositories · arXiv:1910.09799
-
Learning to Make Generalizable and Diverse Predictions for Retrosynthesis 21 Oct 2019 · 0 repositories · arXiv:1910.09688
-
Transformer-CNN: Fast and Reliable tool for QSAR 21 Oct 2019 · 1 repository · arXiv:1911.06603
-
Personalized Graph Neural Networks with Attention Mechanism for Session-Aware Recommendation 20 Oct 2019 · 3 repositories · arXiv:1910.08887
-
Fully Quantized Transformer for Machine Translation 17 Oct 2019 · 0 repositories · arXiv:1910.10485
-
Predicting retrosynthetic pathways using a combined linguistic model and hyper-graph exploration strategy 17 Oct 2019 · 0 repositories · arXiv:1910.08036
-
Question Classification with Deep Contextualized Transformer 17 Oct 2019 · 1 repository · arXiv:1910.10492
-
Efficiency through Auto-Sizing: Notre Dame NLP's Submission to the WNGT 2019 Efficiency Task 16 Oct 2019 · 0 repositories · arXiv:1910.07134
-
Evolution of transfer learning in natural language processing 16 Oct 2019 · 0 repositories · arXiv:1910.07370
-
Imperial College London Submission to VATEX Video Captioning Task 16 Oct 2019 · 0 repositories · arXiv:1910.07482
-
Injecting Hierarchy with U-Net Transformers 16 Oct 2019 · 2 repositories · arXiv:1910.10488
-
Analyzing the Forgetting Problem in the Pretrain-Finetuning of Dialogue Response Models 16 Oct 2019 · 0 repositories · arXiv:1910.07117
-
Transformer ASR with Contextual Block Processing 16 Oct 2019 · 0 repositories · arXiv:1910.07204
-
Using Whole Document Context in Neural Machine Translation 16 Oct 2019 · 0 repositories · arXiv:1910.07481
-
Enhancing the Transformer with Explicit Relational Encoding for Math Problem Solving 15 Oct 2019 · 3 repositories · arXiv:1910.06611Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 1 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 11 harvested samples)
-
Facebook AI's WAT19 Myanmar-English Translation Task Submission 15 Oct 2019 · 0 repositories · arXiv:1910.06848
-
Structured Pruning of a BERT-based Question Answering Model 14 Oct 2019 · 0 repositories · arXiv:1910.06360
-
Q8BERT: Quantized 8Bit BERT 14 Oct 2019 · 5 repositories · arXiv:1910.06188Syntology official (archive's flag): 5 ran · 10 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 2 honoured, 0 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 11 harvested samples) · 3 pointer-only (licence)
-
Transformers without Tears: Improving the Normalization of Self-Attention 14 Oct 2019 · 5 repositories · arXiv:1910.05895Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Stabilizing Transformers for Reinforcement Learning 13 Oct 2019 · 5 repositories · arXiv:1910.06764Syntology 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 2 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
On Recognizing Texts of Arbitrary Shapes with 2D Self-Attention 10 Oct 2019 · 2 repositories · arXiv:1910.04396
-
On the adequacy of untuned warmup for adaptive optimization 9 Oct 2019 · 1 repository · arXiv:1910.04209Syntology 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
PipeMare: Asynchronous Pipeline Parallel DNN Training 9 Oct 2019 · 0 repositories · arXiv:1910.05124
-
HuggingFace's Transformers: State-of-the-art Natural Language Processing 9 Oct 2019 · 9 repositories · arXiv:1910.03771Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
SMArT: Training Shallow Memory-aware Transformers for Robotic Explainability 7 Oct 2019 · 0 repositories · arXiv:1910.02974
-
How Transformer Revitalizes Character-based Neural Machine Translation: An Investigation on Japanese-Vietnamese Translation Systems 5 Oct 2019 · 1 repository · arXiv:1910.02238
-
Neural Zero-Inflated Quality Estimation Model For Automatic Speech Recognition System 3 Oct 2019 · 0 repositories · arXiv:1910.01289
-
Application of Low-resource Machine Translation Techniques to Russian-Tatar Language Pair 1 Oct 2019 · 0 repositories · arXiv:1910.00368
-
Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation 1 Oct 2019 · 1 repository · arXiv:1910.06717
-
Dialogue Transformers 1 Oct 2019 · 1 repository · arXiv:1910.00486
-
Efficiency Metrics for Data-Driven Models: A Text Summarization Case Study 1 Oct 2019 · 0 repositories
-
Entangled Transformer for Image Captioning 1 Oct 2019 · 0 repositories
-
Grammatical Error Correction in Low-Resource Scenarios 1 Oct 2019 · 1 repository · arXiv:1910.00353
-
Better Document-Level Machine Translation with Bayes' Rule 1 Oct 2019 · 0 repositories · arXiv:1910.00553
-
Monotonic Multihead Attention 26 Sep 2019 · 3 repositories · arXiv:1909.12406
-
Universal Graph Transformer Self-Attention Networks 26 Sep 2019 · 1 repository · arXiv:1909.11855Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Attention Convolutional Binary Neural Tree for Fine-Grained Visual Categorization 25 Sep 2019 · 2 repositories · arXiv:1909.11378
-
Reducing Transformer Depth on Demand with Structured Dropout 25 Sep 2019 · 5 repositories · arXiv:1909.11556
-
Knowledge-Enriched Transformer for Emotion Detection in Textual Conversations 24 Sep 2019 · 1 repository · arXiv:1909.10681Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Efficiently Reusing Old Models Across Languages via Transfer Learning 24 Sep 2019 · 0 repositories · arXiv:1909.10955
-
Unified Vision-Language Pre-Training for Image Captioning and VQA 24 Sep 2019 · 3 repositories · arXiv:1909.11059Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 1 honoured, 0 violated, 7 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (of 14 harvested samples) · 14 pointer-only (licence)
-
TinyBERT: Distilling BERT for Natural Language Understanding 23 Sep 2019 · 10 repositories · arXiv:1909.10351Syntology official: harvested, nothing ran · 0 ran · 4 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Self-attention based end-to-end Hindi-English Neural Machine Translation 21 Sep 2019 · 0 repositories · arXiv:1909.09779
-
Adversarial Learning of General Transformations for Data Augmentation 21 Sep 2019 · 0 repositories · arXiv:1909.09801
-
Improved Variational Neural Machine Translation by Promoting Mutual Information 19 Sep 2019 · 0 repositories · arXiv:1909.09237
-
Language models and Automated Essay Scoring 18 Sep 2019 · 1 repository · arXiv:1909.09482
-
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism 17 Sep 2019 · 10 repositories · arXiv:1909.08053Syntology community repositories only · 12 ran (of which 2 constructed an object rather than computing a result; 7 with no instrument failure: 4 honoured, 0 violated, 3 with no contract checked; 5 where Syntology's instrument failed) · 35 unverified (of 47 harvested samples) · 15 pointer-only (licence)
-
Hybrid Neural Models For Sequence Modelling: The Best Of Three Worlds 16 Sep 2019 · 0 repositories · arXiv:1909.07102
-
Multilingual Neural Machine Translation for Zero-Resource Languages 16 Sep 2019 · 1 repository · arXiv:1909.07342
-
Automatically Extracting Challenge Sets for Non local Phenomena in Neural Machine Translation 15 Sep 2019 · 1 repository · arXiv:1909.06814
-
I-MAD: Interpretable Malware Detector Using Galaxy Transformer 15 Sep 2019 · 0 repositories · arXiv:1909.06865
-
Benchmarking the Performance and Energy Efficiency of AI Accelerators for AI Training 15 Sep 2019 · 0 repositories · arXiv:1909.06842
-
Efficiency Metrics for Data-Driven Models: A Text Summarization Case Study 14 Sep 2019 · 0 repositories · arXiv:1909.06618
-
Ouroboros: On Accelerating Training of Transformer-Based Language Models 14 Sep 2019 · 1 repository · arXiv:1909.06695
-
Tree Transformer: Integrating Tree Structures into Self-Attention 14 Sep 2019 · 3 repositories · arXiv:1909.06639Syntology community repositories only · 9 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 10 harvested samples) · 2 pointer-only (licence)
-
A Comparative Study on Transformer vs RNN in Speech Applications 13 Sep 2019 · 2 repositories · arXiv:1909.06317
-
Neural Machine Translation with 4-Bit Precision and Beyond 13 Sep 2019 · 0 repositories · arXiv:1909.06091
-
SANVis: Visual Analytics for Understanding Self-Attention Networks 13 Sep 2019 · 0 repositories · arXiv:1909.09595
-
Scene Graph Parsing by Attention Graph 13 Sep 2019 · 0 repositories · arXiv:1909.06273
-
CUNI System for the Building Educational Applications 2019 Shared Task: Grammatical Error Correction 12 Sep 2019 · 0 repositories · arXiv:1909.05553
-
Global Locality in Biomedical Relation and Event Extraction 11 Sep 2019 · 0 repositories · arXiv:1909.04822
-
How Does BERT Answer Questions? A Layer-Wise Analysis of Transformer Representations 11 Sep 2019 · 2 repositories · arXiv:1909.04925Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Forecaster: A Graph Transformer for Forecasting Spatial and Time-Dependent Data 9 Sep 2019 · 0 repositories · arXiv:1909.04019
-
Pretrained Language Models for Sequential Sentence Classification 9 Sep 2019 · 1 repository · arXiv:1909.04054
-
Span Selection Pre-training for Question Answering 9 Sep 2019 · 1 repository · arXiv:1909.04120Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
KG-BERT: BERT for Knowledge Graph Completion 7 Sep 2019 · 3 repositories · arXiv:1909.03193Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
On Extractive and Abstractive Neural Document Summarization with Transformer Language Models 7 Sep 2019 · 1 repository · arXiv:1909.03186Syntology 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 6 harvested samples)
-
Enhancing Machine Translation with Dependency-Aware Self-Attention 6 Sep 2019 · 1 repository · arXiv:1909.03149
-
Supervised Multimodal Bitransformers for Classifying Images and Text 6 Sep 2019 · 6 repositories · arXiv:1909.02950
-
A Stack-Propagation Framework with Token-Level Intent Detection for Spoken Language Understanding 5 Sep 2019 · 2 repositories · arXiv:1909.02188
-
Accelerating Transformer Decoding via a Hybrid of Self-attention and Recurrent Neural Network 5 Sep 2019 · 0 repositories · arXiv:1909.02279
-
Effective Use of Transformer Networks for Entity Tracking 5 Sep 2019 · 1 repository · arXiv:1909.02635Syntology official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Source Dependency-Aware Transformer with Supervised Self-Attention 5 Sep 2019 · 0 repositories · arXiv:1909.02273
-
Jointly Learning to Align and Translate with Transformer Models 4 Sep 2019 · 1 repository · arXiv:1909.02074
-
Mogrifier LSTM 4 Sep 2019 · 3 repositories · arXiv:1909.01792
-
Encode, Tag, Realize: High-Precision Text Editing 3 Sep 2019 · 5 repositories · arXiv:1909.01187Syntology official (archive's flag): 5 ran · 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 1 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 11 harvested samples) · 1 pointer-only (licence)
-
The Bottom-up Evolution of Representations in the Transformer: A Study with Machine Translation and Language Modeling Objectives 3 Sep 2019 · 0 repositories · arXiv:1909.01380
-
Logic and the 2-Simplicial Transformer 2 Sep 2019 · 0 repositories · arXiv:1909.00668
-
Dependency-Based Relative Positional Encoding for Transformer NMT 1 Sep 2019 · 0 repositories
-
Dependency-Based Self-Attention for Transformer NMT 1 Sep 2019 · 0 repositories
-
Global Entity Disambiguation with BERT 1 Sep 2019 · 1 repository · arXiv:1909.00426
-
Turkish Tweet Classification with Transformer Encoder 1 Sep 2019 · 0 repositories
-
Humor Detection: A Transformer Gets the Last Laugh 31 Aug 2019 · 2 repositories · arXiv:1909.00252Syntology official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 2 pointer-only (licence)
-
Improving Multi-Head Attention with Capsule Networks 31 Aug 2019 · 0 repositories · arXiv:1909.00188
-
Knowledge Enhanced Attention for Robust Natural Language Inference 31 Aug 2019 · 0 repositories · arXiv:1909.00102
-
Modeling Graph Structure in Transformer for Better AMR-to-Text Generation 31 Aug 2019 · 1 repository · arXiv:1909.00136
-
Adaptively Sparse Transformers 30 Aug 2019 · 3 repositories · arXiv:1909.00015Syntology official (archive's flag): 2 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Answering Conversational Questions on Structured Data without Logical Forms 30 Aug 2019 · 0 repositories · arXiv:1908.11787
-
Pre-training A Neural Language Model Improves The Sample Efficiency of an Emergency Room Classification Model 30 Aug 2019 · 0 repositories · arXiv:1909.01136
-
Transformer Dissection: A Unified Understanding of Transformer's Attention via the Lens of Kernel 30 Aug 2019 · 1 repository · arXiv:1908.11775
-
Improving Deep Transformer with Depth-Scaled Initialization and Merged Attention 29 Aug 2019 · 1 repository · arXiv:1908.11365
-
Probing Representations Learned by Multimodal Recurrent and Transformer Models 29 Aug 2019 · 0 repositories · arXiv:1908.11125
-
Regularized Context Gates on Transformer for Machine Translation 29 Aug 2019 · 0 repositories · arXiv:1908.11020