Methods › General › Feedforward Networks › Position-Wise Feed-Forward Layer › Papers, page 132
Position-Wise Feed-Forward Layer
Papers archive 2025-07-28
archive papers tagged: 13,895 · with a code link: 6,514 · where Syntology ran a sample: 2,229 (1,902 with a run with no instrument failure, 327 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,229 of 13,895 tagged: 1,902 with a run with no instrument failure, 327 where every run was a failure of Syntology's instrument)
Page 132 of 139: papers 13,101 to 13,200 of 13,895, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
On Optimal Transformer Depth for Low-Resource Language Translation 9 Apr 2020 · 1 repository · arXiv:2004.04418
-
Sequential View Synthesis with Transformer 9 Apr 2020 · 0 repositories · arXiv:2004.04548
-
Adaptive Transformers in RL 8 Apr 2020 · 1 repository · arXiv:2004.03761
-
Diverse, Controllable, and Keyphrase-Aware: A Corpus and Method for News Multi-Headline Generation 8 Apr 2020 · 1 repository · arXiv:2004.03875
-
On the Effect of Dropping Layers of Pre-trained Transformer Models 8 Apr 2020 · 4 repositories · arXiv:2004.03844Syntology official (archive's flag): 5 ran · 7 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 1 pointer-only (licence)
-
Byte Pair Encoding is Suboptimal for Language Model Pretraining 7 Apr 2020 · 1 repository · arXiv:2004.03720
-
Probabilistic Spatial Transformer Networks 7 Apr 2020 · 1 repository · arXiv:2004.03637
-
Transformers to Learn Hierarchical Contexts in Multiparty Dialogue for Span-based Question Answering 7 Apr 2020 · 1 repository · arXiv:2004.03561
-
A Systematic Analysis of Morphological Content in BERT Models for Multiple Languages 6 Apr 2020 · 1 repository · arXiv:2004.03032
-
AutoToon: Automatic Geometric Warping for Face Cartoon Generation 6 Apr 2020 · 1 repository · arXiv:2004.02377
-
Syntax-driven Iterative Expansion Language Models for Controllable Text Generation 5 Apr 2020 · 0 repositories · arXiv:2004.02211
-
Conversational Question Reformulation via Sequence-to-Sequence Architectures and Pretrained Language Models 4 Apr 2020 · 0 repositories · arXiv:2004.01909
-
A Hierarchical Network for Abstractive Meeting Summarization with Cross-Domain Pretraining 4 Apr 2020 · 3 repositories · arXiv:2004.02016
-
Pre-training for Abstractive Document Summarization by Reinstating Source Text 4 Apr 2020 · 0 repositories · arXiv:2004.01853
-
LiDAR-based Online 3D Video Object Detection with Graph-based Message Passing and Spatiotemporal Transformer Attention 3 Apr 2020 · 1 repository · arXiv:2004.01389
-
Testing pre-trained Transformer models for Lithuanian news clustering 3 Apr 2020 · 0 repositories · arXiv:2004.03461
-
DSTC8-AVSD: Multimodal Semantic Transformer Network with Retrieval Style Word Generator 1 Apr 2020 · 0 repositories · arXiv:2004.08299
-
Better Sign Language Translation with STMC-Transformer 1 Apr 2020 · 1 repository · arXiv:2004.00588Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 1 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
A Swiss German Dictionary: Variation in Speech and Writing 31 Mar 2020 · 0 repositories · arXiv:2004.00139
-
DeepSumm -- Deep Code Summaries using Neural Transformer Architecture 31 Mar 2020 · 0 repositories · arXiv:2004.00998
-
Graph Enhanced Representation Learning for News Recommendation 31 Mar 2020 · 0 repositories · arXiv:2003.14292
-
X-Linear Attention Networks for Image Captioning 31 Mar 2020 · 2 repositories · arXiv:2003.14080
-
A Hierarchical Transformer for Unsupervised Parsing 30 Mar 2020 · 0 repositories · arXiv:2003.13841
-
AriEL: volume coding for sentence generation 30 Mar 2020 · 0 repositories · arXiv:2003.13600
-
Learning Contextualized Sentence Representations for Document-Level Neural Machine Translation 30 Mar 2020 · 0 repositories · arXiv:2003.13205
-
Sign Language Transformers: Joint End-to-end Sign Language Recognition and Translation 30 Mar 2020 · 2 repositories · arXiv:2003.13830Syntology official (archive's flag): 5 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
Abstractive Text Summarization based on Language Model Conditioning and Locality Modeling 29 Mar 2020 · 1 repository · arXiv:2003.13027Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Recursive Non-Autoregressive Graph-to-Graph Transformer for Dependency Parsing with Iterative Refinement 29 Mar 2020 · 1 repository · arXiv:2003.13118
-
Actor-Transformers for Group Activity Recognition 28 Mar 2020 · 0 repositories · arXiv:2003.12737
-
Variational Transformers for Diverse Response Generation 28 Mar 2020 · 2 repositories · arXiv:2003.12738Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
StrokeCoder: Path-Based Image Generation from Single Examples using Transformers 26 Mar 2020 · 0 repositories · arXiv:2003.11958
-
TLDR: Token Loss Dynamic Reweighting for Reducing Repetitive Utterance Generation 26 Mar 2020 · 1 repository · arXiv:2003.11963Syntology official (archive's flag): 15 ran · 15 ran (of which 0 constructed an object rather than computing a result; 14 with no instrument failure: 0 honoured, 0 violated, 14 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 15 harvested samples) · 15 pointer-only (licence)
-
Generalizing Spatial Transformers to Projective Geometry with Applications to 2D/3D Registration 24 Mar 2020 · 1 repository · arXiv:2003.10987
-
Probing Word Translations in the Transformer and Trading Decoder for Encoder Layers 21 Mar 2020 · 0 repositories · arXiv:2003.09586
-
TNT-KID: Transformer-based Neural Tagger for Keyword Identification 20 Mar 2020 · 1 repository · arXiv:2003.09166
-
Detecting Lane and Road Markings at A Distance with Perspective Transformer Layers 19 Mar 2020 · 0 repositories · arXiv:2003.08550
-
Normalized and Geometry-Aware Self-Attention Network for Image Captioning 19 Mar 2020 · 0 repositories · arXiv:2003.08897
-
Temporal Embeddings and Transformer Models for Narrative Text Understanding 19 Mar 2020 · 0 repositories · arXiv:2003.08811
-
Scene Text Recognition via Transformer 18 Mar 2020 · 0 repositories · arXiv:2003.08077
-
Transformer Networks for Trajectory Forecasting 18 Mar 2020 · 1 repository · arXiv:2003.08111Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
Human Activity Recognition from Wearable Sensor Data Using Self-Attention 17 Mar 2020 · 2 repositories · arXiv:2003.09018
-
Multi-modal Dense Video Captioning 17 Mar 2020 · 4 repositories · arXiv:2003.07758
-
PowerNorm: Rethinking Batch Normalization in Transformers 17 Mar 2020 · 1 repository · arXiv:2003.07845
-
TRANS-BLSTM: Transformer with Bidirectional LSTM for Language Understanding 16 Mar 2020 · 0 repositories · arXiv:2003.07000
-
Document Ranking with a Pretrained Sequence-to-Sequence Model 14 Mar 2020 · 2 repositories · arXiv:2003.06713
-
Learning to Encode Position for Transformer with Continuous Dynamical Model 13 Mar 2020 · 1 repository · arXiv:2003.09229Syntology 3 ran (of which 2 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Keyword-Attentive Deep Semantic Matching 11 Mar 2020 · 1 repository · arXiv:2003.11516
-
Hybrid Attention-Based Transformer Block Model for Distant Supervision Relation Extraction 10 Mar 2020 · 0 repositories · arXiv:2003.11518
-
ReZero is All You Need: Fast Convergence at Large Depth 10 Mar 2020 · 13 repositories · arXiv:2003.04887Syntology community repositories only · 6 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 6 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
Capacity of Continuous Channels with Memory via Directed Information Neural Estimator 9 Mar 2020 · 1 repository · arXiv:2003.04179Syntology 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Cross-modal Learning for Multi-modal Video Categorization 7 Mar 2020 · 0 repositories · arXiv:2003.03501
-
TTPP: Temporal Transformer with Progressive Prediction for Efficient Action Anticipation 7 Mar 2020 · 0 repositories · arXiv:2003.03530
-
Teaching Temporal Logics to Neural Networks 6 Mar 2020 · 2 repositories · arXiv:2003.04218Syntology official (archive's flag): 5 ran · 14 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 2 honoured, 0 violated, 9 with no contract checked; 3 where Syntology's instrument failed) · 10 unverified (of 24 harvested samples)
-
State-of-the-Art Augmented NLP Transformer models for direct and single-step retrosynthesis 5 Mar 2020 · 1 repository · arXiv:2003.02804
-
EmpTransfo: A Multi-head Transformer Architecture for Creating Empathetic Dialog Systems 5 Mar 2020 · 1 repository · arXiv:2003.02958
-
AlignTTS: Efficient Feed-Forward Text-to-Speech System without Explicit Alignment 4 Mar 2020 · 2 repositories · arXiv:2003.01950Syntology 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples) · 3 pointer-only (licence)
-
Data Augmentation using Pre-trained Transformer Models 4 Mar 2020 · 4 repositories · arXiv:2003.02245Syntology community repositories only · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples)
-
Controllable Time-Delay Transformer for Real-Time Punctuation Prediction and Disfluency Detection 3 Mar 2020 · 0 repositories · arXiv:2003.01309
-
Heterogeneous Graph Transformer 3 Mar 2020 · 4 repositories · arXiv:2003.01332Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 6 harvested samples) · 3 pointer-only (licence)
-
Meta-Embeddings Based On Self-Attention 3 Mar 2020 · 0 repositories · arXiv:2003.01371
-
Transfer Learning for Context-Aware Spoken Language Understanding 3 Mar 2020 · 0 repositories · arXiv:2003.01305
-
Transformer++ 2 Mar 2020 · 0 repositories · arXiv:2003.04974
-
Unblind Your Apps: Predicting Natural-Language Labels for Mobile GUI Components by Deep Learning 1 Mar 2020 · 1 repository · arXiv:2003.00380Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Automatic Perturbation Analysis for Scalable Certified Robustness and Beyond 28 Feb 2020 · 7 repositories · arXiv:2002.12920Syntology community repositories only · 14 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 2 violated, 11 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 18 harvested samples) · 2 pointer-only (licence)
-
Exploring and Distilling Cross-Modal Information for Image Captioning 28 Feb 2020 · 0 repositories · arXiv:2002.12585
-
Compressing Large-Scale Transformer-Based Models: A Case Study on BERT 27 Feb 2020 · 0 repositories · arXiv:2002.11985
-
Marathi To English Neural Machine Translation With Near Perfect Corpus And Transformers 26 Feb 2020 · 1 repository · arXiv:2002.11643
-
Multi-task Learning with Multi-head Attention for Multi-choice Reading Comprehension 26 Feb 2020 · 0 repositories · arXiv:2003.04992
-
Sparse Sinkhorn Attention 26 Feb 2020 · 1 repository · arXiv:2002.11296
-
Train Large, Then Compress: Rethinking Model Size for Efficient Training and Inference of Transformers 26 Feb 2020 · 2 repositories · arXiv:2002.11794
-
MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers 25 Feb 2020 · 1 repository · arXiv:2002.10957
-
Exploring BERT Parameter Efficiency on the Stanford Question Answering Dataset v2.0 25 Feb 2020 · 0 repositories · arXiv:2002.10670
-
Fixed Encoder Self-Attention Patterns in Transformer-Based Machine Translation 24 Feb 2020 · 0 repositories · arXiv:2002.10260
-
GRET: Global Representation Enhanced Transformer 24 Feb 2020 · 0 repositories · arXiv:2002.10101
-
Sketchformer: Transformer-based Representation for Sketched Structure 24 Feb 2020 · 1 repository · arXiv:2002.10381Syntology 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Addressing Some Limitations of Transformers with Feedback Memory 21 Feb 2020 · 4 repositories · arXiv:2002.09402Syntology official: no sample here; runs from other or unrecorded repositories · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 2 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 2 pointer-only (licence)
-
Learning Dynamic Belief Graphs to Generalize on Text-Based Games 21 Feb 2020 · 1 repository · arXiv:2002.09127
-
Transformer Hawkes Process 21 Feb 2020 · 3 repositories · arXiv:2002.09291Syntology official (archive's flag): 4 ran · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
LAMBERT: Layout-Aware (Language) Modeling for information extraction 19 Feb 2020 · 1 repository · arXiv:2002.08087
-
Molecule Attention Transformer 19 Feb 2020 · 7 repositories · arXiv:2002.08264Syntology official (archive's flag): 2 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 2 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples) · 3 pointer-only (licence)
-
Towards Making the Most of Context in Neural Machine Translation 19 Feb 2020 · 1 repository · arXiv:2002.07982
-
Tree-structured Attention with Hierarchical Accumulation 19 Feb 2020 · 0 repositories · arXiv:2002.08046
-
Conditional Self-Attention for Query-based Summarization 18 Feb 2020 · 0 repositories · arXiv:2002.07338
-
Gradient-Based Adversarial Training on Transformer Networks for Detecting Check-Worthy Factual Claims 18 Feb 2020 · 1 repository · arXiv:2002.07725
-
Hierarchical Transformer Network for Utterance-level Emotion Recognition 18 Feb 2020 · 0 repositories · arXiv:2002.07551
-
Sequential Latent Knowledge Selection for Knowledge-Grounded Dialogue 18 Feb 2020 · 3 repositories · arXiv:2002.07510Syntology official: not harvested · 0 ran · 5 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Uncertainty Estimation in Autoregressive Structured Prediction 18 Feb 2020 · 1 repository · arXiv:2002.07650
-
A Financial Service Chatbot based on Deep Bidirectional Transformers 17 Feb 2020 · 0 repositories · arXiv:2003.04987
-
Controlling Computation versus Quality for Neural Sequence Models 17 Feb 2020 · 0 repositories · arXiv:2002.07106
-
Low-Rank Bottleneck in Multi-head Attention Models 17 Feb 2020 · 0 repositories · arXiv:2002.07028
-
Multi-layer Representation Fusion for Neural Machine Translation 16 Feb 2020 · 1 repository · arXiv:2002.06714Syntology 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Neural Machine Translation with Joint Representation 16 Feb 2020 · 1 repository · arXiv:2002.06546
-
Small energy masking for improved neural network training for end-to-end speech recognition 15 Feb 2020 · 0 repositories · arXiv:2002.06312
-
UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation 15 Feb 2020 · 2 repositories · arXiv:2002.06353Syntology official: harvested, nothing ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
Deep Attentive Study Session Dropout Prediction in Mobile Learning Environment 14 Feb 2020 · 0 repositories · arXiv:2002.11624
-
Stress Test Evaluation of Transformer-based Models in Natural Language Understanding Tasks 14 Feb 2020 · 0 repositories · arXiv:2002.06261
-
Towards an Appropriate Query, Key, and Value Computation for Knowledge Tracing 14 Feb 2020 · 5 repositories · arXiv:2002.07033Syntology 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Transformer on a Diet 14 Feb 2020 · 1 repository · arXiv:2002.06170
-
Sparse and Structured Visual Attention 13 Feb 2020 · 1 repository · arXiv:2002.05556
-
Attentional Speech Recognition Models Misbehave on Out-of-domain Utterances 12 Feb 2020 · 1 repository · arXiv:2002.05150