Methods › General › Feedforward Networks › Position-Wise Feed-Forward Layer › Papers, page 133
Position-Wise Feed-Forward Layer
Papers archive 2025-07-28
archive papers tagged: 13,895 · with a code link: 6,514 · where Syntology ran a sample: 2,229 (1,902 with a run with no instrument failure, 327 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,229 of 13,895 tagged: 1,902 with a run with no instrument failure, 327 where every run was a failure of Syntology's instrument)
Page 133 of 139: papers 13,201 to 13,300 of 13,895, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
End-to-End Face Parsing via Interlinked Convolutional Neural Networks 12 Feb 2020 · 1 repository · arXiv:2002.04831
-
GLU Variants Improve Transformer 12 Feb 2020 · 27 repositories · arXiv:2002.05202Syntology 9 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 3 honoured, 4 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples) · 7 pointer-only (licence)
-
On Layer Normalization in the Transformer Architecture 12 Feb 2020 · 9 repositories · arXiv:2002.04745Syntology 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 2 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 11 harvested samples)
-
Superbloom: Bloom filter meets Transformer 11 Feb 2020 · 0 repositories · arXiv:2002.04723
-
Training with Streaming Annotation 11 Feb 2020 · 0 repositories · arXiv:2002.04165
-
Deep Representation Learning for Dynamical Systems Modeling 10 Feb 2020 · 0 repositories · arXiv:2002.05111
-
End-to-End Multi-speaker Speech Recognition with Transformer 10 Feb 2020 · 0 repositories · arXiv:2002.03921
-
Pre-training Tasks for Embedding-based Large-scale Retrieval 10 Feb 2020 · 0 repositories · arXiv:2002.03932
-
StickyPillars: Robust and Efficient Feature Matching on Point Clouds using Graph Neural Networks 10 Feb 2020 · 0 repositories · arXiv:2002.03983
-
On the distance between two neural networks and the stability of learning 9 Feb 2020 · 2 repositories · arXiv:2002.03432
-
Blank Language Models 8 Feb 2020 · 1 repository · arXiv:2002.03079Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 1 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 1 pointer-only (licence)
-
Multimodal Matching Transformer for Live Commenting 7 Feb 2020 · 0 repositories · arXiv:2002.02649
-
Transformer-Capsule Model for Intent Detection 7 Feb 2020 · 0 repositories
-
Transformer Transducer: A Streamable Speech Recognition Model with Transformer Encoders and RNN-T Loss 7 Feb 2020 · 5 repositories · arXiv:2002.02562Syntology 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 1 pointer-only (licence)
-
Few-Shot Learning as Domain Adaptation: Algorithm and Analysis 6 Feb 2020 · 0 repositories · arXiv:2002.02050
-
perm2vec: Graph Permutation Selection for Decoding of Error Correction Codes using Self-Attention 6 Feb 2020 · 0 repositories · arXiv:2002.02315
-
Aligning the Pretraining and Finetuning Objectives of Language Models 5 Feb 2020 · 0 repositories · arXiv:2002.02000
-
Vocoder-free End-to-End Voice Conversion with Transformer Network 5 Feb 2020 · 1 repository · arXiv:2002.03808
-
Interpretable & Time-Budget-Constrained Contextualization for Re-Ranking 4 Feb 2020 · 1 repository · arXiv:2002.01854
-
Learning Long- and Short-Term User Literal-Preference with Multimodal Hierarchical Transformer Network for Personalized Image Caption 4 Feb 2020 · 0 repositories
-
Multistage Model for Robust Face Alignment Using Deep Neural Networks 4 Feb 2020 · 0 repositories · arXiv:2002.01075
-
IART: Intent-aware Response Ranking with Transformers in Information-seeking Conversation Systems 3 Feb 2020 · 1 repository · arXiv:2002.00571
-
Bridging Text and Video: A Universal Multimodal Transformer for Video-Audio Scene-Aware Dialog 1 Feb 2020 · 1 repository · arXiv:2002.00163Syntology 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Pop Music Transformer: Beat-based Modeling and Generation of Expressive Pop Piano Compositions 1 Feb 2020 · 7 repositories · arXiv:2002.00212
-
Pretrained Transformers for Simple Question Answering over Knowledge Graphs 31 Jan 2020 · 1 repository · arXiv:2001.11985
-
Interpretable Rumor Detection in Microblogs by Attending to User Interactions 29 Jan 2020 · 1 repository · arXiv:2001.10667
-
Applying Recent Innovations from NLP to MOOC Student Course Trajectory Modeling 23 Jan 2020 · 0 repositories · arXiv:2001.08333
-
Multi-level Head-wise Match and Aggregation in Transformer for Textual Sequence Matching 20 Jan 2020 · 0 repositories · arXiv:2001.07234
-
Recommending Themes for Ad Creative Design via Visual-Linguistic Representations 20 Jan 2020 · 1 repository · arXiv:2001.07194
-
A multimodal deep learning approach for named entity recognition from social media 19 Jan 2020 · 0 repositories · arXiv:2001.06888
-
Deep Learning for Hindi Text Classification: A Comparison 19 Jan 2020 · 0 repositories · arXiv:2001.10340
-
Shifted and Squeezed 8-bit Floating Point format for Low-Precision Training of Deep Neural Networks 16 Jan 2020 · 0 repositories · arXiv:2001.05674
-
Insertion-Deletion Transformer 15 Jan 2020 · 0 repositories · arXiv:2001.05540
-
Non-Autoregressive Machine Translation with Disentangled Context Transformer 15 Jan 2020 · 1 repository · arXiv:2001.05136
-
Transformer-based Online CTC/attention End-to-End Speech Recognition Architecture 15 Jan 2020 · 0 repositories · arXiv:2001.08290
-
Auto Completion of User Interface Layout Design Using Transformer-Based Tree Decoders 14 Jan 2020 · 0 repositories · arXiv:2001.05308
-
The problems with using STNs to align CNN feature maps 14 Jan 2020 · 0 repositories · arXiv:2001.05858
-
Reformer: The Efficient Transformer 13 Jan 2020 · 10 repositories · arXiv:2001.04451Syntology community repositories only · 6 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 5 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples)
-
Urdu-English Machine Transliteration using Neural Networks 12 Jan 2020 · 0 repositories · arXiv:2001.05296
-
Spatial-Temporal Transformer Networks for Traffic Flow Forecasting 9 Jan 2020 · 1 repository · arXiv:2001.02908
-
Streaming automatic speech recognition with the transformer model 8 Jan 2020 · 0 repositories · arXiv:2001.02674
-
RECAST: Interactive Auditing of Automatic Toxicity Detection Models 7 Jan 2020 · 0 repositories · arXiv:2001.01819
-
FDFtNet: Facing Off Fake Images using Fake Detection Fine-tuning Network 5 Jan 2020 · 2 repositories · arXiv:2001.01265Syntology official (archive's flag): 4 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 13 harvested samples)
-
Learning Accurate Integer Transformer Machine-Translation Models 3 Jan 2020 · 0 repositories · arXiv:2001.00926
-
Two-Level Transformer and Auxiliary Coherence Modeling for Improved Text Segmentation 3 Jan 2020 · 1 repository · arXiv:2001.00891
-
Representing Unordered Data Using Complex-Weighted Multiset Automata 2 Jan 2020 · 0 repositories · arXiv:2001.00610
-
BERT-AL: BERT for Arbitrarily Long Document Understanding 1 Jan 2020 · 0 repositories
-
NEURAL EXECUTION ENGINES 1 Jan 2020 · 0 repositories
-
Poly-encoders: Architectures and Pre-training Strategies for Fast and Accurate Multi-sentence Scoring 1 Jan 2020 · 2 repositories
-
Deep Attentive Ranking Networks for Learning to Order Sentences 31 Dec 2019 · 0 repositories · arXiv:2001.00056
-
EEG based Continuous Speech Recognition using Transformers 31 Dec 2019 · 0 repositories · arXiv:2001.00501
-
AraNet: A Deep Learning Toolkit for Arabic Social Media 30 Dec 2019 · 1 repository · arXiv:1912.13072
-
All-in-One Image-Grounded Conversational Agents 28 Dec 2019 · 0 repositories · arXiv:1912.12394
-
Encoding word order in complex embeddings 27 Dec 2019 · 1 repository · arXiv:1912.12333
-
Is Attention All What You Need? -- An Empirical Investigation on Convolution-Based Active Memory and Self-Attention 27 Dec 2019 · 1 repository · arXiv:1912.11959
-
Explicit Sparse Transformer: Concentrated Attention Through Explicit Selection 25 Dec 2019 · 2 repositories · arXiv:1912.11637Syntology official: no sample here; runs from other or unrecorded repositories · 3 ran (of which 1 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Improving Abstractive Text Summarization with History Aggregation 24 Dec 2019 · 0 repositories · arXiv:1912.11046
-
Multi-Graph Transformer for Free-Hand Sketch Recognition 24 Dec 2019 · 1 repository · arXiv:1912.11258Syntology official: harvested, nothing ran · 0 ran · 4 unverified (of 4 harvested samples)
-
end-to-end training of a large vocabulary end-to-end speech recognition system 22 Dec 2019 · 0 repositories · arXiv:1912.11040
-
Learning and Evaluating Contextual Embedding of Source Code 21 Dec 2019 · 2 repositories · arXiv:2001.00059
-
Are Transformers universal approximators of sequence-to-sequence functions? 20 Dec 2019 · 0 repositories · arXiv:1912.10077
-
Axial Attention in Multidimensional Transformers 20 Dec 2019 · 3 repositories · arXiv:1912.12180Syntology 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 2 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
ET-USB: Transformer-Based Sequential Behavior Modeling for Inbound Customer Service 20 Dec 2019 · 0 repositories · arXiv:1912.10852
-
Shareable Representations for Search Query Understanding 20 Dec 2019 · 0 repositories · arXiv:2001.04345
-
Temporal Fusion Transformers for Interpretable Multi-horizon Time Series Forecasting 19 Dec 2019 · 36 repositories · arXiv:1912.09363Syntology 8 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 4 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples) · 5 pointer-only (licence)
-
Meshed-Memory Transformer for Image Captioning 17 Dec 2019 · 2 repositories · arXiv:1912.08226Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 0 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples) · 3 pointer-only (licence)
-
BERTQA -- Attention on Steroids 14 Dec 2019 · 0 repositories · arXiv:1912.10435
-
Voice Transformer Network: Sequence-to-Sequence Voice Conversion Using Transformer with Text-to-Speech Pretraining 14 Dec 2019 · 2 repositories · arXiv:1912.06813
-
WaLDORf: Wasteless Language-model Distillation On Reading-comprehension 13 Dec 2019 · 0 repositories · arXiv:1912.06638
-
Encoding Musical Style with Transformer Autoencoders 10 Dec 2019 · 0 repositories · arXiv:1912.05537
-
Learning a Layout Transfer Network for Context Aware Object Detection 9 Dec 2019 · 0 repositories · arXiv:1912.03865
-
Transformer Based Reinforcement Learning For Games 9 Dec 2019 · 0 repositories · arXiv:1912.03918
-
Bidirectional Scene Text Recognition with a Single Decoder 8 Dec 2019 · 1 repository · arXiv:1912.03656
-
Personalized Patent Claim Generation and Measurement 7 Dec 2019 · 0 repositories · arXiv:1912.03502
-
Synchronous Transformers for End-to-End Speech Recognition 6 Dec 2019 · 0 repositories · arXiv:1912.02958
-
Weak Supervision helps Emergence of Word-Object Alignment and improves Vision-Language Tasks 6 Dec 2019 · 0 repositories · arXiv:1912.03063
-
Self-Supervised Contextual Language Representation of Radiology Reports to Improve the Identification of Communication Urgency 5 Dec 2019 · 0 repositories · arXiv:1912.02703
-
AMUSED: A Multi-Stream Vector Representation Method for Use in Natural Dialogue 4 Dec 2019 · 0 repositories · arXiv:1912.10160
-
TU Wien @ TREC Deep Learning '19 -- Simple Contextualization for Re-ranking 3 Dec 2019 · 1 repository · arXiv:1912.01385
-
Audiovisual Transformer Architectures for Large-Scale Classification and Synchronization of Weakly Labeled Audio Events 2 Dec 2019 · 0 repositories · arXiv:1912.02615
-
BLiMP: The Benchmark of Linguistic Minimal Pairs for English 2 Dec 2019 · 4 repositories · arXiv:1912.00582Syntology community repositories only · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Long Distance Relationships without Time Travel: Boosting the Performance of a Sparse Predictive Autoencoder in Sequence Modeling 2 Dec 2019 · 1 repository · arXiv:1912.01116
-
Multi-Scale Self-Attention for Text Classification 2 Dec 2019 · 0 repositories · arXiv:1912.00544
-
Neural Academic Paper Generation 2 Dec 2019 · 1 repository · arXiv:1912.01982
-
Solving Arithmetic Word Problems Automatically Using Transformer and Unambiguous Representations 2 Dec 2019 · 1 repository · arXiv:1912.00871
-
Hybrid 8-bit Floating Point (HFP8) Training and Inference for Deep Neural Networks 1 Dec 2019 · 0 repositories
-
Minimum Bayes Risk Training of RNN-Transducer for End-to-End Speech Recognition 28 Nov 2019 · 0 repositories · arXiv:1911.12487
-
DeFINE: DEep Factorized INput Token Embeddings for Neural Sequence Modeling 27 Nov 2019 · 1 repository · arXiv:1911.12385
-
Do Attention Heads in BERT Track Syntactic Dependencies? 27 Nov 2019 · 1 repository · arXiv:1911.12246
-
SimpleBooks: Long-term dependency book dataset with simplified English vocabulary for word-level language modeling 27 Nov 2019 · 0 repositories · arXiv:1911.12391
-
Taking a Stance on Fake News: Towards Automatic Disinformation Assessment via Deep Bidirectional Transformer Language Models for Stance Detection 27 Nov 2019 · 0 repositories · arXiv:1911.11951
-
Autoencoding Undirected Molecular Graphs With Neural Networks 26 Nov 2019 · 1 repository · arXiv:2001.03517
-
Efficient Attention Mechanism for Visual Dialog that can Handle All the Interactions between Multiple Inputs 26 Nov 2019 · 1 repository · arXiv:1911.11390
-
Low Rank Factorization for Compact Multi-Head Self-Attention 26 Nov 2019 · 1 repository · arXiv:1912.00835
-
Password-conditioned Anonymization and Deanonymization with Face Identity Transformers 26 Nov 2019 · 1 repository · arXiv:1911.11759
-
Relevance-Promoting Language Model for Short-Text Conversation 26 Nov 2019 · 0 repositories · arXiv:1911.11489
-
Learning to Reuse Translations: Guiding Neural Machine Translation with Examples 25 Nov 2019 · 0 repositories · arXiv:1911.10732
-
Who did They Respond to? Conversation Structure Modeling using Masked Hierarchical Transformer 25 Nov 2019 · 1 repository · arXiv:1911.10666
-
Factorized Multimodal Transformer for Multimodal Sequential Learning 22 Nov 2019 · 0 repositories · arXiv:1911.09826
-
Improving N-gram Language Models with Pre-trained Deep Transformer 22 Nov 2019 · 0 repositories · arXiv:1911.10235