Methods › General › Feedforward Networks › Position-Wise Feed-Forward Layer › Papers, page 134
Position-Wise Feed-Forward Layer
Papers archive 2025-07-28
archive papers tagged: 13,895 · with a code link: 6,514 · where Syntology ran a sample: 2,229 (1,902 with a run with no instrument failure, 327 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,229 of 13,895 tagged: 1,902 with a run with no instrument failure, 327 where every run was a failure of Syntology's instrument)
Page 134 of 139: papers 13,301 to 13,400 of 13,895, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Neuron Interaction Based Representation Composition for Neural Machine Translation 22 Nov 2019 · 0 repositories · arXiv:1911.09877
-
Spectral Graph Transformer Networks for Brain Surface Parcellation 22 Nov 2019 · 0 repositories · arXiv:1911.10118
-
WildMix Dataset and Spectro-Temporal Transformer Model for Monoaural Audio Source Separation 21 Nov 2019 · 0 repositories · arXiv:1911.09783
-
MarioNETte: Few-shot Face Reenactment Preserving Identity of Unseen Targets 19 Nov 2019 · 0 repositories · arXiv:1911.08139
-
Graph Transformer for Graph-to-Sequence Learning 18 Nov 2019 · 1 repository · arXiv:1911.07470Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples)
-
MUSE: Parallel Multi-Scale Attention for Sequence to Sequence Learning 17 Nov 2019 · 3 repositories · arXiv:1911.09483Syntology official: not harvested · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Music theme recognition using CNN and self-attention 16 Nov 2019 · 0 repositories · arXiv:1911.07041
-
Evaluating robustness of language models for chief complaint extraction from patient-generated text 15 Nov 2019 · 0 repositories · arXiv:1911.06915
-
Selection-based Question Answering of an MOOC 15 Nov 2019 · 1 repository · arXiv:1911.07629
-
Sequential Recommendation with Relation-Aware Kernelized Self-Attention 15 Nov 2019 · 0 repositories · arXiv:1911.06478
-
Attention on Abstract Visual Reasoning 14 Nov 2019 · 0 repositories · arXiv:1911.05990
-
Iterative Answer Prediction with Pointer-Augmented Multimodal Transformers for TextVQA 14 Nov 2019 · 1 repository · arXiv:1911.06258
-
Character-based NMT with Transformer 12 Nov 2019 · 0 repositories · arXiv:1911.04997
-
SMILES Transformer: Pre-trained Molecular Fingerprint for Low Data Drug Discovery 12 Nov 2019 · 1 repository · arXiv:1911.04738Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Attending to Entities for Better Text Understanding 11 Nov 2019 · 0 repositories · arXiv:1911.04361
-
BP-Transformer: Modelling Long-Range Context via Binary Partitioning 11 Nov 2019 · 2 repositories · arXiv:1911.04070Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 7 harvested samples)
-
Disentangle, align and fuse for multimodal and semi-supervised image segmentation 11 Nov 2019 · 2 repositories · arXiv:1911.04417
-
Long-span language modeling for speech recognition 11 Nov 2019 · 0 repositories · arXiv:1911.04571
-
TANDA: Transfer and Adapt Pre-Trained Transformer Models for Answer Sentence Selection 11 Nov 2019 · 2 repositories · arXiv:1911.04118
-
Distilling Knowledge Learned in BERT for Text Generation 10 Nov 2019 · 2 repositories · arXiv:1911.03829Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Learning to Few-Shot Learn Across Diverse Natural Language Classification Tasks 10 Nov 2019 · 2 repositories · arXiv:1911.03863
-
Listen and Fill in the Missing Letters: Non-Autoregressive Transformer for Speech Recognition 10 Nov 2019 · 0 repositories · arXiv:1911.04908
-
Syntax-Infused Transformer and BERT models for Machine Translation and Natural Language Understanding 10 Nov 2019 · 0 repositories · arXiv:1911.06156
-
TENER: Adapting Transformer Encoder for Named Entity Recognition 10 Nov 2019 · 6 repositories · arXiv:1911.04474Syntology community repositories only · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Two-Headed Monster And Crossed Co-Attention Networks 10 Nov 2019 · 0 repositories · arXiv:1911.03897
-
A Reinforced Generation of Adversarial Examples for Neural Machine Translation 9 Nov 2019 · 1 repository · arXiv:1911.03677
-
Graph-to-Graph Transformer for Transition-based Dependency Parsing 8 Nov 2019 · 1 repository · arXiv:1911.03561
-
Question Generation from Paragraphs: A Tale of Two Hierarchical Models 8 Nov 2019 · 0 repositories · arXiv:1911.03407
-
Resurrecting Submodularity for Neural Text Generation 8 Nov 2019 · 0 repositories · arXiv:1911.03014
-
Towards Hierarchical Importance Attribution: Explaining Compositional Semantics for Neural Sequence Models 8 Nov 2019 · 3 repositories · arXiv:1911.06194
-
Lipschitz Constrained Parameter Initialization for Deep Transformers 8 Nov 2019 · 0 repositories · arXiv:1911.03179
-
Microsoft Research Asia's Systems for WMT19 7 Nov 2019 · 0 repositories · arXiv:1911.06191
-
Porous Lattice-based Transformer Encoder for Chinese NER 7 Nov 2019 · 0 repositories · arXiv:1911.02733
-
Probing Contextualized Sentence Representations with Visual Awareness 7 Nov 2019 · 0 repositories · arXiv:1911.02971
-
An End-to-end Approach for Lexical Stress Detection based on Transformer 6 Nov 2019 · 0 repositories · arXiv:1911.04862
-
CoKE: Contextualized Knowledge Graph Embedding 6 Nov 2019 · 3 repositories · arXiv:1911.02168Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Enriching Conversation Context in Retrieval-based Chatbots 6 Nov 2019 · 0 repositories · arXiv:1911.02290
-
Fast Transformer Decoding: One Write-Head is All You Need 6 Nov 2019 · 4 repositories · arXiv:1911.02150Syntology 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Graph Transformer Networks 6 Nov 2019 · 1 repository · arXiv:1911.06455
-
Learning to Answer by Learning to Ask: Getting the Best of GPT-2 and BERT Worlds 6 Nov 2019 · 0 repositories · arXiv:1911.02365
-
Improving Bidirectional Decoding with Dynamic Target Semantics in Neural Machine Translation 5 Nov 2019 · 0 repositories · arXiv:1911.01597
-
An Algorithm for Routing Capsules in All Domains 2 Nov 2019 · 1 repository · arXiv:1911.00792
-
Machine Translation Evaluation using Bi-directional Entailment 2 Nov 2019 · 0 repositories · arXiv:1911.00681
-
Aggregating Bidirectional Encoder Representations Using MatchLSTM for Sequence Matching 1 Nov 2019 · 0 repositories
-
Automatically Extracting Challenge Sets for Non-Local Phenomena in Neural Machine Translation 1 Nov 2019 · 0 repositories
-
Combining Global Sparse Gradients with Local Gradients in Distributed Neural Network Training 1 Nov 2019 · 0 repositories
-
CVIT's submissions to WAT-2019 1 Nov 2019 · 0 repositories
-
Dialect Text Normalization to Normative Standard Finnish 1 Nov 2019 · 1 repository
-
DialoGPT: Large-Scale Generative Pre-training for Conversational Response Generation 1 Nov 2019 · 6 repositories · arXiv:1911.00536Syntology community repositories only · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 6 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
English to Hindi Multi-modal Neural Machine Translation and Hindi Image Captioning 1 Nov 2019 · 0 repositories
-
Enhanced Transformer Model for Data-to-Text Generation 1 Nov 2019 · 0 repositories
-
Idiap NMT System for WAT 2019 Multimodal Translation Task 1 Nov 2019 · 0 repositories
-
Improving Answer Selection and Answer Triggering using Hard Negatives 1 Nov 2019 · 0 repositories
-
Improving Generalization of Transformer for Speech Recognition with Parallel Schedule Sampling and Relative Positional Embedding 1 Nov 2019 · 0 repositories · arXiv:1911.00203
-
Improving Natural Language Understanding by Reverse Mapping Bytepair Encoding 1 Nov 2019 · 0 repositories
-
Inspecting Unification of Encoding and Matching with Transformer: A Case Study of Machine Reading Comprehension 1 Nov 2019 · 0 repositories
-
Long Warm-up and Self-Training: Training Strategies of NICT-2 NMT System at WAT-2019 1 Nov 2019 · 0 repositories
-
LTRC-MT Simple & Effective Hindi-English Neural Machine Translation Systems at WAT 2019 1 Nov 2019 · 0 repositories
-
Mixed Multi-Head Self-Attention for Neural Machine Translation 1 Nov 2019 · 0 repositories
-
On the Relation between Position Information and Sentence Length in Neural Machine Translation 1 Nov 2019 · 0 repositories
-
Our Neural Machine Translation Systems for WAT 2019 1 Nov 2019 · 0 repositories
-
Recurrent Positional Embedding for Neural Machine Translation 1 Nov 2019 · 0 repositories
-
Recycling a Pre-trained BERT Encoder for Neural Machine Translation 1 Nov 2019 · 0 repositories
-
Sarah's Participation in WAT 2019 1 Nov 2019 · 0 repositories
-
Self-Adaptive Scaling for Learnable Residual Structure 1 Nov 2019 · 0 repositories
-
Supervised neural machine translation based on data augmentation and improved training & inference process 1 Nov 2019 · 0 repositories
-
SYSTRAN @ WAT 2019: Russian-Japanese News Commentary task 1 Nov 2019 · 0 repositories
-
SYSTRAN @ WNGT 2019: DGT Task 1 Nov 2019 · 0 repositories
-
The Concordia NLG Surface Realizer at SRST 2019 1 Nov 2019 · 0 repositories
-
Transformer and seq2seq model for Paraphrase Generation 1 Nov 2019 · 0 repositories
-
Transformer-based Model for Single Documents Neural Summarization 1 Nov 2019 · 0 repositories
-
Transformer Dissection: An Unified Understanding for Transformer's Attention via the Lens of Kernel 1 Nov 2019 · 0 repositories
-
``Transforming'' Delete, Retrieve, Generate Approach for Controlled Text Style Transfer 1 Nov 2019 · 0 repositories
-
Attention Is All You Need for Chinese Word Segmentation 31 Oct 2019 · 1 repository · arXiv:1910.14537
-
Document-level Neural Machine Translation with Associated Memory Network 31 Oct 2019 · 0 repositories · arXiv:1910.14528
-
Image-Conditioned Graph Generation for Road Network Extraction 31 Oct 2019 · 3 repositories · arXiv:1910.14388
-
NAT: Neural Architecture Transformer for Accurate and Compact Architectures 31 Oct 2019 · 1 repository · arXiv:1910.14488
-
Neural Assistant: Joint Action Prediction, Response Generation, and Latent Knowledge Reasoning 31 Oct 2019 · 1 repository · arXiv:1910.14613
-
Parameter Sharing Decoder Pair for Auto Composing 31 Oct 2019 · 0 repositories · arXiv:1910.14270
-
Transfer Learning from Transformers to Fake News Challenge Stance Detection (FNC-1) Task 31 Oct 2019 · 0 repositories · arXiv:1910.14353
-
An Augmented Transformer Architecture for Natural Language Generation Tasks 30 Oct 2019 · 0 repositories · arXiv:1910.13634
-
Lightweight and Efficient End-to-End Speech Recognition Using Low-Rank Transformer 30 Oct 2019 · 0 repositories · arXiv:1910.13923
-
Transformer-based Cascaded Multimodal Speech Translation 29 Oct 2019 · 0 repositories · arXiv:1910.13215
-
An Empirical Study of Generation Order for Machine Translation 29 Oct 2019 · 0 repositories · arXiv:1910.13437
-
Big Bidirectional Insertion Representations for Documents 29 Oct 2019 · 0 repositories · arXiv:1910.13034
-
Transformer-Transducer: End-to-End Speech Recognition with Self-Attention 28 Oct 2019 · 1 repository · arXiv:1910.12977
-
An Adaptive and Momental Bound Method for Stochastic Learning 27 Oct 2019 · 2 repositories · arXiv:1910.12249Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Training ASR models by Generation of Contextual Information 27 Oct 2019 · 0 repositories · arXiv:1910.12367
-
HUBERT Untangles BERT to Improve Transfer across NLP Tasks 25 Oct 2019 · 1 repository · arXiv:1910.12647
-
Mockingjay: Unsupervised Speech Representation Learning with Deep Bidirectional Transformer Encoders 25 Oct 2019 · 7 repositories · arXiv:1910.12638Syntology community repositories only · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 0 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 11 harvested samples) · 1 pointer-only (licence)
-
Towards Online End-to-end Transformer Automatic Speech Recognition 25 Oct 2019 · 0 repositories · arXiv:1910.11871
-
An Empirical Study of Efficient ASR Rescoring with Transformers 24 Oct 2019 · 0 repositories · arXiv:1910.11450
-
ESPnet-TTS: Unified, Reproducible, and Integratable Open Source End-to-End Text-to-Speech Toolkit 24 Oct 2019 · 3 repositories · arXiv:1910.10909Syntology official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 1 pointer-only (licence)
-
Promoting the Knowledge of Source Syntax in Transformer NMT Is Not Needed 24 Oct 2019 · 0 repositories · arXiv:1910.11218
-
A Transformer with Interleaved Self-attention and Convolution for Hybrid Acoustic Models 23 Oct 2019 · 1 repository · arXiv:1910.10352
-
Controlling the Output Length of Neural Machine Translation 23 Oct 2019 · 0 repositories · arXiv:1910.10408
-
Deja-vu: Double Feature Presentation and Iterated Loss in Deep Transformer Networks 23 Oct 2019 · 2 repositories · arXiv:1910.10324
-
TCT: A Cross-supervised Learning Method for Multimodal Sequence Representation 23 Oct 2019 · 0 repositories · arXiv:1911.05186
-
Complex Transformer: A Framework for Modeling Complex-Valued Sequence 22 Oct 2019 · 1 repository · arXiv:1910.10202Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
Exchangeable deep neural networks for set-to-set matching and learning 22 Oct 2019 · 2 repositories · arXiv:1910.09972