Methods › General › Normalization › Layer Normalization › Papers, page 196
Layer Normalization
Papers archive 2025-07-28
archive papers tagged: 24,980 · with a code link: 11,273 · where Syntology ran a sample: 3,471 (2,923 with a run with no instrument failure, 548 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,471 of 24,980 tagged: 2,923 with a run with no instrument failure, 548 where every run was a failure of Syntology's instrument)
Page 196 of 250: papers 19,501 to 19,600 of 24,980, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Dense Contrastive Visual-Linguistic Pretraining 24 Sep 2021 · 0 repositories · arXiv:2109.11778
-
Identification of Enzymatic Active Sites with Unsupervised Language Modeling 24 Sep 2021 · 0 repositories
-
Lacking the embedding of a word? Look it up into a traditional dictionary 24 Sep 2021 · 0 repositories · arXiv:2109.11763
-
Leveraging Pretrained Models for Automatic Summarization of Doctor-Patient Conversations 24 Sep 2021 · 1 repository · arXiv:2109.12174
-
Localizing Infinity-shaped fishes: Sketch-guided object localization in the wild 24 Sep 2021 · 0 repositories · arXiv:2109.11874
-
Long-Range Transformers for Dynamic Spatiotemporal Forecasting 24 Sep 2021 · 2 repositories · arXiv:2109.12218Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
Robustness and Sensitivity of BERT Models Predicting Alzheimer's Disease from Text 24 Sep 2021 · 0 repositories · arXiv:2109.11888
-
Transformers Generalize Linearly 24 Sep 2021 · 1 repository · arXiv:2109.12036
-
Breaking BERT: Understanding its Vulnerabilities for Named Entity Recognition through Adversarial Attack 23 Sep 2021 · 1 repository · arXiv:2109.11308
-
Dependency Structure for News Document Summarization 23 Sep 2021 · 0 repositories · arXiv:2109.11199
-
OH-Former: Omni-Relational High-Order Transformer for Person Re-Identification 23 Sep 2021 · 0 repositories · arXiv:2109.11159
-
Putting Words in BERT's Mouth: Navigating Contextualized Vector Spaces with Pseudowords 23 Sep 2021 · 1 repository · arXiv:2109.11491
-
The Volctrans GLAT System: Non-autoregressive Translation Meets WMT21 23 Sep 2021 · 0 repositories · arXiv:2109.11247
-
Alzheimers Dementia Detection using Acoustic & Linguistic features and Pre-Trained BERT 22 Sep 2021 · 0 repositories · arXiv:2109.11010
-
Controlled Evaluation of Grammatical Knowledge in Mandarin Chinese Language Models 22 Sep 2021 · 1 repository · arXiv:2109.11058
-
DialogueBERT: A Self-Supervised Learning based Dialogue Pre-training Encoder 22 Sep 2021 · 0 repositories · arXiv:2109.10480
-
Hierarchical Multimodal Transformer to Summarize Videos 22 Sep 2021 · 0 repositories · arXiv:2109.10559
-
Improving 360 Monocular Depth Estimation via Non-local Dense Prediction Transformer and Joint Supervised and Self-supervised Learning 22 Sep 2021 · 1 repository · arXiv:2109.10563Syntology 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
KD-VLP: Improving End-to-End Vision-and-Language Pretraining with Object Knowledge Distillation 22 Sep 2021 · 1 repository · arXiv:2109.10504
-
Language Models as Recommender Systems: Evaluations and Limitations 22 Sep 2021 · 0 repositories
-
Predicting Efficiency/Effectiveness Trade-offs for Dense vs. Sparse Retrieval Strategy Selection 22 Sep 2021 · 0 repositories · arXiv:2109.10739
-
Recursively Summarizing Books with Human Feedback 22 Sep 2021 · 0 repositories · arXiv:2109.10862
-
Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers 22 Sep 2021 · 3 repositories · arXiv:2109.10686Syntology community repositories only · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 4 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 12 harvested samples) · 1 pointer-only (licence)
-
T6D-Direct: Transformers for Multi-Object 6D Pose Direct Regression 22 Sep 2021 · 0 repositories · arXiv:2109.10948
-
The NiuTrans Machine Translation Systems for WMT21 22 Sep 2021 · 0 repositories · arXiv:2109.10485
-
Unsupervised Contextualized Document Representation 22 Sep 2021 · 1 repository · arXiv:2109.10509
-
A Comprehensive Review on Summarizing Financial News Using Deep Learning 21 Sep 2021 · 0 repositories · arXiv:2109.10118
-
BERTweetFR : Domain Adaptation of Pre-Trained Language Models for French Tweets 21 Sep 2021 · 0 repositories · arXiv:2109.10234
-
DS-Net++: Dynamic Weight Slicing for Efficient Inference in CNNs and Transformers 21 Sep 2021 · 1 repository · arXiv:2109.10060
-
InvBERT: Reconstructing Text from Contextualized Word Embeddings by inverting the BERT pipeline 21 Sep 2021 · 0 repositories · arXiv:2109.10104
-
LOTR: Face Landmark Localization Using Localization Transformer 21 Sep 2021 · 0 repositories · arXiv:2109.10057
-
Multi-Task Learning with Sentiment, Emotion, and Target Detection to Recognize Hate Speech and Offensive Language 21 Sep 2021 · 0 repositories · arXiv:2109.10255
-
Representation Learning for Short Text Clustering 21 Sep 2021 · 0 repositories · arXiv:2109.09894
-
TrOCR: Transformer-based Optical Character Recognition with Pre-trained Models 21 Sep 2021 · 8 repositories · arXiv:2109.10282Syntology community repositories only · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
A Plug-and-Play Method for Controlled Text Generation 20 Sep 2021 · 1 repository · arXiv:2109.09707
-
BARTpho: Pre-trained Sequence-to-Sequence Models for Vietnamese 20 Sep 2021 · 3 repositories · arXiv:2109.09701
-
BERT Cannot Align Characters 20 Sep 2021 · 0 repositories · arXiv:2109.09700
-
BERT Has Uncommon Sense: Similarity Ranking for Word Sense BERTology 20 Sep 2021 · 1 repository · arXiv:2109.09780
-
Dyadformer: A Multi-modal Transformer for Long-Range Modeling of Dyadic Interactions 20 Sep 2021 · 0 repositories · arXiv:2109.09487
-
iRNN: Integer-only Recurrent Neural Network 20 Sep 2021 · 0 repositories · arXiv:2109.09828
-
MFEViT: A Robust Lightweight Transformer-based Network for Multimodal 2D+3D Facial Expression Recognition 20 Sep 2021 · 0 repositories · arXiv:2109.13086
-
Model Bias in NLP -- Application to Hate Speech Classification using transfer learning techniques 20 Sep 2021 · 0 repositories · arXiv:2109.09725
-
Well Googled is Half Done: Multimodal Forecasting of New Fashion Product Sales with Image-based Google Trends 20 Sep 2021 · 1 repository · arXiv:2109.09824Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
CLIFF: Contrastive Learning for Improving Faithfulness and Factuality in Abstractive Summarization 19 Sep 2021 · 3 repositories · arXiv:2109.09209
-
Do Long-Range Language Models Actually Use Long-Range Context? 19 Sep 2021 · 0 repositories · arXiv:2109.09115
-
MirrorWiC: On Eliciting Word-in-Context Representations from Pretrained Language Models 19 Sep 2021 · 1 repository · arXiv:2109.09237
-
The Seismo-Performer: A Novel Machine Learning Approach for General and Efficient Seismic Phase Recognition from Local Earthquakes in Real Time 19 Sep 2021 · 1 repository
-
Towards Zero-Label Language Learning 19 Sep 2021 · 0 repositories · arXiv:2109.09193
-
Wav-BERT: Cooperative Acoustic and Linguistic Representation Learning for Low-Resource Speech Recognition 19 Sep 2021 · 0 repositories · arXiv:2109.09161
-
What BERT Based Language Models Learn in Spoken Transcripts: An Empirical Study 19 Sep 2021 · 0 repositories · arXiv:2109.09105
-
Complex Temporal Question Answering on Knowledge Graphs 18 Sep 2021 · 1 repository · arXiv:2109.08935
-
DyLex: Incorporating Dynamic Lexicons into BERT for Sequence Labeling 18 Sep 2021 · 1 repository · arXiv:2109.08818
-
UNetFormer: A UNet-like Transformer for Efficient Semantic Segmentation of Remote Sensing Urban Scene Imagery 18 Sep 2021 · 1 repository · arXiv:2109.08937Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
SDTP: Semantic-aware Decoupled Transformer Pyramid for Dense Image Prediction 18 Sep 2021 · 0 repositories · arXiv:2109.08963
-
Text Detoxification using Large Pre-trained Neural Models 18 Sep 2021 · 1 repository · arXiv:2109.08914Syntology official (archive's flag): 4 ran · 4 ran (of which 1 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Towards High-Quality Temporal Action Detection with Sparse Proposals 18 Sep 2021 · 1 repository · arXiv:2109.08847
-
Commonsense Knowledge-Augmented Pretrained Language Models for Causal Reasoning Classification 17 Sep 2021 · 0 repositories
-
Context vs Target Word: Quantifying Biases When Applying Models to Lexical Semantic Datasets 17 Sep 2021 · 0 repositories
-
Continuous Streaming Multi-Talker ASR with Dual-path Transducers 17 Sep 2021 · 0 repositories · arXiv:2109.08555
-
Defending Textual Neural Networks against Black-Box Adversarial Attacks with Stochastic Multi-Expert Patcher 17 Sep 2021 · 0 repositories
-
Digging Errors in NMT: Evaluating and Understanding Model Errors from Hypothesis Distribution 17 Sep 2021 · 0 repositories
-
Does BERT really agree ? Fine-grained Analysis of Lexical Dependence on a Syntactic Task 17 Sep 2021 · 0 repositories
-
Expression Snippet Transformer for Robust Video-based Facial Expression Recognition 17 Sep 2021 · 0 repositories · arXiv:2109.08409
-
Fine-Tuned Transformers Show Clusters of Similar Representations Across Layers 17 Sep 2021 · 0 repositories · arXiv:2109.08406
-
From Known to Unknown: Knowledge-guided Transformer for Time-Series Sales Forecasting in Alibaba 17 Sep 2021 · 0 repositories · arXiv:2109.08381
-
Grounding Natural Language Instructions: Can Large Language Models Capture Spatial Information? 17 Sep 2021 · 1 repository · arXiv:2109.08634
-
Hierarchy-Aware T5 with Path-Adaptive Mask Mechanism for Hierarchical Text Classification 17 Sep 2021 · 0 repositories · arXiv:2109.08585
-
Knowledge Neurons in Pretrained Transformers 17 Sep 2021 · 0 repositories
-
Learning Low-frequency Patterns with A Pre-trained Document-Grounded Conversation Model 17 Sep 2021 · 0 repositories
-
General Cross-Architecture Distillation of Pretrained Language Models into Matrix Embeddings 17 Sep 2021 · 1 repository · arXiv:2109.08449
-
Primer: Searching for Efficient Transformers for Language Modeling 17 Sep 2021 · 4 repositories · arXiv:2109.08668Syntology official: no sample here; runs from other or unrecorded repositories · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 2 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Relating Neural Text Degeneration to Exposure Bias 17 Sep 2021 · 0 repositories · arXiv:2109.08705
-
Remixers: A Mixer-Transformer Architecture with Compositional Operators for Natural Language Understanding 17 Sep 2021 · 0 repositories
-
Scaling Laws vs Model Architectures: How does Inductive Bias Influence Scaling? An Extensive Empirical Study on Language Tasks 17 Sep 2021 · 0 repositories
-
The futility of STILTs for the classification of lexical borrowings in Spanish 17 Sep 2021 · 0 repositories · arXiv:2109.08607
-
The JHU-Microsoft Submission for WMT21 Quality Estimation Shared Task 17 Sep 2021 · 0 repositories · arXiv:2109.08724
-
An End-to-End Transformer Model for 3D Object Detection 16 Sep 2021 · 1 repository · arXiv:2109.08141Syntology 6 ran (of which 3 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 1 pointer-only (licence)
-
Fast-Slow Transformer for Visually Grounding Speech 16 Sep 2021 · 1 repository · arXiv:2109.08186
-
Generative Pre-Training from Molecules 16 Sep 2021 · 1 repository
-
Label-Attention Transformer with Geometrically Coherent Objects for Image Captioning 16 Sep 2021 · 1 repository · arXiv:2109.07799
-
Language Models are Few-shot Multilingual Learners 16 Sep 2021 · 1 repository · arXiv:2109.07684Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Let the CAT out of the bag: Contrastive Attributed explanations for Text 16 Sep 2021 · 0 repositories · arXiv:2109.07983
-
MeLT: Message-Level Transformer with Masked Document Representations as Pre-Training for Stance Detection 16 Sep 2021 · 1 repository · arXiv:2109.08113
-
MOVER: Mask, Over-generate and Rank for Hyperbole Generation 16 Sep 2021 · 1 repository · arXiv:2109.07726
-
RetrievalSum: A Retrieval Enhanced Framework for Abstractive Summarization 16 Sep 2021 · 0 repositories · arXiv:2109.07943
-
Revisiting Tri-training of Dependency Parsers 16 Sep 2021 · 2 repositories · arXiv:2109.08122
-
Scaling Laws for Neural Machine Translation 16 Sep 2021 · 0 repositories · arXiv:2109.07740
-
Sparse Factorization of Large Square Matrices 16 Sep 2021 · 1 repository · arXiv:2109.08184
-
TANet: A new Paradigm for Global Face Super-resolution via Transformer-CNN Aggregation Network 16 Sep 2021 · 0 repositories · arXiv:2109.08174
-
The NiuTrans System for the WMT21 Efficiency Task 16 Sep 2021 · 1 repository · arXiv:2109.08003
-
The NiuTrans System for WNGT 2020 Efficiency Task 16 Sep 2021 · 2 repositories · arXiv:2109.08008Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Utterance-level neural confidence measure for end-to-end children speech recognition 16 Sep 2021 · 0 repositories · arXiv:2109.07750
-
Anchor DETR: Query Design for Transformer-Based Object Detection 15 Sep 2021 · 2 repositories · arXiv:2109.07107
-
Attention Is Indeed All You Need: Semantically Attention-Guided Decoding for Data-to-Text NLG 15 Sep 2021 · 1 repository · arXiv:2109.07043
-
BERT is Robust! A Case Against Synonym-Based Adversarial Examples in Text Classification 15 Sep 2021 · 0 repositories · arXiv:2109.07403
-
EfficientBERT: Progressively Searching Multilayer Perceptron via Warm-up Knowledge Distillation 15 Sep 2021 · 1 repository · arXiv:2109.07222Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Efficient Domain Adaptation of Language Models via Adaptive Tokenization 15 Sep 2021 · 0 repositories · arXiv:2109.07460
-
Enhancing Clinical Information Extraction with Transferred Contextual Embeddings 15 Sep 2021 · 0 repositories · arXiv:2109.07243
-
Complementary Feature Enhanced Network with Vision Transformer for Image Dehazing 15 Sep 2021 · 1 repository · arXiv:2109.07100
-
Improving Text Auto-Completion with Next Phrase Prediction 15 Sep 2021 · 0 repositories · arXiv:2109.07067