Methods › General › Normalization › Layer Normalization › Papers, page 190
Layer Normalization
Papers archive 2025-07-28
archive papers tagged: 24,980 · with a code link: 11,273 · where Syntology ran a sample: 3,471 (2,923 with a run with no instrument failure, 548 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,471 of 24,980 tagged: 2,923 with a run with no instrument failure, 548 where every run was a failure of Syntology's instrument)
Page 190 of 250: papers 18,901 to 19,000 of 24,980, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Improved Robustness of Vision Transformer via PreLayerNorm in Patch Embedding 16 Nov 2021 · 0 repositories · arXiv:2111.08413
-
Improving Compositional Generalization with Self-Training for Data-to-Text Generation 16 Nov 2021 · 0 repositories
-
Improving GPT-3 after deployment with a dynamic memory of feedback 16 Nov 2021 · 0 repositories
-
Improving Neural Models for Radiology Report Retrieval with Lexicon-based Automated Annotation 16 Nov 2021 · 0 repositories
-
Improving Unsupervised Sentence Simplification Using Fine-Tuned Masked Language Models 16 Nov 2021 · 0 repositories
-
Input-specific Attention Subnetworks for Adversarial Detection 16 Nov 2021 · 0 repositories
-
Integrated Semantic and Phonetic Post-correction for Chinese Speech Recognition 16 Nov 2021 · 1 repository · arXiv:2111.08400
-
Interpreting Language Models Through Knowledge Graph Extraction 16 Nov 2021 · 1 repository · arXiv:2111.08546
-
Interpreting the Robustness of Neural NLP Models to Textual Perturbations 16 Nov 2021 · 0 repositories
-
Investigating the Use of BERT Anchors for Bilingual Lexicon Induction with Minimal Supervision 16 Nov 2021 · 0 repositories
-
Is Neural Topic Modelling Better than Clustering? An Empirical Study on Clustering with Contextual Embeddings for Topics 16 Nov 2021 · 0 repositories
-
"Is Whole Word Masking Always Better for Chinese BERT?": Probing on Chinese Grammatical Error Correction 16 Nov 2021 · 0 repositories
-
KinyaBERT: a Morphology-aware Kinyarwanda Language Model 16 Nov 2021 · 0 repositories
-
Knowledge Enhanced Embedding: Improve Model Generalization Through Knowledge Graphs 16 Nov 2021 · 0 repositories
-
Knowledge Graph is in Rescue: Task Oriented Dialogue System for Response Generation without NLU and DM 16 Nov 2021 · 0 repositories
-
Knowledge-guided Transformer for Joint Theme and Emotion Classification of Chinese Classical Poetry 16 Nov 2021 · 0 repositories
-
Language Level Classification on German Texts using a Neural Approach 16 Nov 2021 · 0 repositories
-
Learn More from Less: Improving Conversational Recommender Systems via Contextual and Time-Aware Modeling 16 Nov 2021 · 0 repositories
-
Learning Methods for Solving Astronomy Course Problems 16 Nov 2021 · 0 repositories
-
Learning Non-Autoregressive Models from Search for Unsupervised Sentence Summarization 16 Nov 2021 · 0 repositories
-
Learning to Ignore Adversarial Attacks 16 Nov 2021 · 0 repositories
-
Life after BERT: What do Other Muppets Understand about Language? 16 Nov 2021 · 0 repositories
-
Listen to Both Sides and be Enlightened! -- Hierarchical Modality Fusion Network for Entity and Relation Extraction 16 Nov 2021 · 0 repositories
-
Looking Into the Black Box - How Are Idioms Processed in BERT? 16 Nov 2021 · 0 repositories
-
LordBERT: Embedding Long Text by Segment Ordering with BERT 16 Nov 2021 · 0 repositories
-
Making Transformers Solve Compositional Tasks 16 Nov 2021 · 0 repositories
-
MarCQAp: Effective Context Modeling for Conversational Question Answering 16 Nov 2021 · 0 repositories
-
MarkBERT: Marking Word Boundaries Improves Chinese BERT 16 Nov 2021 · 1 repository
-
MDERank: A Masked Document Embedding Rank Approach for Unsupervised Keyphrase Extraction 16 Nov 2021 · 0 repositories
-
Meeting Summarization with Pre-training and Clustering Methods 16 Nov 2021 · 1 repository · arXiv:2111.08210
-
Metadata Shaping: Natural Language Annotations for the Long Tail 16 Nov 2021 · 0 repositories
-
Modeling Hierarchical Syntax Structure with Triplet Position for Source Code Summarization 16 Nov 2021 · 1 repository
-
Moving the Eiffel Tower to ROME: Tracing and Editing Facts in GPT 16 Nov 2021 · 0 repositories
-
Multi-head or Single-head? An Empirical Comparison for Transformer Training 16 Nov 2021 · 0 repositories
-
N-grammer: Augmenting Transformers with latent n-grams 16 Nov 2021 · 0 repositories
-
Neural Keyphrase Generation: Analysis and Evaluation 16 Nov 2021 · 0 repositories
-
NSP-BERT: A Prompt-based Zero-Shot Learner Through an Original Pre-training Task —— Next Sentence Prediction 16 Nov 2021 · 0 repositories
-
On the Multilingual Capabilities of Very Large-Scale English Language Models 16 Nov 2021 · 0 repositories
-
On the Robustness of Reading Comprehension Models to Entity Renaming 16 Nov 2021 · 0 repositories
-
On Vision Features in Multimodal Machine Translation 16 Nov 2021 · 0 repositories
-
Online Advertising Revenue Forecasting: An Interpretable Deep Learning Approach 16 Nov 2021 · 0 repositories · arXiv:2111.08840
-
Overcoming a Theoretical Limitation of Self-Attention 16 Nov 2021 · 0 repositories
-
PALBERT: Teaching ALBERT to Ponder 16 Nov 2021 · 0 repositories
-
PARE: A Simple and Strong Baseline for Monolingual and Multilingual Distantly Supervised Relation Extraction 16 Nov 2021 · 0 repositories
-
Perturbations in the Wild: Leveraging Human-Written Text Perturbations for Realistic Adversarial Attack and Defense 16 Nov 2021 · 0 repositories
-
PESTO: A Post-User Fusion Network for Rumour Detection on Social Media 16 Nov 2021 · 0 repositories
-
Pinyin-bert: A new solution to Chinese pinyin to character conversion task 16 Nov 2021 · 0 repositories
-
PoliSe: Reinforcing Politeness using User Sentiment for Customer Care Response Generation 16 Nov 2021 · 0 repositories
-
Probing BERT’s priors with serial reproduction chains 16 Nov 2021 · 0 repositories
-
PromptBERT: Improving BERT Sentence Embeddings with Prompts 16 Nov 2021 · 0 repositories
-
Reasoning Like Program Executors 16 Nov 2021 · 0 repositories
-
ReCo: Reliable Multi-hop Causal Reasoning via Structural Causal Recurrent Unit 16 Nov 2021 · 0 repositories
-
Representation of Ambiguity in Pre-Trained Sentence Embeddings 16 Nov 2021 · 0 repositories
-
Retrieval-based Layer-wise Adaptive Transformer for Source Code Summarization 16 Nov 2021 · 0 repositories
-
SAMBERT: Improve Aspect Sentiment Triplet Extraction by Segmenting the Attention Maps of BERT 16 Nov 2021 · 0 repositories
-
Self-Supervised Contrastive Learning with Adversarial Perturbations for Robust Pretrained Language Models 16 Nov 2021 · 0 repositories
-
Sequence-to-sequence AMR Parsing with Ancestor Information 16 Nov 2021 · 0 repositories
-
Sequence-to-Sequence Knowledge Graph Completion and Question Answering 16 Nov 2021 · 0 repositories
-
SHCT: A Successively Hierarchical Conditional Transformer for Controllable Paraphrase Generation 16 Nov 2021 · 0 repositories
-
SHIELD: Defending Textual Neural Networks against Black-Box Adversarial Attacks with Stochastic Multi-Expert Patcher 16 Nov 2021 · 0 repositories
-
ShrinkNAS : Single-Path One-Shot Operator Exploratory Training for Transformer with Dynamic Space Shrinking 16 Nov 2021 · 0 repositories
-
Softmax Bottleneck Makes Language Models Unable to Represent Multi-mode Word Distributions 16 Nov 2021 · 0 repositories
-
Solving Probability and Statistics Problems by Program Synthesis 16 Nov 2021 · 0 repositories · arXiv:2111.08267
-
Solving Probability and Statistics Problems by Program Synthesis 16 Nov 2021 · 0 repositories
-
Sparsifying Transformer Models with Trainable Representation Pooling 16 Nov 2021 · 0 repositories
-
Speaker Profiling in Multi-party Conversations 16 Nov 2021 · 0 repositories
-
Structured Pruning Learns Compact and Accurate Models 16 Nov 2021 · 0 repositories
-
SuperShaper: Task-Agnostic Super Pre-training of BERT Models with Variable Hidden Dimensions 16 Nov 2021 · 0 repositories
-
TACO: Pre-training of Deep Transformers with Attention Convolution using Disentangled Positional Representation 16 Nov 2021 · 0 repositories
-
Teaching BERT to Wait: Balancing Accuracy and Latency for Streaming Disfluency Detection 16 Nov 2021 · 0 repositories
-
Tell me who you are and i'll tell you what to do: A Persona Grounded Task Oriented Dialogue Generation System 16 Nov 2021 · 0 repositories
-
The impact of lexical and grammatical processing on generating code from natural language 16 Nov 2021 · 0 repositories
-
The Power of Prompt Tuning for Low-Resource Semantic Parsing 16 Nov 2021 · 0 repositories
-
Towards Coding Social Science Datasets with Language Models 16 Nov 2021 · 0 repositories
-
Towards Fully Self-Supervised Learning of Knowledge from Unstructured Text 16 Nov 2021 · 0 repositories
-
Towards Improving Topic Models with the BERT-based Neural Topic Encoder 16 Nov 2021 · 0 repositories
-
Understanding Attention in Machine Reading Comprehension 16 Nov 2021 · 0 repositories
-
UNICON: Unsupervised Intent Discovery via Semantic-level Contrastive Learning 16 Nov 2021 · 0 repositories
-
Unsupervised multiple-choice question generation for out-of-domain Q&A fine-tuning 16 Nov 2021 · 0 repositories
-
Weight Squeezing: Reparameterization for Knowledge Transfer and Model Compression 16 Nov 2021 · 0 repositories
-
WeTS: A Benchmark for Translation Suggestion 16 Nov 2021 · 0 repositories
-
What Works and Doesn't Work, A Deep Decoder for Neural Machine Translation 16 Nov 2021 · 0 repositories
-
When classifying grammatical role, BERT doesn't care about word order... except when it matters 16 Nov 2021 · 0 repositories
-
AI in Human-computer Gaming: Techniques, Challenges and Opportunities 15 Nov 2021 · 0 repositories · arXiv:2111.07631
-
Assessing gender bias in medical and scientific masked language models with StereoSet 15 Nov 2021 · 0 repositories · arXiv:2111.08088
-
AUTOMATED AUDIO CAPTIONING BY FINE-TUNING BART WITH AUDIOSET TAGS 15 Nov 2021 · 1 repository
-
Calculating Question Similarity is Enough: A New Method for KBQA Tasks 15 Nov 2021 · 0 repositories · arXiv:2111.07658
-
Energy-optimal Design and Control of Electric Powertrains under Motor Thermal Constraints 15 Nov 2021 · 0 repositories · arXiv:2111.07711
-
Exploring Story Generation with Multi-task Objectives in Variational Autoencoders 15 Nov 2021 · 0 repositories · arXiv:2111.08133
-
FakeTransformer: Exposing Face Forgery From Spatial-Temporal Representation Modeled By Facial Pixel Variations 15 Nov 2021 · 0 repositories · arXiv:2111.07601
-
FastFlow: Unsupervised Anomaly Detection and Localization via 2D Normalizing Flows 15 Nov 2021 · 5 repositories · arXiv:2111.07677Syntology 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
IIITT@Dravidian-CodeMix-FIRE2021: Transliterate or translate? Sentiment analysis of code-mixed text in Dravidian languages 15 Nov 2021 · 1 repository · arXiv:2111.07906
-
Improving Prosody for Unseen Texts in Speech Synthesis by Utilizing Linguistic Information and Noisy Data 15 Nov 2021 · 0 repositories · arXiv:2111.07549
-
Mask-guided Spectral-wise Transformer for Efficient Hyperspectral Image Reconstruction 15 Nov 2021 · 4 repositories · arXiv:2111.07910Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Monaural source separation: From anechoic to reverberant environments 15 Nov 2021 · 0 repositories · arXiv:2111.07578
-
Say What? Collaborative Pop Lyric Generation Using Multitask Transfer Learning 15 Nov 2021 · 0 repositories · arXiv:2111.07592
-
Scaling Law for Recommendation Models: Towards General-purpose User Representations 15 Nov 2021 · 0 repositories · arXiv:2111.11294
-
Local Multi-Head Channel Self-Attention for Facial Expression Recognition 14 Nov 2021 · 1 repository · arXiv:2111.07224
-
"Will You Find These Shortcuts?" A Protocol for Evaluating the Faithfulness of Input Salience Methods for Text Classification 14 Nov 2021 · 0 repositories · arXiv:2111.07367
-
SocialBERT -- Transformers for Online SocialNetwork Language Modelling 13 Nov 2021 · 0 repositories · arXiv:2111.07148