Methods › General › Normalization › Layer Normalization › Papers, page 214
Layer Normalization
Papers archive 2025-07-28
archive papers tagged: 24,980 · with a code link: 11,273 · where Syntology ran a sample: 3,471 (2,923 with a run with no instrument failure, 548 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (3,471 of 24,980 tagged: 2,923 with a run with no instrument failure, 548 where every run was a failure of Syntology's instrument)
Page 214 of 250: papers 21,301 to 21,400 of 24,980, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
MixUp Training Leads to Reduced Overfitting and Improved Calibration for the Transformer Architecture 22 Feb 2021 · 0 repositories · arXiv:2102.11402
-
Parallelizing Legendre Memory Unit Training 22 Feb 2021 · 2 repositories · arXiv:2102.11417
-
Position Information in Transformers: An Overview 22 Feb 2021 · 0 repositories · arXiv:2102.11090
-
RUBERT: A Bilingual Roman Urdu BERT Using Cross Lingual Transfer Learning 22 Feb 2021 · 0 repositories · arXiv:2102.11278
-
UniT: Multimodal Multitask Learning with a Unified Transformer 22 Feb 2021 · 1 repository · arXiv:2102.10772
-
Using Prior Knowledge to Guide BERT's Attention in Semantic Textual Matching Tasks 22 Feb 2021 · 1 repository · arXiv:2102.10934
-
Medical Transformer: Gated Axial-Attention for Medical Image Segmentation 21 Feb 2021 · 2 repositories · arXiv:2102.10662
-
Pre-Training BERT on Arabic Tweets: Practical Considerations 21 Feb 2021 · 0 repositories · arXiv:2102.10684
-
Web-based Application for Detecting Indonesian Clickbait Headlines using IndoBERT 21 Feb 2021 · 0 repositories · arXiv:2102.10601
-
Multilingual Answer Sentence Reranking via Automatically Translated Data 20 Feb 2021 · 0 repositories · arXiv:2102.10250
-
Towards Accurate and Compact Architectures via Neural Architecture Transformer 20 Feb 2021 · 2 repositories · arXiv:2102.10301
-
Calibrate Before Use: Improving Few-Shot Performance of Language Models 19 Feb 2021 · 5 repositories · arXiv:2102.09690Syntology official: harvested, nothing ran · 0 ran · 4 unverified (of 4 harvested samples)
-
Convolutional Normalization 19 Feb 2021 · 0 repositories · arXiv:2102.09685
-
Dialect Identification in Nuanced Arabic Tweets Using Farasa Segmentation and AraBERT 19 Feb 2021 · 0 repositories · arXiv:2102.09749
-
Latent Variable Sequential Set Transformers For Joint Multi-Agent Motion Prediction 19 Feb 2021 · 2 repositories · arXiv:2104.00563Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 1 where Syntology's instrument failed) · 6 unverified (of 17 harvested samples) · 2 pointer-only (licence)
-
Learning Dynamic BERT via Trainable Gate Variables and a Bi-modal Regularizer 19 Feb 2021 · 0 repositories · arXiv:2102.09727
-
Towards Emotion Recognition in Hindi-English Code-Mixed Data: A Transformer Based Approach 19 Feb 2021 · 1 repository · arXiv:2102.09943
-
Using Transformer based Ensemble Learning to classify Scientific Articles 19 Feb 2021 · 2 repositories · arXiv:2102.09991
-
Analysis Of Contextual and Non-Contextual Word Embedding Models For Hindi NER With Web Application For Data Collection 18 Feb 2021 · 1 repository
-
Going Full-TILT Boogie on Document Understanding with Text-Image-Layout Transformer 18 Feb 2021 · 1 repository · arXiv:2102.09550
-
Quiz-Style Question Generation for News Stories 18 Feb 2021 · 2 repositories · arXiv:2102.09094
-
Training Large-Scale News Recommenders with Pretrained Language Models in the Loop 18 Feb 2021 · 1 repository · arXiv:2102.09268
-
UnibucKernel: Geolocating Swiss German Jodels Using Ensemble Learning 18 Feb 2021 · 0 repositories · arXiv:2102.09379
-
Beyond Fully-Connected Layers with Quaternions: Parameterization of Hypercomplex Multiplications with 1/n Parameters 17 Feb 2021 · 3 repositories · arXiv:2102.08597Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Leveraging Query Resolution and Reading Comprehension for Conversational Passage Retrieval 17 Feb 2021 · 0 repositories · arXiv:2102.08795
-
SciDr at SDU-2020: IDEAS -- Identifying and Disambiguating Everyday Acronyms for Scientific Domain 17 Feb 2021 · 2 repositories · arXiv:2102.08818
-
TCN: Table Convolutional Network for Web Table Interpretation 17 Feb 2021 · 1 repository · arXiv:2102.09460
-
THEaiTRE 1.0: Interactive generation of theatre play scripts 17 Feb 2021 · 0 repositories · arXiv:2102.08892
-
COCO-LM: Correcting and Contrasting Text Sequences for Language Model Pretraining 16 Feb 2021 · 2 repositories · arXiv:2102.08473Syntology community repositories only · 5 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
Exploring Transformers in Natural Language Generation: GPT, BERT, and XLNet 16 Feb 2021 · 1 repository · arXiv:2102.08036
-
GradInit: Learning to Initialize Neural Networks for Stable and Efficient Training 16 Feb 2021 · 2 repositories · arXiv:2102.08098Syntology official (archive's flag): 6 ran · 7 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 3 where Syntology's instrument failed) · 6 unverified (of 13 harvested samples) · 12 pointer-only (licence)
-
Have Attention Heads in BERT Learned Constituency Grammar? 16 Feb 2021 · 0 repositories · arXiv:2102.07926
-
Non-Autoregressive Text Generation with Pre-trained Language Models 16 Feb 2021 · 1 repository · arXiv:2102.08220
-
Revisiting Language Encoding in Learning Multilingual Representations 16 Feb 2021 · 1 repository · arXiv:2102.08357
-
TeraPipe: Token-Level Pipeline Parallelism for Training Large-Scale Language Models 16 Feb 2021 · 1 repository · arXiv:2102.07988Syntology official (archive's flag): 7 ran · 7 ran (of which 3 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
DOBF: A Deobfuscation Pre-Training Objective for Programming Languages 15 Feb 2021 · 2 repositories · arXiv:2102.07492
-
Fast End-to-End Speech Recognition via Non-Autoregressive Models and Cross-Modal Knowledge Transferring from BERT 15 Feb 2021 · 0 repositories · arXiv:2102.07594
-
Improved Customer Transaction Classification using Semi-Supervised Knowledge Distillation 15 Feb 2021 · 0 repositories · arXiv:2102.07635
-
Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm 15 Feb 2021 · 0 repositories · arXiv:2102.07350
-
The corruptive force of AI-generated advice 15 Feb 2021 · 0 repositories · arXiv:2102.07536
-
Translational Equivariance in Kernelizable Attention 15 Feb 2021 · 1 repository · arXiv:2102.07680
-
Within-Document Event Coreference with BERT-Based Contextualized Representations 15 Feb 2021 · 0 repositories · arXiv:2102.09600
-
indicnlp@kgp at DravidianLangTech-EACL2021: Offensive Language Identification in Dravidian Languages 14 Feb 2021 · 1 repository · arXiv:2102.07150
-
indicnlp@ kgp at DravidianLangTech-EACL2021: Offensive Language Identification in Dravidian Languages 14 Feb 2021 · 1 repository
-
Characterizing English Variation across Social Media Communities with BERT 12 Feb 2021 · 1 repository · arXiv:2102.06820
-
Dancing along Battery: Enabling Transformer with Run-time Reconfigurability on Mobile Devices 12 Feb 2021 · 0 repositories · arXiv:2102.06336
-
Dynamic Precision Analog Computing for Neural Networks 12 Feb 2021 · 1 repository · arXiv:2102.06365
-
Exploring Classic and Neural Lexical Translation Models for Information Retrieval: Interpretability, Effectiveness, and Efficiency Benefits 12 Feb 2021 · 2 repositories · arXiv:2102.06815
-
Improving Zero-shot Neural Machine Translation on Language-specific Encoders-Decoders 12 Feb 2021 · 0 repositories · arXiv:2102.06578
-
Multiversal views on language models 12 Feb 2021 · 0 repositories · arXiv:2102.06391
-
Optimizing Inference Performance of Transformers on CPUs 12 Feb 2021 · 0 repositories · arXiv:2102.06621
-
Transformer Language Models with LSTM-based Cross-utterance Information Representation 12 Feb 2021 · 1 repository · arXiv:2102.06474
-
Proof Artifact Co-training for Theorem Proving with Language Models 11 Feb 2021 · 4 repositories · arXiv:2102.06203Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Text Compression-aided Transformer Encoding 11 Feb 2021 · 0 repositories · arXiv:2102.05951
-
NAST: Non-Autoregressive Spatial-Temporal Transformer for Time Series Forecasting 10 Feb 2021 · 1 repository · arXiv:2102.05624
-
AuGPT: Auxiliary Tasks and Data Augmentation for End-To-End Dialogue with Pre-Trained Language Models 9 Feb 2021 · 1 repository · arXiv:2102.05126
-
Bayesian Transformer Language Models for Speech Recognition 9 Feb 2021 · 0 repositories · arXiv:2102.04754
-
Conversational Query Rewriting with Self-supervised Learning 9 Feb 2021 · 0 repositories · arXiv:2102.04708
-
Joint Intent Detection and Slot Filling with Wheel-Graph Attention Networks 9 Feb 2021 · 0 repositories · arXiv:2102.04610
-
NewsBERT: Distilling Pre-trained Language Model for Intelligent News Application 9 Feb 2021 · 0 repositories · arXiv:2102.04887
-
Point Cloud Transformers applied to Collider Physics 9 Feb 2021 · 1 repository · arXiv:2102.05073
-
Transfer Learning Approach for Arabic Offensive Language Detection System -- BERT-Based Model 9 Feb 2021 · 0 repositories · arXiv:2102.05708
-
A Hybrid Task-Oriented Dialog System with Domain and Task Adaptive Pretraining 8 Feb 2021 · 0 repositories · arXiv:2102.04506
-
Colorization Transformer 8 Feb 2021 · 2 repositories · arXiv:2102.04432
-
Generating Fake Cyber Threat Intelligence Using Transformer-Based Models 8 Feb 2021 · 0 repositories · arXiv:2102.04351
-
Bias Out-of-the-Box: An Empirical Analysis of Intersectional Occupational Biases in Popular Generative Language Models 8 Feb 2021 · 1 repository · arXiv:2102.04130Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
TransReID: Transformer-based Object Re-Identification 8 Feb 2021 · 4 repositories · arXiv:2102.04378
-
TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation 8 Feb 2021 · 22 repositories · arXiv:2102.04306Syntology official: no sample here; runs from other or unrecorded repositories · 6 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Wake Word Detection with Streaming Transformers 8 Feb 2021 · 0 repositories · arXiv:2102.04488
-
Nyströmformer: A Nyström-Based Algorithm for Approximating Self-Attention 7 Feb 2021 · 10 repositories · arXiv:2102.03902Syntology official: no sample here; runs from other or unrecorded repositories · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 1 pointer-only (licence)
-
Spoiler Alert: Using Natural Language Processing to Detect Spoilers in Book Reviews 7 Feb 2021 · 1 repository · arXiv:2102.03882
-
Jointly Improving Language Understanding and Generation with Quality-Weighted Weak Supervision of Automatic Labeling 6 Feb 2021 · 0 repositories · arXiv:2102.03551
-
Neural Data-to-Text Generation with LM-based Text Augmentation 6 Feb 2021 · 0 repositories · arXiv:2102.03556
-
baller2vec: A Multi-Entity Transformer For Multi-Agent Spatiotemporal Modeling 5 Feb 2021 · 1 repository · arXiv:2102.03291Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
PipeTransformer: Automated Elastic Pipelining for Distributed Training of Transformers 5 Feb 2021 · 1 repository · arXiv:2102.03161Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
RpBERT: A Text-image Relation Propagation-based BERT Model for Multimodal NER 5 Feb 2021 · 1 repository · arXiv:2102.02967
-
Understanding Emails and Drafting Responses -- An Approach Using GPT-3 5 Feb 2021 · 0 repositories · arXiv:2102.03062
-
ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision 5 Feb 2021 · 6 repositories · arXiv:2102.03334Syntology official: harvested, nothing ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified; the one sample that ran constructed an object rather than computing a result (of 4 harvested samples) · 1 pointer-only (licence)
-
1-bit Adam: Communication Efficient Large-Scale Training with Adam's Convergence Speed 4 Feb 2021 · 2 repositories · arXiv:2102.02888
-
Adaptive Semiparametric Language Models 4 Feb 2021 · 0 repositories · arXiv:2102.02557
-
Hierarchical Multi-head Attentive Network for Evidence-aware Fake News Detection 4 Feb 2021 · 1 repository · arXiv:2102.02680
-
Understanding the Capabilities, Limitations, and Societal Impact of Large Language Models 4 Feb 2021 · 0 repositories · arXiv:2102.02503
-
Bootstrapping Multilingual AMR with Contextual Word Alignments 3 Feb 2021 · 0 repositories · arXiv:2102.02189
-
HeBERT & HebEMO: a Hebrew BERT Model and a Tool for Polarity Analysis and Emotion Recognition 3 Feb 2021 · 0 repositories · arXiv:2102.01909
-
MUFASA: Multimodal Fusion Architecture Search for Electronic Health Records 3 Feb 2021 · 0 repositories · arXiv:2102.02340
-
Introduction to Neural Transfer Learning with Transformers for Social Science Text Analysis 3 Feb 2021 · 0 repositories · arXiv:2102.02111
-
Mind the Gap: Assessing Temporal Generalization in Neural Language Models 3 Feb 2021 · 1 repository · arXiv:2102.01951
-
Relaxed Transformer Decoders for Direct Action Proposal Generation 3 Feb 2021 · 2 repositories · arXiv:2102.01894Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
Towards Natural and Controllable Cross-Lingual Voice Conversion Based on Neural TTS Model and Phonetic Posteriorgram 3 Feb 2021 · 0 repositories · arXiv:2102.01991
-
AutoFreeze: Automatically Freezing Model Blocks to Accelerate Fine-tuning 2 Feb 2021 · 1 repository · arXiv:2102.01386Syntology official: no sample here; runs from other or unrecorded repositories · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Clickbait Headline Detection in Indonesian News Sites using Multilingual Bidirectional Encoder Representations from Transformers (M-BERT) 2 Feb 2021 · 0 repositories · arXiv:2102.01497
-
Automated Query Reformulation for Efficient Search based on Query Logs From Stack Overflow 1 Feb 2021 · 1 repository · arXiv:2102.00826
-
GTAE: Graph-Transformer based Auto-Encoders for Linguistic-Constrained Text Style Transfer 1 Feb 2021 · 0 repositories · arXiv:2102.00769
-
Improving Distantly-Supervised Relation Extraction through BERT-based Label & Instance Embeddings 1 Feb 2021 · 1 repository · arXiv:2102.01156
-
"Is depression related to cannabis?": A knowledge-infused model for Entity and Relation Extraction with Limited Supervision 1 Feb 2021 · 0 repositories · arXiv:2102.01222
-
Polyphone Disambiguation in Mandarin Chinese with Semi-Supervised Learning 1 Feb 2021 · 0 repositories · arXiv:2102.00621
-
Scaling Federated Learning for Fine-tuning of Large Language Models 1 Feb 2021 · 0 repositories · arXiv:2102.00875
-
SJ_AJ@DravidianLangTech-EACL2021: Task-Adaptive Pre-Training of Multilingual BERT models for Offensive Language Identification 1 Feb 2021 · 1 repository · arXiv:2102.01051
-
Text-to-hashtag Generation using Seq2seq Learning 1 Feb 2021 · 1 repository · arXiv:2102.00904
-
A Runtime-Based Computational Performance Predictor for Deep Neural Network Training 31 Jan 2021 · 1 repository · arXiv:2102.00527