Methods › General › Regularization › Attention Dropout › Papers, page 86
Attention Dropout
Papers archive 2025-07-28
archive papers tagged: 10,892 · with a code link: 4,634 · where Syntology ran a sample: 1,270 (1,043 with a run with no instrument failure, 227 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (1,270 of 10,892 tagged: 1,043 with a run with no instrument failure, 227 where every run was a failure of Syntology's instrument)
Page 86 of 109: papers 8,501 to 8,600 of 10,892, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Modeling "Newsworthiness" for Lead-Generation Across Corpora 19 Apr 2021 · 0 repositories · arXiv:2104.09653
-
Neural Language Models with Distant Supervision to Identify Major Depressive Disorder from Clinical Notes 19 Apr 2021 · 0 repositories · arXiv:2104.09644
-
OCTIS: Comparing and Optimizing Topic models is Simple! 19 Apr 2021 · 1 repository
-
Operationalizing a National Digital Library: The Case for a Norwegian Transformer Model 19 Apr 2021 · 2 repositories · arXiv:2104.09617
-
Probing for Bridging Inference in Transformer Language Models 19 Apr 2021 · 1 repository · arXiv:2104.09400
-
Sentiment Classification in Swahili Language Using Multilingual BERT 19 Apr 2021 · 0 repositories · arXiv:2104.09006
-
TeamUNCC@LT-EDI-EACL2021: Hope Speech Detection using Transfer Learning with Transformers 19 Apr 2021 · 1 repository
-
A Token-level Reference-free Hallucination Detection Benchmark for Free-form Text Generation 18 Apr 2021 · 2 repositories · arXiv:2104.08704Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
CEAR: Cross-Entity Aware Reranker for Knowledge Base Completion 18 Apr 2021 · 0 repositories · arXiv:2104.08741
-
Dual-View Distilled BERT for Sentence Embedding 18 Apr 2021 · 0 repositories · arXiv:2104.08675
-
Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity 18 Apr 2021 · 2 repositories · arXiv:2104.08786Syntology 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 7 unverified (of 10 harvested samples)
-
FedNLP: Benchmarking Federated Learning Methods for Natural Language Processing Tasks 18 Apr 2021 · 1 repository · arXiv:2104.08815
-
GPT3Mix: Leveraging Large-scale Language Models for Text Augmentation 18 Apr 2021 · 1 repository · arXiv:2104.08826Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Knowledge Neurons in Pretrained Transformers 18 Apr 2021 · 3 repositories · arXiv:2104.08696Syntology official (archive's flag): 1 ran · 2 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Language in a (Search) Box: Grounding Language Learning in Real-World Human-Machine Interaction 18 Apr 2021 · 0 repositories · arXiv:2104.08874
-
MT6: Multilingual Pretrained Text-to-Text Transformer with Translation Pairs 18 Apr 2021 · 2 repositories · arXiv:2104.08692
-
Cross-Task Generalization via Natural Language Crowdsourcing Instructions 18 Apr 2021 · 3 repositories · arXiv:2104.08773Syntology official (archive's flag): 4 ran · 5 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
Rethinking Network Pruning -- under the Pre-train and Fine-tune Paradigm 18 Apr 2021 · 1 repository · arXiv:2104.08682Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample) · 1 pointer-only (licence)
-
SimCSE: Simple Contrastive Learning of Sentence Embeddings 18 Apr 2021 · 23 repositories · arXiv:2104.08821Syntology community repositories only · 17 ran (of which 9 constructed an object rather than computing a result; 12 with no instrument failure: 1 honoured, 0 violated, 11 with no contract checked; 5 where Syntology's instrument failed) · 13 unverified (of 30 harvested samples) · 19 pointer-only (licence)
-
The Power of Scale for Parameter-Efficient Prompt Tuning 18 Apr 2021 · 12 repositories · arXiv:2104.08691Syntology community repositories only · 13 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 15 harvested samples)
-
Zero-shot Cross-lingual Transfer of Neural Machine Translation with Multilingual Pretrained Encoders 18 Apr 2021 · 1 repository · arXiv:2104.08757Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
A multilabel approach to morphosyntactic probing 17 Apr 2021 · 0 repositories · arXiv:2104.08464
-
ASBERT: Siamese and Triplet network embedding for open question answering 17 Apr 2021 · 0 repositories · arXiv:2104.08558
-
Co-BERT: A Context-Aware BERT Retrieval Model Incorporating Local and Query-specific Context 17 Apr 2021 · 0 repositories · arXiv:2104.08523
-
Decrypting Cryptic Crosswords: Semantically Complex Wordplay Puzzles as a Target for NLP 17 Apr 2021 · 2 repositories · arXiv:2104.08620
-
Frequency-based Distortions in Contextualized Word Embeddings 17 Apr 2021 · 0 repositories · arXiv:2104.08465
-
Three-level Hierarchical Transformer Networks for Long-sequence and Multiple Clinical Documents Classification 17 Apr 2021 · 1 repository · arXiv:2104.08444
-
Identifying the Limits of Cross-Domain Knowledge Transfer for Pretrained Models 17 Apr 2021 · 1 repository · arXiv:2104.08410
-
Improving Zero-Shot Cross-Lingual Transfer Learning via Robust Training 17 Apr 2021 · 1 repository · arXiv:2104.08645
-
Multi-source Neural Topic Modeling in Multi-view Embedding Spaces 17 Apr 2021 · 1 repository · arXiv:2104.08551
-
The Topic Confusion Task: A Novel Scenario for Authorship Attribution 17 Apr 2021 · 0 repositories · arXiv:2104.08530
-
UPB at SemEval-2021 Task 5: Virtual Adversarial Training for Toxic Spans Detection 17 Apr 2021 · 0 repositories · arXiv:2104.08635
-
Zero-shot Slot Filling with DPR and RAG 17 Apr 2021 · 2 repositories · arXiv:2104.08610
-
An Adversarially-Learned Turing Test for Dialog Generation Models 16 Apr 2021 · 1 repository · arXiv:2104.08231
-
An Analysis of a BERT Deep Learning Strategy on a Technology Assisted Review Task 16 Apr 2021 · 0 repositories · arXiv:2104.08340
-
Memorisation versus Generalisation in Pre-trained Language Models 16 Apr 2021 · 1 repository · arXiv:2105.00828Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Editing Factual Knowledge in Language Models 16 Apr 2021 · 3 repositories · arXiv:2104.08164Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
Fast, Effective, and Self-Supervised: Transforming Masked Language Models into Universal Lexical and Sentence Encoders 16 Apr 2021 · 1 repository · arXiv:2104.08027Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Membership Inference Attack Susceptibility of Clinical Language Models 16 Apr 2021 · 0 repositories · arXiv:2104.08305
-
Probing Across Time: What Does RoBERTa Know and When? 16 Apr 2021 · 1 repository · arXiv:2104.07885Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Surface Form Competition: Why the Highest Probability Answer Isn't Always Right 16 Apr 2021 · 2 repositories · arXiv:2104.08315Syntology community repositories only · 5 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Temporal Adaptation of BERT and Performance on Downstream Document Classification: Insights from Social Media 16 Apr 2021 · 2 repositories · arXiv:2104.08116
-
Text2App: A Framework for Creating Android Apps from Text Descriptions 16 Apr 2021 · 2 repositories · arXiv:2104.08301
-
Towards Variable-Length Textual Adversarial Attacks 16 Apr 2021 · 0 repositories · arXiv:2104.08139
-
A Sample-Based Training Method for Distantly Supervised Relation Extraction with Pre-Trained Transformers 15 Apr 2021 · 0 repositories · arXiv:2104.07512
-
A Survey of Recent Abstract Summarization Techniques 15 Apr 2021 · 0 repositories · arXiv:2105.00824
-
Are Multilingual BERT models robust? A Case Study on Adversarial Attacks for Multilingual Question Answering 15 Apr 2021 · 0 repositories · arXiv:2104.07646
-
BERT based Transformers lead the way in Extraction of Health Information from Social Media 15 Apr 2021 · 1 repository · arXiv:2104.07367
-
Does BERT Pretrained on Clinical Notes Reveal Sensitive Data? 15 Apr 2021 · 4 repositories · arXiv:2104.07762Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Emotion Dynamics Modeling via BERT 15 Apr 2021 · 0 repositories · arXiv:2104.07252
-
ExplaGraphs: An Explanation Graph Generation Task for Structured Commonsense Reasoning 15 Apr 2021 · 1 repository · arXiv:2104.07644Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples)
-
How to Train BERT with an Academic Budget 15 Apr 2021 · 4 repositories · arXiv:2104.07705
-
NT5?! Training T5 to Perform Numerical Reasoning 15 Apr 2021 · 1 repository · arXiv:2104.07307
-
Natural Language Understanding with Privacy-Preserving BERT 15 Apr 2021 · 0 repositories · arXiv:2104.07504
-
SINA-BERT: A pre-trained Language Model for Analysis of Medical Texts in Persian 15 Apr 2021 · 0 repositories · arXiv:2104.07613
-
Text Guide: Improving the quality of long text classification by a text selection method based on feature importance 15 Apr 2021 · 1 repository · arXiv:2104.07225
-
TorontoCL at CMCL 2021 Shared Task: RoBERTa with Multi-Stage Fine-Tuning for Eye-Tracking Prediction 15 Apr 2021 · 1 repository · arXiv:2104.07244
-
Ultra-High Dimensional Sparse Representations with Binarization for Efficient Text Retrieval 15 Apr 2021 · 0 repositories · arXiv:2104.07198
-
UIT-E10dot3 at SemEval-2021 Task 5: Toxic Spans Detection with Named Entity Recognition and Question-Answering Approaches 15 Apr 2021 · 0 repositories · arXiv:2104.07376
-
An Interpretability Illusion for BERT 14 Apr 2021 · 0 repositories · arXiv:2104.07143
-
Demystifying BERT: Implications for Accelerator Design 14 Apr 2021 · 0 repositories · arXiv:2104.08335
-
Enhancing Word-Level Semantic Representation via Dependency Structure for Expressive Text-to-Speech Synthesis 14 Apr 2021 · 0 repositories · arXiv:2104.06835
-
Disentangling Representations of Text by Masking Transformers 14 Apr 2021 · 0 repositories · arXiv:2104.07155
-
NAREOR: The Narrative Reordering Problem 14 Apr 2021 · 1 repository · arXiv:2104.06669
-
On the Robustness of Intent Classification and Slot Labeling in Goal-oriented Dialog Systems to Real-world Noise 14 Apr 2021 · 1 repository · arXiv:2104.07149
-
Static Embeddings as Efficient Knowledge Bases? 14 Apr 2021 · 1 repository · arXiv:2104.07094
-
1-bit LAMB: Communication Efficient Large-Scale Large-Batch Training with LAMB's Convergence Speed 13 Apr 2021 · 1 repository · arXiv:2104.06069
-
Discourse Probing of Pretrained Language Models 13 Apr 2021 · 1 repository · arXiv:2104.05882
-
Large-Scale Contextualised Language Modelling for Norwegian 13 Apr 2021 · 2 repositories · arXiv:2104.06546Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples)
-
Mediators in Determining what Processing BERT Performs First 13 Apr 2021 · 1 repository · arXiv:2104.06400
-
Semantic maps and metrics for science Semantic maps and metrics for science using deep transformer encoders 13 Apr 2021 · 0 repositories · arXiv:2104.05928
-
Understanding Transformers for Bot Detection in Twitter 13 Apr 2021 · 1 repository · arXiv:2104.06182
-
BERT based freedom to operate patent analysis 12 Apr 2021 · 0 repositories · arXiv:2105.00817
-
Fighting the COVID-19 Infodemic with a Holistic BERT Ensemble 12 Apr 2021 · 1 repository · arXiv:2104.05745
-
Fine-Tuning Transformers for Identifying Self-Reporting Potential Cases and Symptoms of COVID-19 in Tweets 12 Apr 2021 · 1 repository · arXiv:2104.05501
-
Learning to Remove: Towards Isotropic Pre-trained BERT Embedding 12 Apr 2021 · 1 repository · arXiv:2104.05274
-
Multilingual Language Models Predict Human Reading Behavior 12 Apr 2021 · 1 repository · arXiv:2104.05433
-
WHOSe Heritage: Classification of UNESCO World Heritage "Outstanding Universal Value" Documents with Soft Labels 12 Apr 2021 · 1 repository · arXiv:2104.05547
-
Does syntax matter? A strong baseline for Aspect-based Sentiment Analysis with RoBERTa 11 Apr 2021 · 1 repository · arXiv:2104.04986
-
Fine-tuning Encoders for Improved Monolingual and Zero-shot Polylingual Neural Topic Modeling 11 Apr 2021 · 1 repository · arXiv:2104.05064
-
Innovative Bert-based Reranking Language Models for Speech Recognition 11 Apr 2021 · 0 repositories · arXiv:2104.04950
-
UniDrop: A Simple yet Effective Technique to Improve Transformer without Extra Cost 11 Apr 2021 · 0 repositories · arXiv:2104.04946
-
Adapting Language Models for Zero-shot Learning by Meta-tuning on Dataset and Prompt Collections 10 Apr 2021 · 1 repository · arXiv:2104.04670
-
MIPT-NSU-UTMN at SemEval-2021 Task 5: Ensembling Learning with Pre-trained Language Models for Toxic Spans Detection 10 Apr 2021 · 1 repository · arXiv:2104.04739
-
Non-autoregressive Transformer-based End-to-end ASR using BERT 10 Apr 2021 · 0 repositories · arXiv:2104.04805
-
ZS-BERT: Towards Zero-Shot Relation Extraction with Attribute Representation Learning 10 Apr 2021 · 1 repository · arXiv:2104.04697
-
KI-BERT: Infusing Knowledge Context for Better Language and Domain Understanding 9 Apr 2021 · 0 repositories · arXiv:2104.08145
-
The Road to Know-Where: An Object-and-Room Informed Sequential BERT for Indoor Vision-Language Navigation 9 Apr 2021 · 1 repository · arXiv:2104.04167
-
Knowledge-Aware Graph-Enhanced GPT-2 for Dialogue State Tracking 9 Apr 2021 · 1 repository · arXiv:2104.04466
-
Text2Chart: A Multi-Staged Chart Generator from Natural Language Text 9 Apr 2021 · 1 repository · arXiv:2104.04584
-
Transformers: "The End of History" for NLP? 9 Apr 2021 · 0 repositories · arXiv:2105.00813
-
Lone Pine at SemEval-2021 Task 5: Fine-Grained Detection of Hate Speech Using BERToxic 8 Apr 2021 · 1 repository · arXiv:2104.03506
-
Probing BERT in Hyperbolic Spaces 8 Apr 2021 · 1 repository · arXiv:2104.03869Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified; the one sample that ran constructed an object rather than computing a result (of 3 harvested samples) · 3 pointer-only (licence)
-
Uppsala NLP at SemEval-2021 Task 2: Multilingual Language Models for Fine-tuning and Feature Extraction in Word-in-Context Disambiguation 8 Apr 2021 · 0 repositories · arXiv:2104.03767
-
Better Neural Machine Translation by Extracting Linguistic Information from BERT 7 Apr 2021 · 1 repository · arXiv:2104.02831
-
Combining Pre-trained Word Embeddings and Linguistic Features for Sequential Metaphor Identification 7 Apr 2021 · 0 repositories · arXiv:2104.03285
-
Interpreting Verbal Metaphors by Paraphrasing 7 Apr 2021 · 0 repositories · arXiv:2104.03391
-
Speak or Chat with Me: End-to-End Spoken Language Understanding System with Flexible Inputs 7 Apr 2021 · 1 repository · arXiv:2104.05752
-
An Empirical Evaluation of Word Embedding Models for Subjectivity Analysis Tasks 6 Apr 2021 · 1 repository
-
CodeTrans: Towards Cracking the Language of Silicon's Code Through Self-Supervised Deep Learning and High Performance Computing 6 Apr 2021 · 1 repository · arXiv:2104.02443