Methods › General › Regularization › Attention Dropout › Papers, page 78
Attention Dropout
Papers archive 2025-07-28
archive papers tagged: 10,892 · with a code link: 4,634 · where Syntology ran a sample: 1,270 (1,043 with a run with no instrument failure, 227 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (1,270 of 10,892 tagged: 1,043 with a run with no instrument failure, 227 where every run was a failure of Syntology's instrument)
Page 78 of 109: papers 7,701 to 7,800 of 10,892, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Token Pooling in Vision Transformers 8 Oct 2021 · 0 repositories · arXiv:2110.03860
-
A Comparative Study of Transformer-Based Language Models on Extractive Question Answering 7 Oct 2021 · 0 repositories · arXiv:2110.03142
-
Attention is All You Need? Good Embeddings with Statistics are enough:Large Scale Audio Understanding without Transformers/ Convolutions/ BERTs/ Mixers/ Attention/ RNNs or .... 7 Oct 2021 · 0 repositories · arXiv:2110.03183
-
Cross-Language Learning for Entity Matching 7 Oct 2021 · 1 repository · arXiv:2110.03338
-
Universality of Winning Tickets: A Renormalization Group Perspective 7 Oct 2021 · 0 repositories · arXiv:2110.03210
-
8-bit Optimizers via Block-wise Quantization 6 Oct 2021 · 3 repositories · arXiv:2110.02861
-
NUS-IDS at FinCausal 2021: Dependency Tree in Graph Neural Network for Better Cause-Effect Span Detection 6 Oct 2021 · 1 repository · arXiv:2110.02991
-
PoNet: Pooling Network for Efficient Token Mixing in Long Sequences 6 Oct 2021 · 1 repository · arXiv:2110.02442Syntology official: no sample here; runs from other or unrecorded repositories · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
Pretrained Transformers for Offensive Language Identification in Tanglish 6 Oct 2021 · 1 repository · arXiv:2110.02852
-
Analyzing the Impact of COVID-19 on Economy from the Perspective of Users Reviews 5 Oct 2021 · 0 repositories · arXiv:2110.02198
-
ASR Rescoring and Confidence Estimation with ELECTRA 5 Oct 2021 · 0 repositories · arXiv:2110.01857
-
BERT Attends the Conversation: Improving Low-Resource Conversational ASR 5 Oct 2021 · 1 repository · arXiv:2110.02267
-
DistilHuBERT: Speech Representation Learning by Layer-wise Distillation of Hidden-unit BERT 5 Oct 2021 · 1 repository · arXiv:2110.01900
-
FoodChem: A food-chemical relation extraction model 5 Oct 2021 · 1 repository · arXiv:2110.02019
-
Learning Sense-Specific Static Embeddings using Contextualised Word Embeddings as a Proxy 5 Oct 2021 · 0 repositories · arXiv:2110.02204
-
Leveraging the Inductive Bias of Large Language Models for Abstract Textual Reasoning 5 Oct 2021 · 0 repositories · arXiv:2110.02370
-
MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer 5 Oct 2021 · 31 repositories · arXiv:2110.02178Syntology official: no sample here; runs from other or unrecorded repositories · 53 ran (of which 31 constructed an object rather than computing a result; 44 with no instrument failure: 1 honoured, 0 violated, 43 with no contract checked; 9 where Syntology's instrument failed) · 15 unverified (of 68 harvested samples) · 18 pointer-only (licence)
-
ur-iw-hnt at GermEval 2021: An Ensembling Strategy with Multiple BERT Models 5 Oct 2021 · 0 repositories · arXiv:2110.02042
-
Word Acquisition in Neural Language Models 5 Oct 2021 · 1 repository · arXiv:2110.02406
-
DeepA2: A Modular Framework for Deep Argument Analysis with Pretrained Neural Text2Text Language Models 4 Oct 2021 · 1 repository · arXiv:2110.01509
-
Exploiting Pre-Trained ASR Models for Alzheimer's Disease Recognition Through Spontaneous Speech 4 Oct 2021 · 0 repositories · arXiv:2110.01493
-
JuriBERT: A Masked-Language Model Adaptation for French Legal Text 4 Oct 2021 · 1 repository · arXiv:2110.01485
-
Perhaps PTLMs Should Go to School -- A Task to Assess Open Book and Closed Book QA 4 Oct 2021 · 0 repositories · arXiv:2110.01552
-
Adversarial Examples Generation for Reducing Implicit Gender Bias in Pre-trained Models 3 Oct 2021 · 0 repositories · arXiv:2110.01094
-
Unsupervised paradigm for information extraction from transcripts using BERT 3 Oct 2021 · 0 repositories · arXiv:2110.00949
-
Artificial intelligence for Sustainable Energy: A Contextual Topic Modeling and Content Analysis 2 Oct 2021 · 0 repositories · arXiv:2110.00828
-
Swiss-Judgment-Prediction: A Multilingual Legal Judgment Prediction Benchmark 2 Oct 2021 · 1 repository · arXiv:2110.00806
-
BERT4GCN: Using BERT Intermediate Layers to Augment GCN for Aspect-based Sentiment Classification 1 Oct 2021 · 0 repositories · arXiv:2110.00171
-
Improving Punctuation Restoration for Speech Transcripts via External Data 1 Oct 2021 · 0 repositories · arXiv:2110.00560
-
Low Frequency Names Exhibit Bias and Overfitting in Contextualizing Language Models 1 Oct 2021 · 0 repositories · arXiv:2110.00672
-
Span Labeling Approach for Vietnamese and Chinese Word Segmentation 1 Oct 2021 · 0 repositories · arXiv:2110.00156
-
Unpacking the Interdependent Systems of Discrimination: Ableist Bias in NLP Systems through an Intersectional Lens 1 Oct 2021 · 0 repositories · arXiv:2110.00521
-
BERT got a Date: Introducing Transformers to Temporal Tagging 30 Sep 2021 · 1 repository · arXiv:2109.14927
-
COVID-19 Fake News Detection Using Bidirectional Encoder Representations from Transformers Based Models 30 Sep 2021 · 1 repository · arXiv:2109.14816
-
Prose2Poem: The Blessing of Transformers in Translating Prose to Persian Poetry 30 Sep 2021 · 1 repository · arXiv:2109.14934
-
Are BERT Families Zero-Shot Learners? A Study on Their Potential and Limitations 29 Sep 2021 · 0 repositories
-
Collaborative Storytelling with Human Actors and AI Narrators 29 Sep 2021 · 0 repositories · arXiv:2109.14728
-
Compressing Transformer-Based Sequence to Sequence Models With Pre-trained Autoencoders for Text Summarization 29 Sep 2021 · 0 repositories
-
Contrastive Pre-training for Zero-Shot Information Retrieval 29 Sep 2021 · 0 repositories
-
Cross-Architecture Distillation Using Bidirectional CMOW Embeddings 29 Sep 2021 · 0 repositories
-
Efficient Packing: Towards 2x NLP Speed-Up without Loss of Accuracy for BERT 29 Sep 2021 · 0 repositories
-
Embedding models through the lens of Stable Coloring 29 Sep 2021 · 0 repositories
-
ERNIE-SPARSE: Robust Efficient Transformer Through Hierarchically Unifying Isolated Information 29 Sep 2021 · 0 repositories
-
Gradient Broadcast Adaptation: Defending against the backdoor attack in pre-trained models 29 Sep 2021 · 0 repositories
-
Group-based Interleaved Pipeline Parallelism for Large-scale DNN Training 29 Sep 2021 · 1 repository
-
Hierarchical Character Tagger for Short Text Spelling Error Correction 29 Sep 2021 · 0 repositories · arXiv:2109.14259
-
How does BERT address polysemy of Korean adverbial postpositions -ey, -eyse, and -(u)lo? 29 Sep 2021 · 0 repositories
-
Illiterate DALL·E Learns to Compose 29 Sep 2021 · 0 repositories
-
Improving Sentiment Classification Using 0-Shot Generated Labels for Custom Transformer Embeddings 29 Sep 2021 · 0 repositories
-
In defense of dual-encoders for neural ranking 29 Sep 2021 · 0 repositories
-
Language Model Pre-training Improves Generalization in Policy Learning 29 Sep 2021 · 0 repositories
-
Learning Rate Grafting: Transferability of Optimizer Tuning 29 Sep 2021 · 0 repositories
-
Learning Visual-Linguistic Adequacy, Fidelity, and Fluency for Novel Object Captioning 29 Sep 2021 · 0 repositories
-
MaiT: integrating spatial locality into image transformers with attention masks 29 Sep 2021 · 1 repository
-
Mapping Language Models to Grounded Conceptual Spaces 29 Sep 2021 · 0 repositories
-
Modeling label correlations implicitly through latent label encodings for multi-label text classification 29 Sep 2021 · 0 repositories
-
Offline Reinforcement Learning for Large Scale Language Action Spaces 29 Sep 2021 · 0 repositories
-
Rank4Class: Examining Multiclass Classification through the Lens of Learning to Rank 29 Sep 2021 · 0 repositories
-
Robot Intent Recognition Method Based on State Grid Business Office 29 Sep 2021 · 0 repositories
-
ScaLA: Speeding-Up Fine-tuning of Pre-trained Transformer Networks via Efficient and Scalable Adversarial Perturbation 29 Sep 2021 · 0 repositories
-
Scale Efficiently: Insights from Pretraining and Finetuning Transformers 29 Sep 2021 · 0 repositories
-
SeqPATE: Differentially Private Text Generation via Knowledge Distillation 29 Sep 2021 · 0 repositories
-
Specialized Transformers: Faster, Smaller and more Accurate NLP Models 29 Sep 2021 · 0 repositories
-
Training sequence labeling models using prior knowledge 29 Sep 2021 · 0 repositories
-
TransSlowDown: Efficiency Attacks on Neural Machine Translation Systems 29 Sep 2021 · 0 repositories
-
How Different Text-preprocessing Techniques Using The BERT Model Affect The Gender Profiling of Authors 28 Sep 2021 · 0 repositories · arXiv:2109.13890
-
RAFT: A Real-World Few-Shot Text Classification Benchmark 28 Sep 2021 · 1 repository · arXiv:2109.14076
-
Effective Use of Graph Convolution Network and Contextual Sub-Tree forCommodity News Event Extraction 27 Sep 2021 · 1 repository · arXiv:2109.12781
-
Patterns of Lexical Ambiguity in Contextualised Language Models 27 Sep 2021 · 0 repositories · arXiv:2109.13032
-
TURINGBENCH: A Benchmark Environment for Turing Test in the Age of Neural Text Generation 27 Sep 2021 · 3 repositories · arXiv:2109.13296Syntology 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
Understanding and Overcoming the Challenges of Efficient Transformer Quantization 27 Sep 2021 · 1 repository · arXiv:2109.12948Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Improving Question Answering Performance Using Knowledge Distillation and Active Learning 26 Sep 2021 · 1 repository · arXiv:2109.12662Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
On the Prunability of Attention Heads in Multilingual BERT 26 Sep 2021 · 0 repositories · arXiv:2109.12683
-
Finetuning Transformer Models to Build ASAG System 25 Sep 2021 · 0 repositories · arXiv:2109.12300
-
AES Systems Are Both Overstable And Oversensitive: Explaining Why And Proposing Defenses 24 Sep 2021 · 0 repositories · arXiv:2109.11728
-
Dense Contrastive Visual-Linguistic Pretraining 24 Sep 2021 · 0 repositories · arXiv:2109.11778
-
Lacking the embedding of a word? Look it up into a traditional dictionary 24 Sep 2021 · 0 repositories · arXiv:2109.11763
-
Robustness and Sensitivity of BERT Models Predicting Alzheimer's Disease from Text 24 Sep 2021 · 0 repositories · arXiv:2109.11888
-
Breaking BERT: Understanding its Vulnerabilities for Named Entity Recognition through Adversarial Attack 23 Sep 2021 · 1 repository · arXiv:2109.11308
-
Putting Words in BERT's Mouth: Navigating Contextualized Vector Spaces with Pseudowords 23 Sep 2021 · 1 repository · arXiv:2109.11491
-
Alzheimers Dementia Detection using Acoustic & Linguistic features and Pre-Trained BERT 22 Sep 2021 · 0 repositories · arXiv:2109.11010
-
DialogueBERT: A Self-Supervised Learning based Dialogue Pre-training Encoder 22 Sep 2021 · 0 repositories · arXiv:2109.10480
-
Language Models as Recommender Systems: Evaluations and Limitations 22 Sep 2021 · 0 repositories
-
Predicting Efficiency/Effectiveness Trade-offs for Dense vs. Sparse Retrieval Strategy Selection 22 Sep 2021 · 0 repositories · arXiv:2109.10739
-
Recursively Summarizing Books with Human Feedback 22 Sep 2021 · 0 repositories · arXiv:2109.10862
-
Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers 22 Sep 2021 · 3 repositories · arXiv:2109.10686Syntology community repositories only · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 4 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 12 harvested samples) · 1 pointer-only (licence)
-
Unsupervised Contextualized Document Representation 22 Sep 2021 · 1 repository · arXiv:2109.10509
-
A Comprehensive Review on Summarizing Financial News Using Deep Learning 21 Sep 2021 · 0 repositories · arXiv:2109.10118
-
BERTweetFR : Domain Adaptation of Pre-Trained Language Models for French Tweets 21 Sep 2021 · 0 repositories · arXiv:2109.10234
-
InvBERT: Reconstructing Text from Contextualized Word Embeddings by inverting the BERT pipeline 21 Sep 2021 · 0 repositories · arXiv:2109.10104
-
Multi-Task Learning with Sentiment, Emotion, and Target Detection to Recognize Hate Speech and Offensive Language 21 Sep 2021 · 0 repositories · arXiv:2109.10255
-
Representation Learning for Short Text Clustering 21 Sep 2021 · 0 repositories · arXiv:2109.09894
-
A Plug-and-Play Method for Controlled Text Generation 20 Sep 2021 · 1 repository · arXiv:2109.09707
-
BERT Cannot Align Characters 20 Sep 2021 · 0 repositories · arXiv:2109.09700
-
BERT Has Uncommon Sense: Similarity Ranking for Word Sense BERTology 20 Sep 2021 · 1 repository · arXiv:2109.09780
-
Model Bias in NLP -- Application to Hate Speech Classification using transfer learning techniques 20 Sep 2021 · 0 repositories · arXiv:2109.09725
-
MirrorWiC: On Eliciting Word-in-Context Representations from Pretrained Language Models 19 Sep 2021 · 1 repository · arXiv:2109.09237
-
Towards Zero-Label Language Learning 19 Sep 2021 · 0 repositories · arXiv:2109.09193
-
Wav-BERT: Cooperative Acoustic and Linguistic Representation Learning for Low-Resource Speech Recognition 19 Sep 2021 · 0 repositories · arXiv:2109.09161
-
What BERT Based Language Models Learn in Spoken Transcripts: An Empirical Study 19 Sep 2021 · 0 repositories · arXiv:2109.09105