Methods › General › Regularization › Attention Dropout › Papers, page 72
Attention Dropout
Papers archive 2025-07-28
archive papers tagged: 10,892 · with a code link: 4,634 · where Syntology ran a sample: 1,270 (1,043 with a run with no instrument failure, 227 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (1,270 of 10,892 tagged: 1,043 with a run with no instrument failure, 227 where every run was a failure of Syntology's instrument)
Page 72 of 109: papers 7,101 to 7,200 of 10,892, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
A multi-task semi-supervised framework for Text2Graph & Graph2Text 12 Feb 2022 · 1 repository · arXiv:2202.06041
-
Automatic Issue Classifier: A Transfer Learning Framework for Classifying Issue Reports 12 Feb 2022 · 1 repository · arXiv:2202.06149
-
Maximizing Communication Efficiency for Large-scale Training via 0/1 Adam 12 Feb 2022 · 1 repository · arXiv:2202.06009
-
HaT5: Hate Language Identification using Text-to-Text Transfer Transformer 11 Feb 2022 · 0 repositories · arXiv:2202.05690
-
White-Box Attacks on Hate-speech BERT Classifiers in German with Explicit and Implicit Character Level Defense 11 Feb 2022 · 1 repository · arXiv:2202.05778
-
A Multi-task Learning Framework for Product Ranking with BERT 10 Feb 2022 · 0 repositories · arXiv:2202.05317
-
Slovene SuperGLUE Benchmark: Translation and Evaluation 10 Feb 2022 · 0 repositories · arXiv:2202.04994
-
Can Open Domain Question Answering Systems Answer Visual Knowledge Questions? 9 Feb 2022 · 0 repositories · arXiv:2202.04306
-
Social Media as an Instant Source of Feedback on Water Quality 9 Feb 2022 · 0 repositories · arXiv:2202.04462
-
pNLP-Mixer: an Efficient all-MLP Architecture for Language 9 Feb 2022 · 1 repository · arXiv:2202.04350
-
Do Language Models Learn Position-Role Mappings? 8 Feb 2022 · 0 repositories · arXiv:2202.03611
-
HistBERT: A Pre-trained Language Model for Diachronic Lexical Semantic Analysis 8 Feb 2022 · 1 repository · arXiv:2202.03612
-
Logical Reasoning for Task Oriented Dialogue Systems 8 Feb 2022 · 0 repositories · arXiv:2202.04161
-
Semantic features of object concepts generated with GPT-3 8 Feb 2022 · 1 repository · arXiv:2202.03753
-
What are the best systems? New perspectives on NLP Benchmarking 8 Feb 2022 · 1 repository · arXiv:2202.03799
-
Cedille: A large autoregressive French language model 7 Feb 2022 · 1 repository · arXiv:2202.03371
-
OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework 7 Feb 2022 · 4 repositories · arXiv:2202.03052Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Evaluating natural language processing models with generalization metrics that do not need access to any training or testing data 6 Feb 2022 · 1 repository · arXiv:2202.02842
-
Classification on Sentence Embeddings for Legal Assistance 5 Feb 2022 · 0 repositories · arXiv:2202.02639
-
Ethics, Rules of Engagement, and AI: Neural Narrative Mapping Using Large Transformer Language Models 5 Feb 2022 · 0 repositories · arXiv:2202.02647
-
A Benchmark Corpus for the Detection of Automatically Generated Text in Academic Publications 4 Feb 2022 · 1 repository · arXiv:2202.02013
-
StonkBERT: Can Language Models Predict Medium-Run Stock Price Movements? 4 Feb 2022 · 0 repositories · arXiv:2202.02268
-
Temporal Attention for Language Models 4 Feb 2022 · 1 repository · arXiv:2202.02093
-
ASR-Aware End-to-end Neural Diarization 2 Feb 2022 · 0 repositories · arXiv:2202.01286
-
Co-training Improves Prompt-based Learning for Large Language Models 2 Feb 2022 · 1 repository · arXiv:2202.00828Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 10 harvested samples)
-
GatorTron: A Large Clinical Language Model to Unlock Patient Information from Unstructured Electronic Health Records 2 Feb 2022 · 0 repositories · arXiv:2203.03540
-
L3Cube-MahaCorpus and MahaBERT: Marathi Monolingual Corpus, Marathi BERT Language Models, and Resources 2 Feb 2022 · 1 repository · arXiv:2202.01159
-
RescoreBERT: Discriminative Speech Recognition Rescoring with BERT 2 Feb 2022 · 0 repositories · arXiv:2202.01094
-
Robust Training of Neural Networks Using Scale Invariant Architectures 2 Feb 2022 · 0 repositories · arXiv:2202.00980
-
An Adaptive Deep Clustering Pipeline to Inform Text Labeling at Scale 1 Feb 2022 · 0 repositories · arXiv:2202.01211
-
A Semi-Supervised Deep Clustering Pipeline for Mining Intentions From Texts 1 Feb 2022 · 0 repositories · arXiv:2202.00802
-
Improving BERT-based Query-by-Document Retrieval with Multi-Task Optimization 1 Feb 2022 · 0 repositories · arXiv:2202.00373
-
Transformer-based Models of Text Normalization for Speech Applications 1 Feb 2022 · 0 repositories · arXiv:2202.00153
-
BOAT: Bilateral Local Attention Vision Transformer 31 Jan 2022 · 1 repository · arXiv:2201.13027
-
Memory-Efficient Backpropagation through Large Linear Layers 31 Jan 2022 · 2 repositories · arXiv:2201.13195
-
A Frustratingly Simple Approach for End-to-End Image Captioning 30 Jan 2022 · 0 repositories · arXiv:2201.12723
-
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models 28 Jan 2022 · 19 repositories · arXiv:2201.11903Syntology 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples) · 1 pointer-only (licence)
-
Electra: Conditional Generative Model based Predicate-Aware Query Approximation 28 Jan 2022 · 0 repositories · arXiv:2201.12420
-
Describing Differences between Text Distributions with Natural Language 28 Jan 2022 · 1 repository · arXiv:2201.12323Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Clinical-Longformer and Clinical-BigBird: Transformers for long clinical sequences 27 Jan 2022 · 1 repository · arXiv:2201.11838
-
Going Extreme: Comparative Analysis of Hate Speech in Parler and Gab 27 Jan 2022 · 1 repository · arXiv:2201.11770
-
Grad2Task: Improved Few-shot Text Classification Using Gradients for Task Representation 27 Jan 2022 · 1 repository · arXiv:2201.11576
-
A Comprehensive Study of Image Classification Model Sensitivity to Foregrounds, Backgrounds, and Visual Attributes 26 Jan 2022 · 1 repository · arXiv:2201.10766Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
DiscoScore: Evaluating Text Generation with BERT and Discourse Coherence 26 Jan 2022 · 1 repository · arXiv:2201.11176
-
DNNFuser: Generative Pre-Trained Transformer as a Generalized Mapper for Layer Fusion in DNN Accelerators 26 Jan 2022 · 0 repositories · arXiv:2201.11218
-
FiNCAT: Financial Numeral Claim Analysis Tool 26 Jan 2022 · 1 repository · arXiv:2202.00631
-
Neural Grapheme-to-Phoneme Conversion with Pre-trained Grapheme Models 26 Jan 2022 · 1 repository · arXiv:2201.10716
-
Self-supervised 3D Semantic Representation Learning for Vision-and-Language Navigation 26 Jan 2022 · 0 repositories · arXiv:2201.10788
-
Synchromesh: Reliable code generation from pre-trained language models 26 Jan 2022 · 2 repositories · arXiv:2201.11227Syntology 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
BERTHA: Video Captioning Evaluation Via Transfer-Learned Human Assessment 25 Jan 2022 · 1 repository · arXiv:2201.10243
-
Improving non-autoregressive end-to-end speech recognition with pre-trained acoustic and language models 25 Jan 2022 · 0 repositories · arXiv:2201.10103
-
Pre-Trained Language Transformers are Universal Image Classifiers 25 Jan 2022 · 0 repositories · arXiv:2201.10182
-
Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data Selection 25 Jan 2022 · 0 repositories · arXiv:2201.10474
-
Emotion-based Modeling of Mental Disorders on Social Media 24 Jan 2022 · 0 repositories · arXiv:2201.09451
-
Polyphone disambiguation and accent prediction using pre-trained language models in Japanese TTS front-end 24 Jan 2022 · 0 repositories · arXiv:2201.09427
-
Synthetic Books 24 Jan 2022 · 0 repositories · arXiv:2201.09518
-
Unified Multimodal Punctuation Restoration Framework for Mixed-Modality Corpus 24 Jan 2022 · 1 repository · arXiv:2202.00468
-
A Large and Diverse Arabic Corpus for Language Modeling 23 Jan 2022 · 0 repositories · arXiv:2201.09227
-
An Application of Pseudo-Log-Likelihoods to Natural Language Scoring 23 Jan 2022 · 0 repositories · arXiv:2201.09377
-
A Comparative Study on Language Models for Task-Oriented Dialogue Systems 21 Jan 2022 · 1 repository · arXiv:2201.08687
-
Black-box Prompt Learning for Pre-trained Language Models 21 Jan 2022 · 1 repository · arXiv:2201.08531Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Dual Contrastive Learning: Text Classification via Label-Aware Data Augmentation 21 Jan 2022 · 2 repositories · arXiv:2201.08702
-
Less is Less: When Are Snippets Insufficient for Human vs Machine Relevance Estimation? 21 Jan 2022 · 0 repositories · arXiv:2201.08721
-
Cheating Automatic Short Answer Grading: On the Adversarial Usage of Adjectives and Adverbs 20 Jan 2022 · 1 repository · arXiv:2201.08318
-
Sentiment Analysis: Predicting Yelp Scores 20 Jan 2022 · 0 repositories · arXiv:2201.07999
-
TerViT: An Efficient Ternary Vision Transformer 20 Jan 2022 · 0 repositories · arXiv:2201.08050
-
Transfer Learning Approaches for Building Cross-Language Dense Retrieval Models 20 Jan 2022 · 1 repository · arXiv:2201.08471Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
GAP-Gen: Guided Automatic Python Code Generation 19 Jan 2022 · 1 repository · arXiv:2201.08810
-
Near-Optimal Sparse Allreduce for Distributed Deep Learning 19 Jan 2022 · 1 repository · arXiv:2201.07598
-
Q-ViT: Fully Differentiable Quantization for Vision Transformer 19 Jan 2022 · 1 repository · arXiv:2201.07703Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 2 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
TourBERT: A pretrained language model for the tourism industry 19 Jan 2022 · 0 repositories · arXiv:2201.07449
-
CoAuthor: Designing a Human-AI Collaborative Writing Dataset for Exploring Language Model Capabilities 18 Jan 2022 · 1 repository · arXiv:2201.06796
-
Hierarchical Neural Network Approaches for Long Document Classification 18 Jan 2022 · 0 repositories · arXiv:2201.06774
-
BERT vs ALBERT explained 17 Jan 2022 · 0 repositories
-
MuLVE, A Multi-Language Vocabulary Evaluation Data Set 17 Jan 2022 · 0 repositories · arXiv:2201.06286
-
Unintended Bias in Language Model-driven Conversational Recommendation 17 Jan 2022 · 0 repositories · arXiv:2201.06224
-
A Balanced Data Approach for Evaluating Cross-Lingual Transfer: Mapping the Linguistic Blood Bank 16 Jan 2022 · 0 repositories
-
A Multi-Granularity Opinion Summarization Method 16 Jan 2022 · 0 repositories
-
A Study of Pre-trained Language Models for Analogy Generation 16 Jan 2022 · 0 repositories
-
A Study of the Attention Abnormality in Trojaned BERTs 16 Jan 2022 · 0 repositories
-
AllWOZ: Towards Multilingual Task-Oriented Dialog Systems for All 16 Jan 2022 · 0 repositories
-
An Exploitation of Heterogeneous Graph Neural Network for Extractive Long Document Summarization 16 Jan 2022 · 0 repositories
-
Applying SoftTriple Loss for Supervised Language Model Fine Tuning 16 Jan 2022 · 0 repositories
-
AraBART: a Pretrained Arabic Sequence-to-Sequence Model for Abstractive Summarization 16 Jan 2022 · 0 repositories
-
Are Pretrained Multilingual Models Equally Fair Across Languages? 16 Jan 2022 · 0 repositories
-
Auto-regressive Text Generation with Pre-Trained Language Models: An Empirical Study on Question-type Short Text Generation 16 Jan 2022 · 0 repositories
-
AutoAttention: Automatic Attention Head Selection Through Differentiable Pruning 16 Jan 2022 · 0 repositories
-
Bridge the Gap Between CV and NLP! A Gradient-based Textual Adversarial Attack Framework 16 Jan 2022 · 0 repositories
-
Can BERT Conduct Logical Reasoning? On the Difficulty of Learning to Reason from Data 16 Jan 2022 · 0 repositories
-
Context-Aware Prompt: Customize A Unique Prompt For Each Input 16 Jan 2022 · 0 repositories
-
DECK: Behavioral Tests to Improve Interpretability and Generalizability of BERT Models Detecting Depression from Text 16 Jan 2022 · 0 repositories
-
Divide and Conquer: Text Semantic Matching with Disentangled Keywords and Intents 16 Jan 2022 · 0 repositories
-
Do BERTs Learn to Use Browser User Interface? Exploring Multi-Step Tasks with Unified Vision-and-Language BERTs 16 Jan 2022 · 0 repositories
-
Efficient Hierarchical Domain Adaptation for Pretrained Language Models 16 Jan 2022 · 0 repositories
-
Efficient Zero-Shot Semantic Parsing with Paraphrasing from Pretrained Language Models 16 Jan 2022 · 0 repositories
-
EiCi: A New Method of Dynamic Embedding Incorporating Contextual Information in Chinese NER 16 Jan 2022 · 0 repositories
-
Elastic Weight Consolidation for Reduction of Catastrophic Forgetting in GPT-2 16 Jan 2022 · 0 repositories
-
Event Detection via Derangement Reading Comprehension 16 Jan 2022 · 0 repositories
-
Experiments with adversarial attacks on text genres 16 Jan 2022 · 0 repositories
-
Exploring Example Selection for Few-shot Text-to-SQL Semantic Parsing 16 Jan 2022 · 0 repositories