Methods › General › Regularization › Attention Dropout › Papers, page 70
Attention Dropout
Papers archive 2025-07-28
archive papers tagged: 10,892 · with a code link: 4,634 · where Syntology ran a sample: 1,270 (1,043 with a run with no instrument failure, 227 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (1,270 of 10,892 tagged: 1,043 with a run with no instrument failure, 227 where every run was a failure of Syntology's instrument)
Page 70 of 109: papers 6,901 to 7,000 of 10,892, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Building Markovian Generative Architectures over Pretrained LM Backbones for Efficient Task-Oriented Dialog Systems 13 Apr 2022 · 2 repositories · arXiv:2204.06452
-
TangoBERT: Reducing Inference Cost by using Cascaded Architecture 13 Apr 2022 · 0 repositories · arXiv:2204.06271
-
Team ÚFAL at CMCL 2022 Shared Task: Figuring out the correct recipe for predicting Eye-Tracking features using Pretrained Language Models 11 Apr 2022 · 0 repositories · arXiv:2204.04998
-
Tokenwise Contrastive Pretraining for Finer Speech-to-BERT Alignment in End-to-End Speech-to-Intent Systems 11 Apr 2022 · 0 repositories · arXiv:2204.05188
-
Towards Generalizable Semantic Product Search by Text Similarity Pre-training on Search Click Logs 11 Apr 2022 · 0 repositories · arXiv:2204.05231
-
Uniform Complexity for Text Generation 11 Apr 2022 · 1 repository · arXiv:2204.05185
-
Fake news detection using parallel BERT deep neural networks 10 Apr 2022 · 0 repositories · arXiv:2204.04793
-
Few-Shot Cross-lingual Transfer for Coarse-grained De-identification of Code-Mixed Clinical Texts 10 Apr 2022 · 1 repository · arXiv:2204.04775
-
Pushing on Personality Detection from Verbal Behavior: A Transformer Meets Text Contours of Psycholinguistic Features 10 Apr 2022 · 0 repositories · arXiv:2204.04629
-
FoundationLayerNorm: Scaling BERT and GPT to 1,000 Layers 9 Apr 2022 · 0 repositories · arXiv:2204.04477
-
Modeling Multi-Granularity Hierarchical Features for Relation Extraction 9 Apr 2022 · 1 repository · arXiv:2204.04437Syntology official: no sample here; runs from other or unrecorded repositories · 2 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Are We Really Making Much Progress in Text Classification? A Comparative Review 8 Apr 2022 · 1 repository · arXiv:2204.03954
-
Contextual Representation Learning beyond Masked Language Modeling 8 Apr 2022 · 1 repository · arXiv:2204.04163
-
Enhance Incomplete Utterance Restoration by Joint Learning Token Extraction and Text Generation 8 Apr 2022 · 1 repository · arXiv:2204.03958
-
Infusing Knowledge from Wikipedia to Enhance Stance Detection 8 Apr 2022 · 2 repositories · arXiv:2204.03839
-
MMTAfrica: Multilingual Machine Translation for African Languages 8 Apr 2022 · 1 repository · arXiv:2204.04306
-
Towards Understanding Large-Scale Discourse Structures in Pre-Trained and Fine-Tuned Language Models 8 Apr 2022 · 0 repositories · arXiv:2204.04289
-
Accelerating Attention through Gradient-Based Learned Runtime Pruning 7 Apr 2022 · 0 repositories · arXiv:2204.03227
-
Autoencoding Language Model Based Ensemble Learning for Commonsense Validation and Explanation 7 Apr 2022 · 0 repositories · arXiv:2204.03324
-
BERTuit: Understanding Spanish language in Twitter through a native transformer 7 Apr 2022 · 0 repositories · arXiv:2204.03465
-
CoCoSoDa: Effective Contrastive Learning for Code Search 7 Apr 2022 · 0 repositories · arXiv:2204.03293
-
PALBERT: Teaching ALBERT to Ponder 7 Apr 2022 · 1 repository · arXiv:2204.03276Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; the one sample that ran constructed an object rather than computing a result (of 2 harvested samples)
-
Pretraining Text Encoders with Adversarial Mixture of Training Signal Generators 7 Apr 2022 · 1 repository · arXiv:2204.03243Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 8 harvested samples) · 2 pointer-only (licence)
-
Testing the limits of natural language models for predicting human language judgments 7 Apr 2022 · 1 repository · arXiv:2204.03592
-
ByT5 model for massively multilingual grapheme-to-phoneme conversion 6 Apr 2022 · 1 repository · arXiv:2204.03067
-
drsphelps at SemEval-2022 Task 2: Learning idiom representations using BERTRAM 6 Apr 2022 · 0 repositories · arXiv:2204.02821
-
Knowledge Infused Decoding 6 Apr 2022 · 1 repository · arXiv:2204.03084Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Paying More Attention to Self-attention: Improving Pre-trained Language Models via Attention Guiding 6 Apr 2022 · 0 repositories · arXiv:2204.02922
-
Using Synthetic Data for Conversational Response Generation in Low-resource Settings 6 Apr 2022 · 0 repositories · arXiv:2204.02653
-
Abstractive summarization of hospitalisation histories with transformer networks 5 Apr 2022 · 0 repositories · arXiv:2204.02208
-
An Exploratory Study on Code Attention in BERT 5 Apr 2022 · 0 repositories · arXiv:2204.10200
-
Data Augmentation for Intent Classification with Off-the-shelf Large Language Models 5 Apr 2022 · 1 repository · arXiv:2204.01959Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples) · 2 pointer-only (licence)
-
How Different are Pre-trained Transformers for Text Ranking? 5 Apr 2022 · 1 repository · arXiv:2204.07233
-
Multilinguals at SemEval-2022 Task 11: Transformer Based Architecture for Complex NER 5 Apr 2022 · 1 repository · arXiv:2204.02173
-
POS-BERT: Point Cloud One-Stage BERT Pre-Training 3 Apr 2022 · 1 repository · arXiv:2204.00989
-
BERT-Assisted Semantic Annotation Correction for Emotion-Related Questions 2 Apr 2022 · 1 repository · arXiv:2204.00916
-
Efficient comparison of sentence embeddings 2 Apr 2022 · 0 repositories · arXiv:2204.00820
-
Analyzing how BERT performs entity matching 1 Apr 2022 · 1 repository
-
CharacterBERT and Self-Teaching for Improving the Robustness of Dense Retrievers on Queries with Typos 1 Apr 2022 · 1 repository · arXiv:2204.00716
-
Cyberbullying detection across social media platforms via platform-aware adversarial encoding 1 Apr 2022 · 0 repositories · arXiv:2204.00334
-
Effect and Analysis of Large-scale Language Model Rescoring on Competitive ASR Systems 1 Apr 2022 · 0 repositories · arXiv:2204.00212
-
Feature Structure Distillation with Centered Kernel Alignment in BERT Transferring 1 Apr 2022 · 1 repository · arXiv:2204.08922
-
Monarch: Expressive Structured Matrices for Efficient and Accurate Training 1 Apr 2022 · 2 repositories · arXiv:2204.00595Syntology community repositories only · 15 ran (of which 0 constructed an object rather than computing a result; 14 with no instrument failure: 0 honoured, 1 violated, 13 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 15 harvested samples)
-
Syntax-informed Question Answering with Heterogeneous Graph Transformer 1 Apr 2022 · 0 repositories · arXiv:2204.09655
-
A Baseline Readability Model for Cebuano 31 Mar 2022 · 1 repository · arXiv:2203.17225
-
A Character-level Span-based Model for Mandarin Prosodic Structure Prediction 31 Mar 2022 · 1 repository · arXiv:2203.16922
-
CatIss: An Intelligent Tool for Categorizing Issues Reports using Transformers 31 Mar 2022 · 1 repository · arXiv:2203.17196
-
CREATE: A Benchmark for Chinese Short Video Retrieval and Title Generation 31 Mar 2022 · 0 repositories · arXiv:2203.16763
-
ESGBERT: Language Model to Help with Classification Tasks Related to Companies Environmental, Social, and Governance Practices 31 Mar 2022 · 0 repositories · arXiv:2203.16788
-
Generative Pre-Trained Transformers for Biologically Inspired Design 31 Mar 2022 · 0 repositories · arXiv:2204.09714
-
Leveraging pre-trained language models for conversational information seeking from text 31 Mar 2022 · 0 repositories · arXiv:2204.03542
-
Mixed-Phoneme BERT: Improving BERT with Mixed Phoneme and Sup-Phoneme Representations for Text to Speech 31 Mar 2022 · 0 repositories · arXiv:2203.17190
-
Incorporating Dynamic Semantics into Pre-Trained Language Model for Aspect-based Sentiment Analysis 30 Mar 2022 · 0 repositories · arXiv:2203.16369
-
Transformer Language Models without Positional Encodings Still Learn Positional Information 30 Mar 2022 · 1 repository · arXiv:2203.16634Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
A Fast Post-Training Pruning Framework for Transformers 29 Mar 2022 · 2 repositories · arXiv:2204.09656Syntology official (archive's flag): 1 ran · 3 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Improving Persian Relation Extraction Models by Data Augmentation 29 Mar 2022 · 0 repositories · arXiv:2203.15323
-
LinkBERT: Pretraining Language Models with Document Links 29 Mar 2022 · 1 repository · arXiv:2203.15827Syntology official: harvested, nothing ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 11 unverified (of 14 harvested samples)
-
mc-BEiT: Multi-choice Discretization for Image BERT Pre-training 29 Mar 2022 · 1 repository · arXiv:2203.15371Syntology official (archive's flag): 3 ran · 3 ran (of which 3 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified; every one of the 3 samples that ran constructed an object rather than computing a result (of 5 harvested samples) · 5 pointer-only (licence)
-
Training Compute-Optimal Large Language Models 29 Mar 2022 · 2 repositories · arXiv:2203.15556Syntology 8 ran (of which 3 constructed an object rather than computing a result; 5 with no instrument failure: 2 honoured, 0 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (of 11 harvested samples) · 4 pointer-only (licence)
-
ANNA: Enhanced Language Representation for Question Answering 28 Mar 2022 · 0 repositories · arXiv:2203.14507
-
Automated Progressive Learning for Efficient Training of Vision Transformers 28 Mar 2022 · 2 repositories · arXiv:2203.14509Syntology official (archive's flag): 7 ran · 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 1 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (of 13 harvested samples) · 4 pointer-only (licence)
-
Continuous Metric Learning For Transferable Speech Emotion Recognition and Embedding Across Low-resource Languages 28 Mar 2022 · 0 repositories · arXiv:2203.14867
-
Hierarchical Transformer Model for Scientific Named Entity Recognition 28 Mar 2022 · 1 repository · arXiv:2203.14710
-
UTSA NLP at SemEval-2022 Task 4: An Exploration of Simple Ensembles of Transformers, Convolutional, and Recurrent Neural Networks 28 Mar 2022 · 0 repositories · arXiv:2203.14920
-
Example-based Hypernetworks for Out-of-Distribution Generalization 27 Mar 2022 · 1 repository · arXiv:2203.14276
-
Pyramid-BERT: Reducing Complexity via Successive Core-set based Token Selection 27 Mar 2022 · 0 repositories · arXiv:2203.14380
-
StruBERT: Structure-aware BERT for Table Search and Matching 27 Mar 2022 · 1 repository · arXiv:2203.14278Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 1 honoured, 0 violated, 11 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 13 harvested samples)
-
Autoregressive Linguistic Steganography Based on BERT and Consistency Coding 26 Mar 2022 · 0 repositories · arXiv:2203.13972
-
L3Cube-MahaHate: A Tweet-based Marathi Hate Speech Detection Dataset and BERT models 25 Mar 2022 · 1 repository · arXiv:2203.13778
-
MKQ-BERT: Quantized BERT with 4-bits Weights and Activations 25 Mar 2022 · 0 repositories · arXiv:2203.13483
-
Predicting Clinical Intent from Free Text Electronic Health Records 25 Mar 2022 · 0 repositories · arXiv:2204.09594
-
Bailando: 3D Dance Generation by Actor-Critic GPT with Choreographic Memory 24 Mar 2022 · 1 repository · arXiv:2203.13055Syntology official (archive's flag): 4 ran · 4 ran (of which 3 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Ensembling and Knowledge Distilling of Large Sequence Taggers for Grammatical Error Correction 24 Mar 2022 · 1 repository · arXiv:2203.13064
-
mcBERT: Momentum Contrastive Learning with BERT for Zero-Shot Slot Filling 24 Mar 2022 · 0 repositories · arXiv:2203.12940
-
minicons: Enabling Flexible Behavioral and Representational Analyses of Transformer Language Models 24 Mar 2022 · 1 repository · arXiv:2203.13112
-
Mono vs Multilingual BERT: A Case Study in Hindi and Marathi Named Entity Recognition 24 Mar 2022 · 0 repositories · arXiv:2203.12907
-
Recommendation as Language Processing (RLP): A Unified Pretrain, Personalized Prompt & Predict Paradigm (P5) 24 Mar 2022 · 2 repositories · arXiv:2203.13366
-
Token Dropping for Efficient BERT Pretraining 24 Mar 2022 · 0 repositories · arXiv:2203.13240
-
Adversarial Training for Improving Model Robustness? Look at Both Prediction and Interpretation 23 Mar 2022 · 1 repository · arXiv:2203.12709Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
ERNIE-SPARSE: Learning Hierarchical Efficient Transformer Through Regularized Self-Attention 23 Mar 2022 · 0 repositories · arXiv:2203.12276
-
Input-specific Attention Subnetworks for Adversarial Detection 23 Mar 2022 · 0 repositories · arXiv:2203.12298Syntology 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
A Computational Approach to Understand Mental Health from Reddit: Knowledge-aware Multitask Learning Framework 22 Mar 2022 · 0 repositories · arXiv:2203.11856
-
Are You Misinformed? A Study of Covid-Related Fake News in Bengali on Facebook 22 Mar 2022 · 0 repositories · arXiv:2203.11669
-
BERT-ASC: Auxiliary-Sentence Construction for Implicit Aspect Learning in Sentiment Analysis 22 Mar 2022 · 1 repository · arXiv:2203.11702Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples) · 1 pointer-only (licence)
-
Factual Consistency of Multilingual Pretrained Language Models 22 Mar 2022 · 1 repository · arXiv:2203.11552
-
Self-supervision through Random Segments with Autoregressive Coding (RandSAC) 22 Mar 2022 · 0 repositories · arXiv:2203.12054
-
Transformer based ensemble for emotion detection 22 Mar 2022 · 0 repositories · arXiv:2203.11899
-
Under the Hood of Transformer Networks for Trajectory Forecasting 22 Mar 2022 · 0 repositories · arXiv:2203.11878
-
A Slot Is Not Built in One Utterance: Spoken Language Dialogs with Sub-Slots 21 Mar 2022 · 1 repository · arXiv:2203.10759
-
An Intellectual Property Entity Recognition Method Based on Transformer and Technological Word Information 21 Mar 2022 · 0 repositories · arXiv:2203.10717
-
AraBART: a Pretrained Arabic Sequence-to-Sequence Model for Abstractive Summarization 21 Mar 2022 · 0 repositories · arXiv:2203.10945
-
Compression of Generative Pre-trained Language Models via Quantization 21 Mar 2022 · 0 repositories · arXiv:2203.10705
-
DQ-BART: Efficient Sequence-to-Sequence Model via Joint Distillation and Quantization 21 Mar 2022 · 2 repositories · arXiv:2203.11239Syntology official: harvested, nothing ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
Neural Token Segmentation for High Token-Internal Complexity 21 Mar 2022 · 0 repositories · arXiv:2203.10845
-
Semantic Similarity Computing for Scientific Academic Conferences fused with domain features 21 Mar 2022 · 0 repositories · arXiv:2203.12593
-
Towards Explainable Evaluation Metrics for Natural Language Generation 21 Mar 2022 · 1 repository · arXiv:2203.11131
-
Build a Robust QA System with Transformer-based Mixture of Experts 20 Mar 2022 · 1 repository · arXiv:2204.09598
-
Cluster & Tune: Boost Cold Start Performance in Text Classification 20 Mar 2022 · 1 repository · arXiv:2203.10581Syntology official: harvested, nothing ran · 0 ran · 3 unverified (of 3 harvested samples)
-
g2pW: A Conditional Weighted Softmax BERT for Polyphone Disambiguation in Mandarin 20 Mar 2022 · 1 repository · arXiv:2203.10430
-
How does the pre-training objective affect what large language models learn about linguistic properties? 20 Mar 2022 · 1 repository · arXiv:2203.10415