Methods › General › Regularization › Attention Dropout › Papers, page 63
Attention Dropout
Papers archive 2025-07-28
archive papers tagged: 10,892 · with a code link: 4,634 · where Syntology ran a sample: 1,270 (1,043 with a run with no instrument failure, 227 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (1,270 of 10,892 tagged: 1,043 with a run with no instrument failure, 227 where every run was a failure of Syntology's instrument)
Page 63 of 109: papers 6,201 to 6,300 of 10,892, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Collateral facilitation in humans and language models 9 Nov 2022 · 1 repository · arXiv:2211.05198Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Cross-lingual Transfer Learning for Check-worthy Claim Identification over Twitter 9 Nov 2022 · 0 repositories · arXiv:2211.05087
-
Large Language Models with Controllable Working Memory 9 Nov 2022 · 0 repositories · arXiv:2211.05110
-
Mask More and Mask Later: Efficient Pre-training of Masked Language Models by Disentangling the [MASK] Token 9 Nov 2022 · 1 repository · arXiv:2211.04898
-
Sentiment Analysis of Persian Language: Review of Algorithms, Approaches and Datasets 9 Nov 2022 · 0 repositories · arXiv:2212.06041
-
A Multimodal Approach for Dementia Detection from Spontaneous Speech with Tensor Fusion Layer 8 Nov 2022 · 0 repositories · arXiv:2211.04368
-
Active Example Selection for In-Context Learning 8 Nov 2022 · 1 repository · arXiv:2211.04486Syntology official (archive's flag): 10 ran · 10 ran (of which 3 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 6 where Syntology's instrument failed) · 6 unverified (of 16 harvested samples)
-
Conciseness: An Overlooked Language Task 8 Nov 2022 · 0 repositories · arXiv:2211.04126
-
Discover, Explanation, Improvement: An Automatic Slice Detection Framework for Natural Language Processing 8 Nov 2022 · 0 repositories · arXiv:2211.04476
-
AD-BERT: Using Pre-trained contextualized embeddings to Predict the Progression from Mild Cognitive Impairment to Alzheimer's Disease 7 Nov 2022 · 0 repositories · arXiv:2212.06042
-
Suffix Retrieval-Augmented Language Modeling 6 Nov 2022 · 1 repository · arXiv:2211.03053
-
BERT-Deep CNN: State-of-the-Art for Sentiment Analysis of COVID-19 Tweets 4 Nov 2022 · 0 repositories · arXiv:2211.09733
-
BERT for Long Documents: A Case Study of Automated ICD Coding 4 Nov 2022 · 0 repositories · arXiv:2211.02519
-
Continuous Prompt Tuning Based Textual Entailment Model for E-commerce Entity Typing 4 Nov 2022 · 1 repository · arXiv:2211.02483
-
Crosslingual Generalization through Multitask Finetuning 3 Nov 2022 · 1 repository · arXiv:2211.01786Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Fine-Tuning Language Models via Epistemic Neural Networks 3 Nov 2022 · 1 repository · arXiv:2211.01568Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Using Large Pre-Trained Language Model to Assist FDA in Premarket Medical Device 3 Nov 2022 · 0 repositories · arXiv:2212.01217
-
BECTRA: Transducer-based End-to-End ASR with BERT-Enhanced Encoder 2 Nov 2022 · 0 repositories · arXiv:2211.00792
-
eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers 2 Nov 2022 · 2 repositories · arXiv:2211.01324Syntology 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 0 violated, 9 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples)
-
Multi-level Distillation of Semantic Knowledge for Pre-training Multilingual Language Model 2 Nov 2022 · 0 repositories · arXiv:2211.01200
-
Processing Long Legal Documents with Pre-trained Transformers: Modding LegalBERT and Longformer 2 Nov 2022 · 0 repositories · arXiv:2211.00974
-
Data Level Lottery Ticket Hypothesis for Vision Transformers 2 Nov 2022 · 1 repository · arXiv:2211.01484
-
ClassActionPrediction: A Challenging Benchmark for Legal Judgment Prediction of Class Action Cases in the US 1 Nov 2022 · 1 repository · arXiv:2211.00582
-
FRSUM: Towards Faithful Abstractive Summarization via Enhancing Factual Robustness 1 Nov 2022 · 0 repositories · arXiv:2211.00294
-
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small 1 Nov 2022 · 7 repositories · arXiv:2211.00593Syntology community repositories only · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 13 harvested samples)
-
Investigating Content-Aware Neural Text-To-Speech MOS Prediction Using Prosodic and Linguistic Features 1 Nov 2022 · 0 repositories · arXiv:2211.00342
-
Two-stage LLM Fine-tuning with Less Specialization and More Generalization 1 Nov 2022 · 0 repositories · arXiv:2211.00635
-
Reduce, Reuse, Recycle: Improving Training Efficiency with Distillation 1 Nov 2022 · 0 repositories · arXiv:2211.00683
-
T5lephone: Bridging Speech and Text Self-supervised Models for Spoken Language Understanding via Phoneme level T5 1 Nov 2022 · 1 repository · arXiv:2211.00586
-
Text-Only Training for Image Captioning using Noise-Injected CLIP 1 Nov 2022 · 4 repositories · arXiv:2211.00575Syntology official (archive's flag): 2 ran · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
Efficient Document Retrieval by End-to-End Refining and Quantizing BERT Embedding with Contrastive Product Quantization 31 Oct 2022 · 1 repository · arXiv:2210.17170
-
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers 31 Oct 2022 · 17 repositories · arXiv:2210.17323Syntology official (archive's flag): 1 ran · 5 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 3 where Syntology's instrument failed) · 10 unverified (of 15 harvested samples) · 1 pointer-only (licence)
-
Leveraging Pre-trained Models for Failure Analysis Triplets Generation 31 Oct 2022 · 0 repositories · arXiv:2210.17497
-
QuaLA-MiniLM: a Quantized Length Adaptive MiniLM 31 Oct 2022 · 2 repositories · arXiv:2210.17114
-
SDCL: Self-Distillation Contrastive Learning for Chinese Spell Checking 31 Oct 2022 · 0 repositories · arXiv:2210.17168
-
SSD-LM: Semi-autoregressive Simplex-based Diffusion Language Model for Text Generation and Modular Control 31 Oct 2022 · 2 repositories · arXiv:2210.17432
-
Towards Zero-Shot and Few-Shot Table Question Answering using GPT-3 31 Oct 2022 · 0 repositories · arXiv:2210.17284
-
Learning to Decompose: Hypothetical Question Decomposition Based on Comparable Texts 30 Oct 2022 · 0 repositories · arXiv:2210.16865
-
Parameter-Efficient Tuning Makes a Good Classification Head 30 Oct 2022 · 1 repository · arXiv:2210.16771
-
BERT Meets CTC: New Formulation of End-to-End Speech Recognition with Pre-trained Masked Language Model 29 Oct 2022 · 0 repositories · arXiv:2210.16663
-
Empirical Evaluation of Post-Training Quantization Methods for Language Tasks 29 Oct 2022 · 0 repositories · arXiv:2210.16621
-
Exploiting prompt learning with pre-trained language models for Alzheimer's Disease detection 29 Oct 2022 · 1 repository · arXiv:2210.16539
-
BEBERT: Efficient and Robust Binary Ensemble BERT 28 Oct 2022 · 1 repository · arXiv:2210.15976
-
Feature Engineering vs BERT on Twitter Data 28 Oct 2022 · 0 repositories · arXiv:2210.16168
-
On the Use of Modality-Specific Large-Scale Pre-Trained Encoders for Multimodal Sentiment Analysis 28 Oct 2022 · 0 repositories · arXiv:2210.15937
-
Probing for targeted syntactic knowledge through grammatical error detection 28 Oct 2022 · 1 repository · arXiv:2210.16228
-
BERT-Flow-VAE: A Weakly-supervised Model for Multi-Label Text Classification 27 Oct 2022 · 0 repositories · arXiv:2210.15225
-
COCO-DR: Combating Distribution Shifts in Zero-Shot Dense Retrieval with Contrastive and Distributionally Robust Learning 27 Oct 2022 · 1 repository · arXiv:2210.15212Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; the one sample that ran constructed an object rather than computing a result (of 2 harvested samples)
-
COST-EFF: Collaborative Optimization of Spatial and Temporal Efficiency with Slenderized Multi-exit Language Models 27 Oct 2022 · 1 repository · arXiv:2210.15523
-
Fast DistilBERT on CPUs 27 Oct 2022 · 1 repository · arXiv:2211.07715
-
FCTalker: Fine and Coarse Grained Context Modeling for Expressive Conversational Speech Synthesis 27 Oct 2022 · 1 repository · arXiv:2210.15360
-
Masked Vision-Language Transformer in Fashion 27 Oct 2022 · 1 repository · arXiv:2210.15110
-
TRScore: A Novel GPT-based Readability Scorer for ASR Segmentation and Punctuation model evaluation and selection 27 Oct 2022 · 0 repositories · arXiv:2210.15104
-
Unsupervised Boundary-Aware Language Model Pretraining for Chinese Sequence Labeling 27 Oct 2022 · 2 repositories · arXiv:2210.15231
-
Automatic extraction of materials and properties from superconductors scientific literature 26 Oct 2022 · 2 repositories · arXiv:2210.15600
-
Beyond English-Centric Bitexts for Better Multilingual Language Representation Learning 26 Oct 2022 · 0 repositories · arXiv:2210.14867
-
Bi-Link: Bridging Inductive Link Predictions from Text via Contrastive Learning of Transformers and Prompts 26 Oct 2022 · 0 repositories · arXiv:2210.14463
-
Don't Prompt, Search! Mining-based Zero-Shot Learning with Language Models 26 Oct 2022 · 0 repositories · arXiv:2210.14803
-
Exploring Robustness of Prefix Tuning in Noisy Data: A Case Study in Financial Sentiment Analysis 26 Oct 2022 · 0 repositories · arXiv:2211.05584
-
Leveraging Affirmative Interpretations from Negation Improves Natural Language Understanding 26 Oct 2022 · 1 repository · arXiv:2210.14486
-
How Long Is Enough? Exploring the Optimal Intervals of Long-Range Clinical Note Language Modeling 25 Oct 2022 · 1 repository · arXiv:2211.07713
-
IELM: An Open Information Extraction Benchmark for Pre-Trained Language Models 25 Oct 2022 · 0 repositories · arXiv:2210.14128
-
XRICL: Cross-lingual Retrieval-Augmented In-Context Learning for Cross-lingual Text-to-SQL Semantic Parsing 25 Oct 2022 · 0 repositories · arXiv:2210.13693
-
Inferring Past Human Actions in Homes with Abductive Reasoning 24 Oct 2022 · 1 repository · arXiv:2210.13984
-
Effective Pre-Training Objectives for Transformer-based Autoencoders 24 Oct 2022 · 0 repositories · arXiv:2210.13536
-
Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task 24 Oct 2022 · 4 repositories · arXiv:2210.13382Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
Entity-level Sentiment Analysis in Contact Center Telephone Conversations 24 Oct 2022 · 0 repositories · arXiv:2210.13401
-
Explaining Translationese: why are Neural Classifiers Better and what do they Learn? 24 Oct 2022 · 0 repositories · arXiv:2210.13391
-
Exploring Euphemism Detection in Few-Shot and Zero-Shot Settings 24 Oct 2022 · 1 repository · arXiv:2210.12926
-
Perfectly Secure Steganography Using Minimum Entropy Coupling 24 Oct 2022 · 2 repositories · arXiv:2210.14889Syntology community repositories only · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 4 harvested samples)
-
The Better Your Syntax, the Better Your Semantics? Probing Pretrained Language Models for the English Comparative Correlative 24 Oct 2022 · 0 repositories · arXiv:2210.13181
-
A BERT-based Deep Learning Approach for Reputation Analysis in Social Media 23 Oct 2022 · 0 repositories · arXiv:2211.01954
-
Data Augmentation for Automated Essay Scoring using Transformer Models 23 Oct 2022 · 0 repositories · arXiv:2210.12809
-
Discriminative Language Model as Semantic Consistency Scorer for Prompt-based Few-Shot Text Classification 23 Oct 2022 · 0 repositories · arXiv:2210.12763
-
Leveraging Large Language Models for Multiple Choice Question Answering 22 Oct 2022 · 1 repository · arXiv:2210.12353Syntology official (archive's flag): 6 ran · 6 ran (of which 3 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
Meta-learning Pathologies from Radiology Reports using Variance Aware Prototypical Networks 22 Oct 2022 · 0 repositories · arXiv:2210.13979
-
A Causal Framework to Quantify the Robustness of Mathematical Reasoning with Language Models 21 Oct 2022 · 1 repository · arXiv:2210.12023Syntology official (archive's flag): 5 ran · 5 ran (of which 2 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 0 violated, 2 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Amos: An Adam-style Optimizer with Adaptive Weight Decay towards Model-Oriented Scale 21 Oct 2022 · 1 repository · arXiv:2210.11693Syntology official (archive's flag): 14 ran · 14 ran (of which 0 constructed an object rather than computing a result; 14 with no instrument failure: 0 honoured, 0 violated, 14 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 18 harvested samples)
-
Decoding a Neural Retriever's Latent Space for Query Suggestion 21 Oct 2022 · 1 repository · arXiv:2210.12084
-
Diffuser: Efficient Transformers with Multi-hop Attention Diffusion for Long Sequences 21 Oct 2022 · 1 repository · arXiv:2210.11794
-
Discovering Differences in the Representation of People using Contextualized Semantic Axes 21 Oct 2022 · 1 repository · arXiv:2210.12170
-
LittleBird: Efficient Faster & Longer Transformer for Question Answering 21 Oct 2022 · 0 repositories · arXiv:2210.11870
-
Probing with Noise: Unpicking the Warp and Weft of Embeddings 21 Oct 2022 · 1 repository · arXiv:2210.12206
-
SLING: Sino Linguistic Evaluation of Large Language Models 21 Oct 2022 · 1 repository · arXiv:2210.11689Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples)
-
SpaBERT: A Pretrained Language Model from Geographic Data for Geo-Entity Representation 21 Oct 2022 · 0 repositories · arXiv:2210.12213
-
WikiWhy: Answering and Explaining Cause-and-Effect Questions 21 Oct 2022 · 0 repositories · arXiv:2210.12152
-
3DALL-E: Integrating Text-to-Image AI in 3D Design Workflows 20 Oct 2022 · 0 repositories · arXiv:2210.11603
-
Composing Ensembles of Pre-trained Models via Iterative Consensus 20 Oct 2022 · 0 repositories · arXiv:2210.11522
-
General Image Descriptors for Open World Image Retrieval using ViT CLIP 20 Oct 2022 · 1 repository · arXiv:2210.11141Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Scaling Instruction-Finetuned Language Models 20 Oct 2022 · 9 repositories · arXiv:2210.11416Syntology 8 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 7 where Syntology's instrument failed) · 9 unverified (of 17 harvested samples) · 2 pointer-only (licence)
-
A Unified Neural Network Model for Readability Assessment with Feature Projection and Length-Balanced Loss 19 Oct 2022 · 1 repository · arXiv:2210.10305
-
BioGPT: Generative Pre-trained Transformer for Biomedical Text Generation and Mining 19 Oct 2022 · 4 repositories · arXiv:2210.10341
-
Language Model Decomposition: Quantifying the Dependency and Correlation of Language Models 19 Oct 2022 · 1 repository · arXiv:2210.10289
-
Self-supervised Graph Masking Pre-training for Graph-to-Text Generation 19 Oct 2022 · 1 repository · arXiv:2210.10599
-
Tempo: Accelerating Transformer-Based Model Training through Memory Footprint Reduction 19 Oct 2022 · 1 repository · arXiv:2210.10246Syntology official: harvested, nothing ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Towards a neural architecture of language: Deep learning versus logistics of access in neural architectures for compositional processing 19 Oct 2022 · 0 repositories · arXiv:2210.10543
-
ELASTIC: Numerical Reasoning with Adaptive Symbolic Compiler 18 Oct 2022 · 1 repository · arXiv:2210.10105Syntology official (archive's flag): 2 ran · 2 ran (of which 1 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Swinv2-Imagen: Hierarchical Vision Transformer Diffusion Models for Text-to-Image Generation 18 Oct 2022 · 0 repositories · arXiv:2210.09549
-
Systematicity in GPT-3's Interpretation of Novel English Noun Compounds 18 Oct 2022 · 0 repositories · arXiv:2210.09492
-
Team Flow at DRC2022: Pipeline System for Travel Destination Recommendation Task in Spoken Dialogue 18 Oct 2022 · 0 repositories · arXiv:2210.09518