Methods › General › Regularization › Attention Dropout › Papers, page 58
Attention Dropout
Papers archive 2025-07-28
archive papers tagged: 10,892 · with a code link: 4,634 · where Syntology ran a sample: 1,270 (1,043 with a run with no instrument failure, 227 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (1,270 of 10,892 tagged: 1,043 with a run with no instrument failure, 227 where every run was a failure of Syntology's instrument)
Page 58 of 109: papers 5,701 to 5,800 of 10,892, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Multimodal Pre-training Framework for Sequential Recommendation via Contrastive Learning 21 Mar 2023 · 0 repositories · arXiv:2303.11879
-
Sparse-IFT: Sparse Iso-FLOP Transformations for Maximizing Training Efficiency 21 Mar 2023 · 2 repositories · arXiv:2303.11525Syntology official (archive's flag): 21 ran · 21 ran (of which 0 constructed an object rather than computing a result; 20 with no instrument failure: 0 honoured, 0 violated, 20 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 24 harvested samples) · 2 pointer-only (licence)
-
Capabilities of GPT-4 on Medical Challenge Problems 20 Mar 2023 · 1 repository · arXiv:2303.13375
-
Character, Word, or Both? Revisiting the Segmentation Granularity for Chinese Pre-trained Language Models 20 Mar 2023 · 1 repository · arXiv:2303.10893
-
Mind meets machine: Unravelling GPT-4's cognitive psychology 20 Mar 2023 · 0 repositories · arXiv:2303.11436
-
Bangla Grammatical Error Detection Using T5 Transformer Model 19 Mar 2023 · 2 repositories · arXiv:2303.10612
-
CTRAN: CNN-Transformer-based Network for Natural Language Understanding 19 Mar 2023 · 1 repository · arXiv:2303.10606
-
PACO: Provocation Involving Action, Culture, and Oppression 19 Mar 2023 · 0 repositories · arXiv:2303.12808
-
A Comprehensive Capability Analysis of GPT-3 and GPT-3.5 Series Models 18 Mar 2023 · 0 repositories · arXiv:2303.10420
-
An Empirical Study of Pre-trained Language Models in Simple Knowledge Graph Question Answering 18 Mar 2023 · 1 repository · arXiv:2303.10368
-
NoisyHate: Mining Online Human-Written Perturbations for Realistic Robustness Benchmarking of Content Moderation Models 18 Mar 2023 · 0 repositories · arXiv:2303.10430
-
SPDF: Sparse Pre-training and Dense Fine-tuning for Large Language Models 18 Mar 2023 · 0 repositories · arXiv:2303.10464
-
GADformer: A Transparent Transformer Model for Group Anomaly Detection on Trajectories 17 Mar 2023 · 1 repository · arXiv:2303.09841
-
GPTs are GPTs: An Early Look at the Labor Market Impact Potential of Large Language Models 17 Mar 2023 · 0 repositories · arXiv:2303.10130
-
Trained on 100 million words and still in shape: BERT meets British National Corpus 17 Mar 2023 · 2 repositories · arXiv:2303.09859Syntology official (archive's flag): 3 ran · 3 ran (of which 3 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified; every one of the 3 samples that ran constructed an object rather than computing a result (of 7 harvested samples) · 7 pointer-only (licence)
-
Block-wise Bit-Compression of Transformer-based Models 16 Mar 2023 · 0 repositories · arXiv:2303.09184
-
Can Generative Pre-trained Transformers (GPT) Pass Assessments in Higher Education Programming Courses? 16 Mar 2023 · 0 repositories · arXiv:2303.09325
-
Exploring Distributional Shifts in Large Language Models for Code Analysis 16 Mar 2023 · 0 repositories · arXiv:2303.09128
-
Instance-Conditioned GAN Data Augmentation for Representation Learning 16 Mar 2023 · 0 repositories · arXiv:2303.09677
-
Jump to Conclusions: Short-Cutting Transformers With Linear Transformations 16 Mar 2023 · 2 repositories · arXiv:2303.09435Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Measuring Improvement of F₁-Scores in Detection of Self-Admitted Technical Debt 16 Mar 2023 · 0 repositories · arXiv:2303.09617
-
SmartBERT: A Promotion of Dynamic Early Exiting Mechanism for Accelerating BERT Inference 16 Mar 2023 · 0 repositories · arXiv:2303.09266Syntology 3 ran (of which 3 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; every one of the 3 samples that ran constructed an object rather than computing a result (of 4 harvested samples)
-
Towards the Scalable Evaluation of Cooperativeness in Language Models 16 Mar 2023 · 0 repositories · arXiv:2303.13360
-
TypeT5: Seq2seq Type Inference using Static Analysis 16 Mar 2023 · 1 repository · arXiv:2303.09564Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples)
-
Automated Interactive Domain-Specific Conversational Agents that Understand Human Dialogs 15 Mar 2023 · 0 repositories · arXiv:2303.08941
-
Efficient Uncertainty Estimation with Gaussian Process for Reliable Dialog Response Retrieval 15 Mar 2023 · 0 repositories · arXiv:2303.08599
-
GCRE-GPT: A Generative Model for Comparative Relation Extraction 15 Mar 2023 · 0 repositories · arXiv:2303.08601
-
PRESTO: A Multilingual Dataset for Parsing Realistic Task-Oriented Dialogs 15 Mar 2023 · 1 repository · arXiv:2303.08954
-
SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models 15 Mar 2023 · 1 repository · arXiv:2303.08896Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 3 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 5 unverified (of 11 harvested samples)
-
Do Transformers Parse while Predicting the Masked Word? 14 Mar 2023 · 0 repositories · arXiv:2303.08117
-
Can ChatGPT Replace Traditional KBQA Models? An In-depth Analysis of the Question Answering Performance of the GPT LLM Family 14 Mar 2023 · 2 repositories · arXiv:2303.07992
-
Features matching using natural language processing 14 Mar 2023 · 0 repositories · arXiv:2303.12804
-
Finding the Needle in a Haystack: Unsupervised Rationale Extraction from Long Text Classifiers 14 Mar 2023 · 0 repositories · arXiv:2303.07991
-
MEDBERT.de: A Comprehensive German BERT Model for the Medical Domain 14 Mar 2023 · 0 repositories · arXiv:2303.08179
-
Neuro-symbolic Commonsense Social Reasoning 14 Mar 2023 · 3 repositories · arXiv:2303.08264
-
RE-MOVE: An Adaptive Policy Design for Robotic Navigation Tasks in Dynamic Environments via Language-Based Feedback 14 Mar 2023 · 0 repositories · arXiv:2303.07622
-
Deep Learning Approach for Classifying the Aggressive Comments on Social Media: Machine Translated Data Vs Real Life Data 13 Mar 2023 · 0 repositories · arXiv:2303.07484
-
Large Language Models in the Workplace: A Case Study on Prompt Engineering for Job Type Classification 13 Mar 2023 · 0 repositories · arXiv:2303.07142
-
Transformer-based approaches to Sentiment Detection 13 Mar 2023 · 0 repositories · arXiv:2303.07292
-
Large Language Models Know Your Contextual Search Intent: A Prompting Framework for Conversational Search 12 Mar 2023 · 2 repositories · arXiv:2303.06573
-
LUKE-Graph: A Transformer-based Approach with Gated Relational Graph Attention for Cloze-style Reading Comprehension 12 Mar 2023 · 0 repositories · arXiv:2303.06675
-
Proactive Prioritization of App Issues via Contrastive Learning 12 Mar 2023 · 1 repository · arXiv:2303.06586
-
Learning Combinatorial Prompts for Universal Controllable Image Captioning 11 Mar 2023 · 0 repositories · arXiv:2303.06338
-
Algorithmic Ghost in the Research Shell: Large Language Models and Academic Knowledge Creation in Management Research 10 Mar 2023 · 0 repositories · arXiv:2303.07304
-
Is In-hospital Meta-information Useful for Abstractive Discharge Summary Generation? 10 Mar 2023 · 0 repositories · arXiv:2303.06002
-
Research on CPI Prediction Based on Natural Language Processing 10 Mar 2023 · 0 repositories · arXiv:2303.05666
-
ChatGPT may Pass the Bar Exam soon, but has a Long Way to Go for the LexGLUE benchmark 9 Mar 2023 · 1 repository · arXiv:2304.12202
-
ICL-D3IE: In-Context Learning with Diverse Demonstrations Updating for Document Information Extraction 9 Mar 2023 · 1 repository · arXiv:2303.05063
-
Large Language Models (GPT) Struggle to Answer Multiple-Choice Questions about Code 9 Mar 2023 · 0 repositories · arXiv:2303.08033
-
ChatGPT Participates in a Computer Science Exam 8 Mar 2023 · 1 repository · arXiv:2303.09461
-
Cost-Effective Hyperparameter Optimization for Large Language Model Generation Inference 8 Mar 2023 · 3 repositories · arXiv:2303.04673Syntology community repositories only · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Stealing the Decoding Algorithms of Language Models 8 Mar 2023 · 1 repository · arXiv:2303.04729
-
X-Pruner: eXplainable Pruning for Vision Transformers 8 Mar 2023 · 1 repository · arXiv:2303.04935
-
A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPT 7 Mar 2023 · 1 repository · arXiv:2303.04226
-
ADELT: Transpilation Between Deep Learning Frameworks 7 Mar 2023 · 0 repositories · arXiv:2303.03593
-
Classifying Text-Based Conspiracy Tweets related to COVID-19 using Contextualized Word Embeddings 7 Mar 2023 · 0 repositories · arXiv:2303.03706
-
German BERT Model for Legal Named Entity Recognition 7 Mar 2023 · 0 repositories · arXiv:2303.05388
-
Gradient-Free Structured Pruning with Unlabeled Data 7 Mar 2023 · 0 repositories · arXiv:2303.04185
-
Spelling convention sensitivity in neural language models 6 Mar 2023 · 0 repositories · arXiv:2303.03457
-
Towards Zero-Shot Functional Compositionality of Language Models 6 Mar 2023 · 1 repository · arXiv:2303.03103
-
Video Question Answering Using CLIP-Guided Visual-Text Attention 6 Mar 2023 · 0 repositories · arXiv:2303.03131
-
Industry Risk Assessment via Hierarchical Financial Data Using Stock Market Sentiment Indicators 5 Mar 2023 · 0 repositories · arXiv:2303.02707
-
Robust affine point matching via quadratic assignment on Grassmannians 5 Mar 2023 · 3 repositories · arXiv:2303.02698
-
Training-Free Acceleration of ViTs with Delayed Spatial Merging 4 Mar 2023 · 1 repository · arXiv:2303.02331
-
Early Warning Signals of Social Instabilities in Twitter Data 3 Mar 2023 · 0 repositories · arXiv:2303.05401
-
Exploring Data Augmentation Methods on Social Media Corpora 3 Mar 2023 · 0 repositories · arXiv:2303.02198
-
Multi label classification of Artificial Intelligence related patents using Modified D2SBERT and Sentence Attention mechanism 3 Mar 2023 · 0 repositories · arXiv:2303.03165
-
Pre-trained Model Representations and their Robustness against Noise for Speech Emotion Analysis 3 Mar 2023 · 0 repositories · arXiv:2303.03177
-
Prompt, Generate, then Cache: Cascade of Foundation Models makes Strong Few-shot Learners 3 Mar 2023 · 3 repositories · arXiv:2303.02151Syntology official: not harvested · 0 ran · 1 unverified (of 1 harvested sample)
-
Prophet: Prompting Large Language Models with Complementary Answer Heuristics for Knowledge-based Visual Question Answering 3 Mar 2023 · 1 repository · arXiv:2303.01903Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples)
-
TrojText: Test-time Invisible Textual Trojan Insertion 3 Mar 2023 · 1 repository · arXiv:2303.02242Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Will Affective Computing Emerge from Foundation Models and General AI? A First Evaluation on ChatGPT 3 Mar 2023 · 0 repositories · arXiv:2303.03186
-
Adopting the Multi-answer Questioning Task with an Auxiliary Metric for Extreme Multi-label Text Classification Utilizing the Label Hierarchy 2 Mar 2023 · 0 repositories · arXiv:2303.01064
-
Can BERT Refrain from Forgetting on Sequential Tasks? A Probing Study 2 Mar 2023 · 1 repository · arXiv:2303.01081
-
Evaluating Parameter-Efficient Transfer Learning Approaches on SURE Benchmark for Speech Understanding 2 Mar 2023 · 1 repository · arXiv:2303.03267
-
INO at Factify 2: Structure Coherence based Multi-Modal Fact Verification 2 Mar 2023 · 1 repository · arXiv:2303.01510
-
Sparse MoE as the New Dropout: Scaling Dense and Self-Slimmable Transformers 2 Mar 2023 · 1 repository · arXiv:2303.01610Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
WiCE: Real-World Entailment for Claims in Wikipedia 2 Mar 2023 · 2 repositories · arXiv:2303.01432Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
A Framework for Neurosymbolic Robot Action Planning using Large Language Models 1 Mar 2023 · 1 repository · arXiv:2303.00438Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Competence-Based Analysis of Language Models 1 Mar 2023 · 0 repositories · arXiv:2303.00333
-
Domain-adapted large language models for classifying nuclear medicine reports 1 Mar 2023 · 0 repositories · arXiv:2303.01258
-
How Robust is GPT-3.5 to Predecessors? A Comprehensive Study on Language Understanding Tasks 1 Mar 2023 · 0 repositories · arXiv:2303.00293
-
N-best T5: Robust ASR Error Correction using Multiple Input Hypotheses and Constrained Decoding Space 1 Mar 2023 · 0 repositories · arXiv:2303.00456
-
ToxVis: Enabling Interpretability of Implicit vs. Explicit Toxicity Detection Models with Interactive Visualization 1 Mar 2023 · 0 repositories · arXiv:2303.09402
-
Are Character-level Translations Worth the Wait? Comparing ByT5 and mT5 for Machine Translation 28 Feb 2023 · 1 repository · arXiv:2302.14220
-
Automatically Classifying Emotions based on Text: A Comparative Exploration of Different Datasets 28 Feb 2023 · 0 repositories · arXiv:2302.14727
-
Zero-Shot Cross-Lingual Summarization via Large Language Models 28 Feb 2023 · 0 repositories · arXiv:2302.14229
-
Information-Restricted Neural Language Models Reveal Different Brain Regions' Sensitivity to Semantics, Syntax and Context 28 Feb 2023 · 1 repository · arXiv:2302.14389Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Large Language Models Are State-of-the-Art Evaluators of Translation Quality 28 Feb 2023 · 4 repositories · arXiv:2302.14520
-
Sampled Transformer for Point Sets 28 Feb 2023 · 0 repositories · arXiv:2302.14346
-
Text classification dataset and analysis for Uzbek language 28 Feb 2023 · 1 repository · arXiv:2302.14494
-
Weighted Sampling for Masked Language Modeling 28 Feb 2023 · 0 repositories · arXiv:2302.14225
-
Elementwise Language Representation 27 Feb 2023 · 0 repositories · arXiv:2302.13475
-
Inseq: An Interpretability Toolkit for Sequence Generation Models 27 Feb 2023 · 2 repositories · arXiv:2302.13942
-
LLaMA: Open and Efficient Foundation Language Models 27 Feb 2023 · 57 repositories · arXiv:2302.13971Syntology official: no sample here; runs from other or unrecorded repositories · 37 ran (of which 9 constructed an object rather than computing a result; 25 with no instrument failure: 3 honoured, 0 violated, 22 with no contract checked; 12 where Syntology's instrument failed) · 21 unverified (of 58 harvested samples) · 4 pointer-only (licence)
-
Reward Design with Language Models 27 Feb 2023 · 1 repository · arXiv:2303.00001
-
Systematic Rectification of Language Models via Dead-end Analysis 27 Feb 2023 · 1 repository · arXiv:2302.14003Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
Using Auxiliary Tasks In Multimodal Fusion Of Wav2vec 2.0 And BERT For Multimodal Emotion Recognition 27 Feb 2023 · 0 repositories · arXiv:2302.13661
-
Comparing Sentence-Level Suggestions to Message-Level Suggestions in AI-Mediated Communication 26 Feb 2023 · 0 repositories · arXiv:2302.13382
-
Efficient Ensemble for Multimodal Punctuation Restoration using Time-Delay Neural Network 26 Feb 2023 · 1 repository · arXiv:2302.13376