Methods › General › Regularization › Attention Dropout › Papers, page 51
Attention Dropout
Papers archive 2025-07-28
archive papers tagged: 10,892 · with a code link: 4,634 · where Syntology ran a sample: 1,270 (1,043 with a run with no instrument failure, 227 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (1,270 of 10,892 tagged: 1,043 with a run with no instrument failure, 227 where every run was a failure of Syntology's instrument)
Page 51 of 109: papers 5,001 to 5,100 of 10,892, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
MAT: Mixed-Strategy Game of Adversarial Training in Fine-tuning 27 Jun 2023 · 0 repositories · arXiv:2306.15826
-
SparseOptimizer: Sparsify Language Models through Moreau-Yosida Regularization and Accelerate via Compiler Co-design 27 Jun 2023 · 0 repositories · arXiv:2306.15656
-
Unleashing the Power of User Reviews: Exploring Airline Choices at Catania Airport, Italy 27 Jun 2023 · 0 repositories · arXiv:2306.15541
-
Constraint-aware and Ranking-distilled Token Pruning for Efficient Transformer Inference 26 Jun 2023 · 1 repository · arXiv:2306.14393
-
Exploring the Robustness of Large Language Models for Solving Programming Problems 26 Jun 2023 · 0 repositories · arXiv:2306.14583
-
LongCoder: A Long-Range Pre-trained Language Model for Code Completion 26 Jun 2023 · 1 repository · arXiv:2306.14893Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Addressing Cold Start Problem for End-to-end Automatic Speech Scoring 25 Jun 2023 · 0 repositories · arXiv:2306.14310
-
Interactive Design by Integrating a Large Pre-Trained Language Model and Building Information Modeling 25 Jun 2023 · 0 repositories · arXiv:2306.14165
-
Let's Do a Thought Experiment: Using Counterfactuals to Improve Moral Reasoning 25 Jun 2023 · 0 repositories · arXiv:2306.14308
-
Revolutionizing Cyber Threat Detection with Large Language Models: A privacy-preserving BERT-based Lightweight Model for IoT/IIoT Devices 25 Jun 2023 · 0 repositories · arXiv:2306.14263
-
Switch-BERT: Learning to Model Multimodal Interactions by Switching Attention and Input 25 Jun 2023 · 0 repositories · arXiv:2306.14182
-
Comparison of Pre-trained Language Models for Turkish Address Parsing 24 Jun 2023 · 0 repositories · arXiv:2306.13947
-
IERL: Interpretable Ensemble Representation Learning -- Combining CrowdSourced Knowledge and Distributed Semantic Representations 24 Jun 2023 · 0 repositories · arXiv:2306.13865
-
Is Pre-training Truly Better Than Meta-Learning? 24 Jun 2023 · 0 repositories · arXiv:2306.13841
-
L3Cube-MahaSent-MD: A Multi-domain Marathi Sentiment Analysis Dataset and Transformer Models 24 Jun 2023 · 1 repository · arXiv:2306.13888
-
Large Language Models as Sous Chefs: Revising Recipes with GPT-3 24 Jun 2023 · 1 repository · arXiv:2306.13986
-
Large Sequence Models for Sequential Decision-Making: A Survey 24 Jun 2023 · 0 repositories · arXiv:2306.13945
-
Math Word Problem Solving by Generating Linguistic Variants of Problem Statements 24 Jun 2023 · 1 repository · arXiv:2306.13899
-
My Boli: Code-mixed Marathi-English Corpora, Pretrained Language Models and Evaluation Benchmarks 24 Jun 2023 · 1 repository · arXiv:2306.14030
-
On the Uses of Large Language Models to Interpret Ambiguous Cyberattack Descriptions 24 Jun 2023 · 0 repositories · arXiv:2306.14062
-
Partitioning-Guided K-Means: Extreme Empty Cluster Resolution for Extreme Model Compression 24 Jun 2023 · 0 repositories · arXiv:2306.14031
-
Abstractive Text Summarization for Resumes With Cutting Edge NLP Transformers and LSTM 23 Jun 2023 · 0 repositories · arXiv:2306.13315
-
On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes 23 Jun 2023 · 0 repositories · arXiv:2306.13649
-
Incorporating Graph Information in Transformer-based AMR Parsing 23 Jun 2023 · 1 repository · arXiv:2306.13467
-
LLM-Assisted Content Analysis: Using Large Language Models to Support Deductive Coding 23 Jun 2023 · 0 repositories · arXiv:2306.14924
-
Resume Information Extraction via Post-OCR Text Processing 23 Jun 2023 · 0 repositories · arXiv:2306.13775
-
System-Level Natural Language Feedback 23 Jun 2023 · 1 repository · arXiv:2306.13588Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale 23 Jun 2023 · 1 repository · arXiv:2306.15687Syntology 10 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 3 honoured, 3 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples) · 3 pointer-only (licence)
-
Cross-lingual Cross-temporal Summarization: Dataset, Models, Evaluation 22 Jun 2023 · 1 repository · arXiv:2306.12916
-
Named entity recognition in resumes 22 Jun 2023 · 0 repositories · arXiv:2306.13062
-
Prompt to GPT-3: Step-by-Step Thinking Instructions for Humor Generation 22 Jun 2023 · 1 repository · arXiv:2306.13195
-
FlakyFix: Using Large Language Models for Predicting Flaky Test Fix Categories and Test Code Repair 21 Jun 2023 · 0 repositories · arXiv:2307.00012
-
Investigating Pre-trained Language Models on Cross-Domain Datasets, a Step Closer to General AI 21 Jun 2023 · 0 repositories · arXiv:2306.12205
-
Solving and Generating NPR Sunday Puzzles with Large Language Models 21 Jun 2023 · 1 repository · arXiv:2306.12255
-
Which Spurious Correlations Impact Reasoning in NLI Models? A Visual Interactive Diagnosis through Data-Constrained Counterfactuals 21 Jun 2023 · 0 repositories · arXiv:2306.12146
-
A Novel Counterfactual Data Augmentation Method for Aspect-Based Sentiment Analysis 20 Jun 2023 · 0 repositories · arXiv:2306.11260
-
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models 20 Jun 2023 · 0 repositories · arXiv:2306.11698
-
Deep Fusion: Efficient Network Training via Pre-trained Initializations 20 Jun 2023 · 0 repositories · arXiv:2306.11903
-
Event Stream GPT: A Data Pre-processing and Modeling Library for Generative, Pre-trained Transformers over Continuous-time Sequences of Complex Events 20 Jun 2023 · 1 repository · arXiv:2306.11547Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 10 harvested samples)
-
InRank: Incremental Low-Rank Learning 20 Jun 2023 · 1 repository · arXiv:2306.11250
-
Learning to Generate Better Than Your LLM 20 Jun 2023 · 1 repository · arXiv:2306.11816
-
Textbooks Are All You Need 20 Jun 2023 · 0 repositories · arXiv:2306.11644
-
A Preliminary Study of ChatGPT on News Recommendation: Personalization, Provider Fairness, Fake News 19 Jun 2023 · 1 repository · arXiv:2306.10702
-
BayLing: Bridging Cross-lingual Alignment and Instruction Following through Interactive Translation for Large Language Models 19 Jun 2023 · 1 repository · arXiv:2306.10968
-
Fine-Tuning Language Models for Scientific Writing Support 19 Jun 2023 · 1 repository · arXiv:2306.10974
-
Generative Sequential Recommendation with GPTRec 19 Jun 2023 · 0 repositories · arXiv:2306.11114
-
SynerGPT: In-Context Learning for Personalized Drug Synergy Prediction and Drug Design 19 Jun 2023 · 0 repositories · arXiv:2307.11694
-
Instant Soup: Cheap Pruning Ensembles in A Single Pass Can Draw Lottery Tickets from Large Models 18 Jun 2023 · 1 repository · arXiv:2306.10460Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Summarization from Leaderboards to Practice: Choosing A Representation Backbone and Ensuring Robustness 18 Jun 2023 · 0 repositories · arXiv:2306.10555
-
Enhancing social network hate detection using back translation and GPT-3 augmentations during training and test-time 17 Jun 2023 · 1 repository
-
Is Self-Repair a Silver Bullet for Code Generation? 16 Jun 2023 · 1 repository · arXiv:2306.09896
-
GPT4 is Slightly Helpful for Peer-Review Assistance: A Pilot Study 16 Jun 2023 · 2 repositories · arXiv:2307.05492
-
Investigating Masking-based Data Generation in Language Models 16 Jun 2023 · 0 repositories · arXiv:2307.00008
-
Revealing the impact of social circumstances on the selection of cancer therapy through natural language processing of social work notes 16 Jun 2023 · 0 repositories · arXiv:2306.09877
-
BED: Bi-Encoder-Based Detectors for Out-of-Distribution Detection 15 Jun 2023 · 1 repository · arXiv:2306.08852
-
ChessGPT: Bridging Policy Learning and Language Modeling 15 Jun 2023 · 1 repository · arXiv:2306.09200Syntology official (archive's flag): 14 ran · 14 ran (of which 8 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 1 violated, 9 with no contract checked; 4 where Syntology's instrument failed) · 6 unverified (of 20 harvested samples)
-
Distillation Strategies for Discriminative Speech Recognition Rescoring 15 Jun 2023 · 0 repositories · arXiv:2306.09452
-
Explore, Establish, Exploit: Red Teaming Language Models from Scratch 15 Jun 2023 · 3 repositories · arXiv:2306.09442Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Exploring the MIT Mathematics and EECS Curriculum Using Large Language Models 15 Jun 2023 · 0 repositories · arXiv:2306.08997
-
Mapping Researcher Activity based on Publication Data by means of Transformers 15 Jun 2023 · 0 repositories · arXiv:2306.09049
-
SLAMB: Accelerated Large Batch Training with Sparse Communication 15 Jun 2023 · 1 repository
-
Stochastic Re-weighted Gradient Descent via Distributionally Robust Optimization 15 Jun 2023 · 0 repositories · arXiv:2306.09222
-
The pop song generator: designing an online course to teach collaborative, creative AI 15 Jun 2023 · 0 repositories · arXiv:2306.10069
-
Thrilled by Your Progress! Large Language Models (GPT-4) No Longer Struggle to Pass Assessments in Higher Education Programming Courses 15 Jun 2023 · 0 repositories · arXiv:2306.10073
-
A semantically enhanced dual encoder for aspect sentiment triplet extraction 14 Jun 2023 · 1 repository · arXiv:2306.08373
-
Assessing the Effectiveness of GPT-3 in Detecting False Political Statements: A Case Study on the LIAR Dataset 14 Jun 2023 · 1 repository · arXiv:2306.08190
-
Building a Corpus for Biomedical Relation Extraction of Species Mentions 14 Jun 2023 · 0 repositories · arXiv:2306.08403
-
Language models are not naysayers: An analysis of language models on negation benchmarks 14 Jun 2023 · 1 repository · arXiv:2306.08189
-
Towards AGI in Computer Vision: Lessons Learned from GPT and Large Language Models 14 Jun 2023 · 0 repositories · arXiv:2306.08641
-
World-to-Words: Grounded Open Vocabulary Acquisition through Fast Mapping in Vision-Language Models 14 Jun 2023 · 1 repository · arXiv:2306.08685Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples)
-
Enhancing Social Network Hate Detection Using Back Translation and GPT-3 Augmentations During Training and Test-Time 13 Jun 2023 · 1 repository
-
FLamE: Few-shot Learning from Natural Language Explanations 13 Jun 2023 · 0 repositories · arXiv:2306.08042
-
GEmo-CLAP: Gender-Attribute-Enhanced Contrastive Language-Audio Pretraining for Accurate Speech Emotion Recognition 13 Jun 2023 · 0 repositories · arXiv:2306.07848
-
Human-Like Intuitive Behavior and Reasoning Biases Emerged in Language Models -- and Disappeared in GPT-4 13 Jun 2023 · 0 repositories · arXiv:2306.07622
-
Improving Zero-Shot Detection of Low Prevalence Chest Pathologies using Domain Pre-trained Language Models 13 Jun 2023 · 1 repository · arXiv:2306.08000
-
Monolingual and Cross-Lingual Knowledge Transfer for Topic Classification 13 Jun 2023 · 0 repositories · arXiv:2306.07797
-
A Survey of Vision-Language Pre-training from the Lens of Multimodal Machine Translation 12 Jun 2023 · 0 repositories · arXiv:2306.07198
-
Imbalanced Multi-label Classification for Business-related Text with Moderately Large Label Spaces 12 Jun 2023 · 0 repositories · arXiv:2306.07046
-
Linear Classifier: An Often-Forgotten Baseline for Text Classification 12 Jun 2023 · 1 repository · arXiv:2306.07111Syntology official: no sample here; runs from other or unrecorded repositories · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples)
-
Multimodal Audio-textual Architecture for Robust Spoken Language Understanding 12 Jun 2023 · 0 repositories · arXiv:2306.06819
-
On the N-gram Approximation of Pre-trained Language Models 12 Jun 2023 · 0 repositories · arXiv:2306.06892
-
Recursion of Thought: A Divide-and-Conquer Approach to Multi-Context Reasoning with Language Models 12 Jun 2023 · 1 repository · arXiv:2306.06891
-
The BEA 2023 Shared Task on Generating AI Teacher Responses in Educational Dialogues 12 Jun 2023 · 0 repositories · arXiv:2306.06941
-
UniPoll: A Unified Social Media Poll Generation Framework via Multi-Objective Optimization 12 Jun 2023 · 1 repository · arXiv:2306.06851
-
Waffling around for Performance: Visual Classification with Random Words and Broad Concepts 12 Jun 2023 · 2 repositories · arXiv:2306.07282Syntology official (archive's flag): 4 ran · 5 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
CoTran: An LLM-based Code Translator using Reinforcement Learning with Feedback from Compiler and Symbolic Execution 11 Jun 2023 · 1 repository · arXiv:2306.06755
-
EaSyGuide : ESG Issue Identification Framework leveraging Abilities of Generative Large Language Models 11 Jun 2023 · 1 repository · arXiv:2306.06662
-
Inductive reasoning in humans and large language models 11 Jun 2023 · 1 repository · arXiv:2306.06548
-
RoBERTweet: A BERT Language Model for Romanian Tweets 11 Jun 2023 · 0 repositories · arXiv:2306.06598
-
Enhancing Low Resource NER Using Assisting Language And Transfer Learning 10 Jun 2023 · 0 repositories · arXiv:2306.06477
-
Medical Data Augmentation via ChatGPT: A Case Study on Medication Identification and Medication Event Classification 10 Jun 2023 · 0 repositories · arXiv:2306.07297
-
A Unified Generative Approach to Product Attribute-Value Identification 9 Jun 2023 · 0 repositories · arXiv:2306.05605
-
COVER: A Heuristic Greedy Adversarial Attack on Prompt-based Learning in Language Models 9 Jun 2023 · 0 repositories · arXiv:2306.05659
-
End-to-End Neural Network Compression via ℓ₁/ℓ₂ Regularized Latency Surrogates 9 Jun 2023 · 0 repositories · arXiv:2306.05785
-
Exploring the Responses of Large Language Models to Beginner Programmers' Help Requests 9 Jun 2023 · 0 repositories · arXiv:2306.05715
-
GPT-Calls: Enhancing Call Segmentation and Tagging by Generating Synthetic Conversations via Large Language Models 9 Jun 2023 · 0 repositories · arXiv:2306.07941
-
Implementing BERT and fine-tuned RobertA to detect AI generated news by ChatGPT 9 Jun 2023 · 0 repositories · arXiv:2306.07401
-
Language Models Can Learn Exceptions to Syntactic Rules 9 Jun 2023 · 1 repository · arXiv:2306.05969
-
Prodigy: An Expeditiously Adaptive Parameter-Free Learner 9 Jun 2023 · 1 repository · arXiv:2306.06101
-
Reliability Check: An Analysis of GPT-3's Response to Sensitive Topics and Prompt Wording 9 Jun 2023 · 2 repositories · arXiv:2306.06199