Methods › General › Learning Rate Schedules › Cosine Annealing › Papers, page 7
Cosine Annealing
Papers archive 2025-07-28
archive papers tagged: 3,965 · with a code link: 1,734 · where Syntology ran a sample: 627 (513 with a run with no instrument failure, 114 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (627 of 3,965 tagged: 513 with a run with no instrument failure, 114 where every run was a failure of Syntology's instrument)
Page 7 of 40: papers 601 to 700 of 3,965, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Agent Skill Acquisition for Large Language Models via CycleQD 16 Oct 2024 · 1 repository · arXiv:2410.14735Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Context-Scaling versus Task-Scaling in In-Context Learning 16 Oct 2024 · 0 repositories · arXiv:2410.12783
-
FusionLLM: A Decentralized LLM Training System on Geo-distributed GPUs with Adaptive Compression 16 Oct 2024 · 0 repositories · arXiv:2410.12707
-
Kallini et al. (2024) do not compare impossible languages with constituency-based ones 16 Oct 2024 · 0 repositories · arXiv:2410.12271
-
ShapefileGPT: A Multi-Agent Large Language Model Framework for Automated Shapefile Processing 16 Oct 2024 · 0 repositories · arXiv:2410.12376
-
Stabilize the Latent Space for Image Autoregressive Modeling: A Unified Perspective 16 Oct 2024 · 1 repository · arXiv:2410.12490Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples)
-
Table-LLM-Specialist: Language Model Specialists for Tables using Iterative Generator-Validator Fine-tuning 16 Oct 2024 · 0 repositories · arXiv:2410.12164
-
When Not to Answer: Evaluating Prompts on GPT Models for Effective Abstention in Unanswerable Math Word Problems 16 Oct 2024 · 0 repositories · arXiv:2410.13029
-
Deciphering the Chaos: Enhancing Jailbreak Attacks via Adversarial Prompt Translation 15 Oct 2024 · 1 repository · arXiv:2410.11317Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Evidence of Cognitive Deficits andDevelopmental Advances in Generative AI: A Clock Drawing Test Analysis 15 Oct 2024 · 0 repositories · arXiv:2410.11756
-
"Is Hate Lost in Translation?": Evaluation of Multilingual LGBTQIA+ Hate Speech Detection 15 Oct 2024 · 0 repositories · arXiv:2410.11230
-
MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language Models 15 Oct 2024 · 1 repository · arXiv:2410.11710
-
Nonlinear Gaussian process tomography with imposed non-negativity constraints on physical quantities for plasma diagnostics 15 Oct 2024 · 0 repositories · arXiv:2410.11454
-
SEER: Self-Aligned Evidence Extraction for Retrieval-Augmented Generation 15 Oct 2024 · 0 repositories · arXiv:2410.11315
-
Code-Mixer Ya Nahi: Novel Approaches to Measuring Multilingual LLMs' Code-Mixing Capabilities 14 Oct 2024 · 0 repositories · arXiv:2410.11079
-
Double Jeopardy and Climate Impact in the Use of Large Language Models: Socio-economic Disparities and Reduced Utility for Non-English Speakers 14 Oct 2024 · 1 repository · arXiv:2410.10665
-
One Language, Many Gaps: Evaluating Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks 14 Oct 2024 · 1 repository · arXiv:2410.11005Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples)
-
Performance in a dialectal profiling task of LLMs for varieties of Brazilian Portuguese 14 Oct 2024 · 0 repositories · arXiv:2410.10991
-
Rethinking Legal Judgement Prediction in a Realistic Scenario in the Era of Large Language Models 14 Oct 2024 · 1 repository · arXiv:2410.10542
-
Towards Better Multi-head Attention via Channel-wise Sample Permutation 14 Oct 2024 · 1 repository · arXiv:2410.10914Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Can In-context Learning Really Generalize to Out-of-distribution Tasks? 13 Oct 2024 · 0 repositories · arXiv:2410.09695
-
Evaluating Gender Bias of LLMs in Making Morality Judgements 13 Oct 2024 · 0 repositories · arXiv:2410.09992
-
Investigating Implicit Bias in Large Language Models: A Large-Scale Study of Over 50 LLMs 13 Oct 2024 · 0 repositories · arXiv:2410.12864
-
M2M-Gen: A Multimodal Framework for Automated Background Music Generation in Japanese Manga Using Large Language Models 13 Oct 2024 · 0 repositories · arXiv:2410.09928
-
\llinstruct: An Instruction-tuned model for English Language Proficiency Assessments 12 Oct 2024 · 0 repositories · arXiv:2410.09314
-
AttnGCG: Enhancing Jailbreaking Attacks on LLMs with Attention Manipulation 11 Oct 2024 · 1 repository · arXiv:2410.09040
-
Extra Global Attention Designation Using Keyword Detection in Sparse Transformer Architectures 11 Oct 2024 · 0 repositories · arXiv:2410.08971
-
Fine-Tuning In-House Large Language Models to Infer Differential Diagnosis from Radiology Reports 11 Oct 2024 · 0 repositories · arXiv:2410.09234
-
Humanity in AI: Detecting the Personality of Large Language Models 11 Oct 2024 · 0 repositories · arXiv:2410.08545
-
Observing the Southern US Culture of Honor Using Large-Scale Social Media Analysis 11 Oct 2024 · 0 repositories · arXiv:2410.13887
-
SocialGaze: Improving the Integration of Human Social Norms in Large Language Models 11 Oct 2024 · 1 repository · arXiv:2410.08698
-
Synth-SONAR: Sonar Image Synthesis with Enhanced Diversity and Realism via Dual Diffusion Models and GPT Prompting 11 Oct 2024 · 1 repository · arXiv:2410.08612
-
Adam Exploits ℓ_∞-geometry of Loss Landscape via Coordinate-wise Adaptivity 10 Oct 2024 · 1 repository · arXiv:2410.08198Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples)
-
FLIER: Few-shot Language Image Models Embedded with Latent Representations 10 Oct 2024 · 0 repositories · arXiv:2410.07648
-
The Rise of AI-Generated Content in Wikipedia 10 Oct 2024 · 1 repository · arXiv:2410.08044Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
AutoFeedback: An LLM-based Framework for Efficient and Accurate API Request Generation 9 Oct 2024 · 0 repositories · arXiv:2410.06943
-
Capturing Bias Diversity in LLMs 9 Oct 2024 · 0 repositories · arXiv:2410.12839
-
Generative Model for Less-Resourced Language with 1 billion parameters 9 Oct 2024 · 0 repositories · arXiv:2410.06898
-
Large Language Models as Code Executors: An Exploratory Study 9 Oct 2024 · 0 repositories · arXiv:2410.06667
-
MentalArena: Self-play Training of Language Models for Diagnosis and Treatment of Mental Health Disorders 9 Oct 2024 · 1 repository · arXiv:2410.06845
-
SAGE: Scalable Ground Truth Evaluations for Large Sparse Autoencoders 9 Oct 2024 · 0 repositories · arXiv:2410.07456
-
A second-order-like optimizer with adaptive gradient scaling for deep learning 8 Oct 2024 · 1 repository · arXiv:2410.05871
-
Auto-Evolve: Enhancing Large Language Model's Performance via Self-Reasoning Framework 8 Oct 2024 · 0 repositories · arXiv:2410.06328
-
Coevolving with the Other You: Fine-Tuning LLM with Sequential Cooperative Multi-Agent Reinforcement Learning 8 Oct 2024 · 1 repository · arXiv:2410.06101Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 11 harvested samples) · 1 pointer-only (licence)
-
Leveraging free energy in pretraining model selection for improved fine-tuning 8 Oct 2024 · 0 repositories · arXiv:2410.05612
-
AnyAttack: Towards Large-scale Self-supervised Adversarial Attacks on Vision-language Models 7 Oct 2024 · 0 repositories · arXiv:2410.05346
-
LPZero: Language Model Zero-cost Proxy Search from Zero 7 Oct 2024 · 0 repositories · arXiv:2410.04808
-
Narrative-of-Thought: Improving Temporal Reasoning of Large Language Models via Recounted Narratives 7 Oct 2024 · 1 repository · arXiv:2410.05558
-
On Instruction-Finetuning Neural Machine Translation Models 7 Oct 2024 · 0 repositories · arXiv:2410.05553
-
FAMMA: A Benchmark for Financial Domain Multilingual Multimodal Question Answering 6 Oct 2024 · 1 repository · arXiv:2410.04526Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective 6 Oct 2024 · 1 repository · arXiv:2410.04466
-
Large Language Models for Knowledge-Free Network Management: Feasibility Study and Opportunities 6 Oct 2024 · 0 repositories · arXiv:2410.17259
-
ProtocoLLM: Automatic Evaluation Framework of LLMs on Domain-Specific Scientific Protocol Formulation Tasks 6 Oct 2024 · 0 repositories · arXiv:2410.04601
-
Gamified crowd-sourcing of high-quality data for visual fine-tuning 5 Oct 2024 · 0 repositories · arXiv:2410.04038
-
Take It Easy: Label-Adaptive Self-Rationalization for Fact Verification and Explanation Generation 5 Oct 2024 · 1 repository · arXiv:2410.04002Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Crafting Narrative Closures: Zero-Shot Learning with SSM Mamba for Short Story Ending Generation 4 Oct 2024 · 0 repositories · arXiv:2410.10848
-
Cross-lingual Transfer for Automatic Question Generation by Learning Interrogative Structures in Target Languages 4 Oct 2024 · 0 repositories · arXiv:2410.03197
-
How Language Models Prioritize Contextual Grammatical Cues? 4 Oct 2024 · 1 repository · arXiv:2410.03447
-
Steering Large Language Models between Code Execution and Textual Reasoning 4 Oct 2024 · 1 repository · arXiv:2410.03524
-
Structured List-Grounded Question Answering 4 Oct 2024 · 0 repositories · arXiv:2410.03950
-
Towards Linguistically-Aware and Language-Independent Tokenization for Large Language Models (LLMs) 4 Oct 2024 · 0 repositories · arXiv:2410.03568
-
Using Prompts to Guide Large Language Models in Imitating a Real Person's Language Style 4 Oct 2024 · 0 repositories · arXiv:2410.03848
-
AlphaIntegrator: Transformer Action Search for Symbolic Integration Proofs 3 Oct 2024 · 0 repositories · arXiv:2410.02666
-
CodeJudge: Evaluating Code Generation with Large Language Models 3 Oct 2024 · 1 repository · arXiv:2410.02184Syntology official (archive's flag): 13 ran · 15 ran (of which 0 constructed an object rather than computing a result; 14 with no instrument failure: 2 honoured, 0 violated, 12 with no contract checked; 1 where Syntology's instrument failed) · 8 unverified (of 23 harvested samples) · 2 pointer-only (licence)
-
LLaVA-Critic: Learning to Evaluate Multimodal Models 3 Oct 2024 · 0 repositories · arXiv:2410.02712
-
Plots Unlock Time-Series Understanding in Multimodal Models 3 Oct 2024 · 0 repositories · arXiv:2410.02637
-
Visual Editing with LLM-based Tool Chaining: An Efficient Distillation Approach for Real-Time Applications 3 Oct 2024 · 1 repository · arXiv:2410.02952
-
Automatic deductive coding in discourse analysis: an application of large language models in learning analytics 2 Oct 2024 · 1 repository · arXiv:2410.01240
-
Emotion-Aware Embedding Fusion in LLMs (Flan-T5, LLAMA 2, DeepSeek-R1, and ChatGPT 4) for Intelligent Response Generation 2 Oct 2024 · 0 repositories · arXiv:2410.01306
-
Enhancing LLM Fine-tuning for Text-to-SQLs by SQL Quality Measurement 2 Oct 2024 · 0 repositories · arXiv:2410.01869
-
On The Adaptation of Unlimiformer for Decoder-Only Transformers 2 Oct 2024 · 0 repositories · arXiv:2410.01637
-
Quantifying Generalization Complexity for Large Language Models 2 Oct 2024 · 1 repository · arXiv:2410.01769Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Seeing Eye to AI: Human Alignment via Gaze-Based Response Rewards for Large Language Models 2 Oct 2024 · 1 repository · arXiv:2410.01532Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
AlignSum: Data Pyramid Hierarchical Fine-tuning for Aligning with Human Summarization Preference 1 Oct 2024 · 1 repository · arXiv:2410.00409Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Decoding Hate: Exploring Language Models' Reactions to Hate Speech 1 Oct 2024 · 0 repositories · arXiv:2410.00775
-
Language Enhanced Model for Eye (LEME): An Open-Source Ophthalmology-Specific Large Language Model 1 Oct 2024 · 0 repositories · arXiv:2410.03740
-
Sparse Attention Decomposition Applied to Circuit Tracing 1 Oct 2024 · 1 repository · arXiv:2410.00340
-
A Looming Replication Crisis in Evaluating Behavior in Language Models? Evidence and Solutions 30 Sep 2024 · 0 repositories · arXiv:2409.20303
-
Adapting LLMs for the Medical Domain in Portuguese: A Study on Fine-Tuning and Model Evaluation 30 Sep 2024 · 0 repositories · arXiv:2410.00163
-
Evaluating the fairness of task-adaptive pretraining on unlabeled test data before few-shot text classification 30 Sep 2024 · 1 repository · arXiv:2410.00179
-
Modelando procesos cognitivos de la lectura natural con GPT-2 30 Sep 2024 · 0 repositories · arXiv:2409.20174
-
Analog In-Memory Computing Attention Mechanism for Fast and Energy-Efficient Large Language Models 28 Sep 2024 · 1 repository · arXiv:2409.19315
-
Charting the Future: Using Chart Question-Answering for Scalable Evaluation of LLM-Driven Data Visualizations 27 Sep 2024 · 0 repositories · arXiv:2409.18764
-
Cottention: Linear Transformers With Cosine Attention 27 Sep 2024 · 1 repository · arXiv:2409.18747Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Experimental Evaluation of Machine Learning Models for Goal-oriented Customer Service Chatbot with Pipeline Architecture 27 Sep 2024 · 0 repositories · arXiv:2409.18568
-
Comparing Unidirectional, Bidirectional, and Word2vec Models for Discovering Vulnerabilities in Compiled Lifted Code 26 Sep 2024 · 0 repositories · arXiv:2409.17513
-
Efficient In-Domain Question Answering for Resource-Constrained Environments 26 Sep 2024 · 0 repositories · arXiv:2409.17648
-
MaskLLM: Learnable Semi-Structured Sparsity for Large Language Models 26 Sep 2024 · 1 repository · arXiv:2409.17481Syntology official (archive's flag): 5 ran · 5 ran (of which 1 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 11 unverified (of 16 harvested samples) · 16 pointer-only (licence)
-
T3: A Novel Zero-shot Transfer Learning Framework Iteratively Training on an Assistant Task for a Target Task 26 Sep 2024 · 0 repositories · arXiv:2409.17640
-
The application of GPT-4 in grading design university students' assignment and providing feedback: An exploratory study 26 Sep 2024 · 0 repositories · arXiv:2409.17698
-
A Prompting-Based Representation Learning Method for Recommendation with Large Language Models 25 Sep 2024 · 0 repositories · arXiv:2409.16674
-
Deep Learning and Machine Learning, Advancing Big Data Analytics and Management: Handy Appetizer 25 Sep 2024 · 0 repositories · arXiv:2409.17120
-
Severity Prediction in Mental Health: LLM-based Creation, Analysis, Evaluation of a Novel Multilingual Dataset 25 Sep 2024 · 0 repositories · arXiv:2409.17397
-
Using LLM for Real-Time Transcription and Summarization of Doctor-Patient Interactions into ePuskesmas in Indonesia 25 Sep 2024 · 0 repositories · arXiv:2409.17054
-
AI Can Be Cognitively Biased: An Exploratory Study on Threshold Priming in LLM-Based Batch Relevance Assessment 24 Sep 2024 · 0 repositories · arXiv:2409.16022
-
Data Augmentation for Sparse Multidimensional Learning Performance Data Using Generative AI 24 Sep 2024 · 1 repository · arXiv:2409.15631
-
Effectiveness of Cross-linguistic Extraction of Genetic Information using Generative Large Language Models 24 Sep 2024 · 1 repository
-
Selection of Prompt Engineering Techniques for Code Generation through Predicting Code Complexity 24 Sep 2024 · 0 repositories · arXiv:2409.16416
-
Self-attention as an attractor network: transient memories without backpropagation 24 Sep 2024 · 1 repository · arXiv:2409.16112
-
Synatra: Turning Indirect Knowledge into Direct Demonstrations for Digital Agents at Scale 24 Sep 2024 · 0 repositories · arXiv:2409.15637