Methods › General › Distillation › SFT › Papers, page 3
Shrink and Fine-Tune
SFT
Papers archive 2025-07-28
archive papers tagged: 415 · with a code link: 204 · where Syntology ran a sample: 103 (86 with a run with no instrument failure, 17 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (103 of 415 tagged: 86 with a run with no instrument failure, 17 where every run was a failure of Syntology's instrument)
Page 3 of 5: papers 201 to 300 of 415, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Self-Improvement in Language Models: The Sharpening Mechanism 2 Dec 2024 · 0 repositories · arXiv:2412.01951
-
Continual SFT Matches Multimodal RLHF with Negative Supervision 22 Nov 2024 · 0 repositories · arXiv:2411.14797
-
Learning from "Silly" Questions Improves Large Language Models, But Only Slightly 21 Nov 2024 · 0 repositories · arXiv:2411.14121
-
Multimodal large language model for wheat breeding: a new exploration of smart breeding 20 Nov 2024 · 0 repositories · arXiv:2411.15203
-
On Active Privacy Auditing in Supervised Fine-tuning for White-Box Language Models 11 Nov 2024 · 0 repositories · arXiv:2411.07070
-
IOPO: Empowering LLMs with Complex Instruction Following via Input-Output Preference Optimization 9 Nov 2024 · 1 repository · arXiv:2411.06208
-
Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning 7 Nov 2024 · 0 repositories · arXiv:2411.05193
-
A Post-Training Enhanced Optimization Approach for Small Language Models 5 Nov 2024 · 0 repositories · arXiv:2411.02939
-
LLMs for Domain Generation Algorithm Detection 5 Nov 2024 · 0 repositories · arXiv:2411.03307
-
MetRex: A Benchmark for Verilog Code Metric Reasoning Using LLMs 5 Nov 2024 · 1 repository · arXiv:2411.03471Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
On the Loss of Context-awareness in General Instruction Fine-tuning 5 Nov 2024 · 1 repository · arXiv:2411.02688Syntology official: no sample here; runs from other or unrecorded repositories · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
L3Ms -- Lagrange Large Language Models 28 Oct 2024 · 0 repositories · arXiv:2410.21533
-
LongReward: Improving Long-context Large Language Models with AI Feedback 28 Oct 2024 · 1 repository · arXiv:2410.21252Syntology official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 1 pointer-only (licence)
-
Reducing the Scope of Language Models 28 Oct 2024 · 0 repositories · arXiv:2410.21597
-
UFT: Unifying Fine-Tuning of SFT and RLHF/DPO/UNA through a Generalized Implicit Reward Function 28 Oct 2024 · 0 repositories · arXiv:2410.21438
-
Fictitious Synthetic Data Can Improve LLM Factuality via Prerequisite Learning 25 Oct 2024 · 1 repository · arXiv:2410.19290
-
ToolFlow: Boosting LLM Tool-Calling Through Natural and Coherent Dialogue Synthesis 24 Oct 2024 · 0 repositories · arXiv:2410.18447
-
VoiceTextBlender: Augmenting Large Language Models with Speech Capabilities via Single-Stage Joint Speech-Text Supervised Fine-Tuning 23 Oct 2024 · 1 repository · arXiv:2410.17485
-
DEAN: Deactivating the Coupled Neurons to Mitigate Fairness-Privacy Conflicts in Large Language Models 22 Oct 2024 · 1 repository · arXiv:2410.16672
-
Bayesian scaling laws for in-context learning 21 Oct 2024 · 1 repository · arXiv:2410.16531Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Mitigating Forgetting in LLM Supervised Fine-Tuning and Preference Learning 20 Oct 2024 · 1 repository · arXiv:2410.15483
-
Training Language Models to Critique With Multi-agent Feedback 20 Oct 2024 · 0 repositories · arXiv:2410.15287
-
Persistent Pre-Training Poisoning of LLMs 17 Oct 2024 · 0 repositories · arXiv:2410.13722
-
RAG-DDR: Optimizing Retrieval-Augmented Generation Using Differentiable Data Rewards 17 Oct 2024 · 1 repository · arXiv:2410.13509Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples)
-
SemiEvol: Semi-supervised Fine-tuning for LLM Adaptation 17 Oct 2024 · 2 repositories · arXiv:2410.14745
-
RosePO: Aligning LLM-based Recommenders with Human Values 16 Oct 2024 · 0 repositories · arXiv:2410.12519
-
Improving the Language Understanding Capabilities of Large Language Models Using Reinforcement Learning 14 Oct 2024 · 0 repositories · arXiv:2410.11020
-
Self-Data Distillation for Recovering Quality in Pruned Large Language Models 13 Oct 2024 · 0 repositories · arXiv:2410.09982
-
\llinstruct: An Instruction-tuned model for English Language Proficiency Assessments 12 Oct 2024 · 0 repositories · arXiv:2410.09314
-
Rethinking Data Selection at Scale: Random Selection is Almost All You Need 12 Oct 2024 · 1 repository · arXiv:2410.09335Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Packing Analysis: Packing Is More Appropriate for Large Models or Datasets in Supervised Fine-tuning 10 Oct 2024 · 1 repository · arXiv:2410.08081Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Preference Fine-Tuning for Factuality in Chest X-Ray Interpretation Models Without Human Feedback 9 Oct 2024 · 0 repositories · arXiv:2410.07025
-
Towards Self-Improvement of LLMs via MCTS: Leveraging Stepwise Knowledge with Curriculum Preference Learning 9 Oct 2024 · 0 repositories · arXiv:2410.06508
-
Utilize the Flow before Stepping into the Same River Twice: Certainty Represented Knowledge Flow for Refusal-Aware Instruction Tuning 9 Oct 2024 · 1 repository · arXiv:2410.06913
-
Self-rationalization improves LLM as a fine-grained judge 7 Oct 2024 · 0 repositories · arXiv:2410.05495
-
SFTMix: Elevating Language Model Instruction Tuning with Mixup Recipe 7 Oct 2024 · 0 repositories · arXiv:2410.05248
-
How to Train Long-Context Language Models (Effectively) 3 Oct 2024 · 1 repository · arXiv:2410.02660Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 2 pointer-only (licence)
-
An Exploration of Self-Supervised Mutual Information Alignment for Multi-Task Settings 2 Oct 2024 · 1 repository · arXiv:2410.01704
-
FlashMask: Efficient and Rich Mask Extension of FlashAttention 2 Oct 2024 · 1 repository · arXiv:2410.01359Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Locret: Enhancing Eviction in Long-Context LLM Inference with Trained Retaining Heads on Consumer-Grade Devices 2 Oct 2024 · 1 repository · arXiv:2410.01805
-
OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data 2 Oct 2024 · 1 repository · arXiv:2410.01560Syntology 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 15 harvested samples)
-
Mitigating Training Imbalance in LLM Fine-Tuning via Selective Parameter Merging 1 Oct 2024 · 0 repositories · arXiv:2410.03743
-
Modulated Intervention Preference Optimization (MIPO): Keep the Easy, Refine the Difficult 26 Sep 2024 · 0 repositories · arXiv:2409.17545
-
60 Data Points are Sufficient to Fine-Tune LLMs for Question-Answering 24 Sep 2024 · 0 repositories · arXiv:2409.15825
-
Supervised Fine-Tuning Achieve Rapid Task Adaption Via Alternating Attention Head Activation Patterns 24 Sep 2024 · 0 repositories · arXiv:2409.15820
-
Backtracking Improves Generation Safety 22 Sep 2024 · 0 repositories · arXiv:2409.14586
-
A Case Study of Web App Coding with OpenAI Reasoning Models 19 Sep 2024 · 1 repository · arXiv:2409.13773
-
Fine Tuning Large Language Models for Medicine: The Role and Importance of Direct Preference Optimization 19 Sep 2024 · 0 repositories · arXiv:2409.12741
-
Training Language Models to Self-Correct via Reinforcement Learning 19 Sep 2024 · 2 repositories · arXiv:2409.12917
-
Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement 18 Sep 2024 · 0 repositories · arXiv:2409.12122
-
Generated Data with Fake Privacy: Hidden Dangers of Fine-tuning Large Language Models on Generated Data 12 Sep 2024 · 0 repositories · arXiv:2409.11423
-
Beyond IID: Optimizing Instruction Learning from the Perspective of Instruction Interaction and Dependency 11 Sep 2024 · 0 repositories · arXiv:2409.07045
-
Selective Self-Rehearsal: A Fine-Tuning Approach to Improve Generalization in Large Language Models 7 Sep 2024 · 0 repositories · arXiv:2409.04787
-
The representation landscape of few-shot learning and fine-tuning in large language models 5 Sep 2024 · 1 repository · arXiv:2409.03662Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-context QA 4 Sep 2024 · 1 repository · arXiv:2409.02897Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Surface Flux Transport Modeling using Physics Informed Neural Networks 3 Sep 2024 · 0 repositories · arXiv:2409.01744
-
Imitating Language via Scalable Inverse Reinforcement Learning 2 Sep 2024 · 0 repositories · arXiv:2409.01369
-
Entropic Distribution Matching in Supervised Fine-tuning of LLMs: Less Overfitting and Better Diversity 29 Aug 2024 · 0 repositories · arXiv:2408.16673
-
SciLitLLM: How to Adapt LLMs for Scientific Literature Understanding 28 Aug 2024 · 1 repository · arXiv:2408.15545Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 13 harvested samples)
-
Instruct-SkillMix: A Powerful Pipeline for LLM Instruction Tuning 27 Aug 2024 · 1 repository · arXiv:2408.14774Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Quality or Quantity? On Data Scale and Diversity in Adapting Large Language Models for Low-Resource Translation 23 Aug 2024 · 0 repositories · arXiv:2408.12780
-
Minor SFT loss for LLM fine-tune to increase performance and reduce model deviation 20 Aug 2024 · 0 repositories · arXiv:2408.10642
-
Threshold Filtering Packing for Supervised Fine-Tuning: Training Related Samples within Packs 18 Aug 2024 · 0 repositories · arXiv:2408.09327
-
API-guided Dataset Synthesis to Finetune Large Code Models 15 Aug 2024 · 0 repositories · arXiv:2408.08343
-
LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs 13 Aug 2024 · 3 repositories · arXiv:2408.07055
-
Sampling Foundational Transformer: A Theoretical Perspective 11 Aug 2024 · 0 repositories · arXiv:2408.05822
-
Intermediate direct preference optimization 6 Aug 2024 · 0 repositories · arXiv:2408.02923
-
A Framework for Fine-Tuning LLMs using Heterogeneous Feedback 5 Aug 2024 · 0 repositories · arXiv:2408.02861
-
Large Language Models for Anomaly Detection in Computational Workflows: from Supervised Fine-Tuning to In-Context Learning 24 Jul 2024 · 1 repository · arXiv:2407.17545
-
The Hitchhiker's Guide to Human Alignment with *PO 21 Jul 2024 · 0 repositories · arXiv:2407.15229
-
Does Refusal Training in LLMs Generalize to the Past Tense? 16 Jul 2024 · 1 repository · arXiv:2407.11969Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Enhancing Training Efficiency Using Packing with Flash Attention 12 Jul 2024 · 0 repositories · arXiv:2407.09105
-
Model Surgery: Modulating LLM's Behavior Via Simple Parameter Editing 11 Jul 2024 · 1 repository · arXiv:2407.08770Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
Skywork-Math: Data Scaling Laws for Mathematical Reasoning in Large Language Models -- The Story Goes On 11 Jul 2024 · 0 repositories · arXiv:2407.08348
-
Efficient and Accurate Memorable Conversation Model using DPO based on sLLM 9 Jul 2024 · 0 repositories · arXiv:2407.06537
-
LIONs: An Empirically Optimized Approach to Align Language Models 9 Jul 2024 · 1 repository · arXiv:2407.06542
-
Unlocking the Potential of Model Merging for Low-Resource Languages 4 Jul 2024 · 0 repositories · arXiv:2407.03994
-
52B to 1T: Lessons Learned via Tele-FLM Series 3 Jul 2024 · 0 repositories · arXiv:2407.02783
-
Raw Text is All you Need: Knowledge-intensive Multi-turn Instruction Tuning for Large Language Model 3 Jul 2024 · 0 repositories · arXiv:2407.03040
-
Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning 30 Jun 2024 · 0 repositories · arXiv:2407.00617
-
Step-Controlled DPO: Leveraging Stepwise Error for Enhanced Mathematical Reasoning 30 Jun 2024 · 1 repository · arXiv:2407.00782Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Understanding and Mitigating Language Confusion in LLMs 28 Jun 2024 · 1 repository · arXiv:2406.20052Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Suri: Multi-constraint Instruction Following for Long-form Text Generation 27 Jun 2024 · 1 repository · arXiv:2406.19371Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 1 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 15 harvested samples) · 15 pointer-only (licence)
-
Open-vocabulary Mobile Manipulation in Unseen Dynamic Environments with 3D Semantic Maps 26 Jun 2024 · 0 repositories · arXiv:2406.18115
-
ARES: Alternating Reinforcement Learning and Supervised Fine-Tuning for Enhanced Multi-Modal Chain-of-Thought Reasoning Through Diverse AI Feedback 25 Jun 2024 · 1 repository · arXiv:2407.00087
-
PAFT: A Parallel Training Paradigm for Effective LLM Fine-Tuning 25 Jun 2024 · 0 repositories · arXiv:2406.17923
-
Does Cross-Cultural Alignment Change the Commonsense Morality of Language Models? 24 Jun 2024 · 0 repositories · arXiv:2406.16316
-
RuleR: Improving LLM Controllability by Rule-based Data Recycling 22 Jun 2024 · 3 repositories · arXiv:2406.15938
-
Gradient-Mask Tuning Elevates the Upper Limits of LLM Performance 21 Jun 2024 · 0 repositories · arXiv:2406.15330
-
SeCoKD: Aligning Large Language Models for In-Context Learning with Fewer Shots 20 Jun 2024 · 0 repositories · arXiv:2406.14208
-
Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models 19 Jun 2024 · 1 repository · arXiv:2406.13542Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Aqulia-Med LLM: Pioneering Full-Process Open-Source Medical Language Models 18 Jun 2024 · 0 repositories · arXiv:2406.12182
-
Recognition of Dynamic Hand Gestures in Long Distance using a Web-Camera for Robot Guidance 18 Jun 2024 · 0 repositories · arXiv:2406.12424
-
A Systematic Analysis of Large Language Models as Soft Reasoners: The Case of Syllogistic Inferences 17 Jun 2024 · 1 repository · arXiv:2406.11341
-
Preserving Knowledge in Large Language Model with Model-Agnostic Self-Decompression 17 Jun 2024 · 0 repositories · arXiv:2406.11354
-
Self-Evolution Fine-Tuning for Policy Optimization 16 Jun 2024 · 0 repositories · arXiv:2406.10813
-
SyntheT2C: Generating Synthetic Data for Fine-Tuning Large Language Models on the Text2Cypher Task 15 Jun 2024 · 1 repository · arXiv:2406.10710
-
Unlock the Correlation between Supervised Fine-Tuning and Reinforcement Learning in Training Code Large Language Models 14 Jun 2024 · 0 repositories · arXiv:2406.10305
-
Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing 12 Jun 2024 · 2 repositories · arXiv:2406.08464Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
PLUM: Improving Code LMs with Execution-Guided On-Policy Preference Learning Driven By Synthetic Test Cases 11 Jun 2024 · 0 repositories · arXiv:2406.06887