Methods › General › Distillation › SFT › Papers, page 5
Shrink and Fine-Tune
SFT
Papers archive 2025-07-28
archive papers tagged: 415 · with a code link: 204 · where Syntology ran a sample: 103 (86 with a run with no instrument failure, 17 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (103 of 415 tagged: 86 with a run with no instrument failure, 17 where every run was a failure of Syntology's instrument)
Page 5 of 5: papers 401 to 415 of 415, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
#InsTag: Instruction Tagging for Analyzing Supervised Fine-tuning of Large Language Models 14 Aug 2023 · 1 repository · arXiv:2308.07074
-
Zhongjing: Enhancing the Chinese Medical Capabilities of Large Language Model through Expert Feedback and Real-world Multi-turn Dialogue 7 Aug 2023 · 1 repository · arXiv:2308.03549
-
Okapi: Instruction-tuned Large Language Models in Multiple Languages with Reinforcement Learning from Human Feedback 29 Jul 2023 · 2 repositories · arXiv:2307.16039Syntology official: harvested, nothing ran · 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 1 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 11 harvested samples) · 7 pointer-only (licence)
-
MediaGPT : A Large Language Model For Chinese Media 20 Jul 2023 · 0 repositories · arXiv:2307.10930
-
Secrets of RLHF in Large Language Models Part I: PPO 11 Jul 2023 · 1 repository · arXiv:2307.04964
-
Preference Ranking Optimization for Human Alignment 30 Jun 2023 · 1 repository · arXiv:2306.17492
-
On the Uses of Large Language Models to Interpret Ambiguous Cyberattack Descriptions 24 Jun 2023 · 0 repositories · arXiv:2306.14062
-
Fine-tuning Language Models with Generative Adversarial Reward Modelling 9 May 2023 · 0 repositories · arXiv:2305.06176
-
RRHF: Rank Responses to Align Language Models with Human Feedback without tears 11 Apr 2023 · 1 repository · arXiv:2304.05302Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
SFT-KD-Recon: Learning a Student-friendly Teacher for Knowledge Distillation in Magnetic Resonance Image Reconstruction 11 Apr 2023 · 1 repository · arXiv:2304.05057
-
Optimizing DDPM Sampling with Shortcut Fine-Tuning 31 Jan 2023 · 1 repository · arXiv:2301.13362Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 5 unverified (of 9 harvested samples)
-
Self-Filtering: A Noise-Aware Sample Selection for Label Noise with Confidence Penalization 24 Aug 2022 · 1 repository · arXiv:2208.11351Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Cognitive Modeling of Semantic Fluency Using Transformers 20 Aug 2022 · 0 repositories · arXiv:2208.09719
-
p-Adic Statistical Field Theory and Deep Belief Networks 28 Jul 2022 · 0 repositories · arXiv:2207.13877
-
Pre-trained Summarization Distillation 24 Oct 2020 · 1 repository · arXiv:2010.13002