Methods › General › Distillation › SFT › Papers where code ran, page 1
Shrink and Fine-Tune
SFT
Papers archive 2025-07-28
archive papers tagged: 415 · with a code link: 204 · where Syntology ran a sample: 103 (86 with a run with no instrument failure, 17 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (103 of 415 tagged: 86 with a run with no instrument failure, 17 where every run was a failure of Syntology's instrument)
Syntology We ran code from the paper's repository; we did not isolate this method inside it.
Page 1 of 2: papers 1 to 100 of the 103 tagged papers where Syntology ran at least one harvested sample (86 with a run with no instrument failure, 17 where every run was a failure of Syntology's instrument), newest first by the archive's date (ties by arXiv id). This is a filter on Syntology's record ordered by date only, not a ranking; a run is not a correctness claim. A paper missing from this list is not a recorded non-run: it may have no arXiv id, no harvested code, or only samples that have not run yet.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
A Practical Two-Stage Recipe for Mathematical LLMs: Maximizing Accuracy with SFT and Efficiency with Reinforcement Learning 11 Jul 2025 · 1 repository · arXiv:2507.08267Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training 7 Jul 2025 · 0 repositories · arXiv:2507.05386Syntology 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning 30 Jun 2025 · 1 repository · arXiv:2506.24119Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples)
-
LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning 23 Jun 2025 · 1 repository · arXiv:2506.18841Syntology 12 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 6 where Syntology's instrument failed) · 2 unverified (of 14 harvested samples) · 1 pointer-only (licence)
-
QFFT, Question-Free Fine-Tuning for Adaptive Reasoning 15 Jun 2025 · 1 repository · arXiv:2506.12860Syntology official (archive's flag): 14 ran · 14 ran (of which 0 constructed an object rather than computing a result; 14 with no instrument failure: 0 honoured, 0 violated, 14 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 14 harvested samples) · 7 pointer-only (licence)
-
cadrille: Multi-modal CAD Reconstruction with Online Reinforcement Learning 28 May 2025 · 1 repository · arXiv:2505.22914Syntology 2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; every one of the 2 samples that ran constructed an object rather than computing a result (of 3 harvested samples)
-
SimpleDeepSearcher: Deep Information Seeking via Web-Powered Reasoning Trajectory Synthesis 22 May 2025 · 3 repositories · arXiv:2505.16834Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
UFT: Unifying Supervised and Reinforcement Fine-Tuning 22 May 2025 · 1 repository · arXiv:2505.16984Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 0 violated, 2 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 9 harvested samples)
-
R1-Searcher++: Incentivizing the Dynamic Knowledge Acquisition of LLMs via Reinforcement Learning 22 May 2025 · 3 repositories · arXiv:2505.17005Syntology community repositories only · 9 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 10 harvested samples)
-
JOLT-SQL: Joint Loss Tuning of Text-to-SQL with Confusion-aware Noisy Schema Sampling 20 May 2025 · 1 repository · arXiv:2505.14305Syntology official: no sample here; runs from other or unrecorded repositories · 16 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 3 violated, 9 with no contract checked; 4 where Syntology's instrument failed) · 2 unverified (of 18 harvested samples) · 1 pointer-only (licence)
-
DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning 20 May 2025 · 1 repository · arXiv:2505.14362Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 1 honoured, 0 violated, 8 with no contract checked; 3 where Syntology's instrument failed) · 5 unverified (of 17 harvested samples) · 4 pointer-only (licence)
-
VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-Tuning 18 May 2025 · 1 repository · arXiv:2505.12434Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 3 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples) · 2 pointer-only (licence)
-
GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning 16 May 2025 · 1 repository · arXiv:2505.11049Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 12 harvested samples) · 1 pointer-only (licence)
-
Compile Scene Graphs with Reinforcement Learning 18 Apr 2025 · 1 repository · arXiv:2504.13617Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
ToolRL: Reward is All Tool Learning Needs 16 Apr 2025 · 1 repository · arXiv:2504.13958Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
MM-IFEngine: Towards Multimodal Instruction Following 10 Apr 2025 · 1 repository · arXiv:2504.07957Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples)
-
AnesBench: Multi-Dimensional Evaluation of LLM Reasoning in Anesthesiology 3 Apr 2025 · 1 repository · arXiv:2504.02404Syntology official: no sample here; runs from other or unrecorded repositories · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Do Theory of Mind Benchmarks Need Explicit Human-like Reasoning in Language Models? 2 Apr 2025 · 1 repository · arXiv:2504.01698Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 1 honoured, 0 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 15 harvested samples)
-
Rec-R1: Bridging Generative Large Language Models and User-Centric Recommendation Systems via Reinforcement Learning 31 Mar 2025 · 1 repository · arXiv:2503.24289Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Exploring the Effect of Reinforcement Learning on Video Understanding: Insights from SEED-Bench-R1 31 Mar 2025 · 1 repository · arXiv:2503.24376Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 3 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples)
-
Video-R1: Reinforcing Video Reasoning in MLLMs 27 Mar 2025 · 1 repository · arXiv:2503.21776Syntology community repositories only · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning 26 Mar 2025 · 1 repository · arXiv:2503.20752Syntology 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 8 harvested samples)
-
OpenVLThinker: An Early Exploration to Complex Vision-Language Reasoning via Iterative Self-Improvement 21 Mar 2025 · 1 repository · arXiv:2503.17352Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 12 unverified (of 21 harvested samples) · 2 pointer-only (licence)
-
Think or Not Think: A Study of Explicit Thinking in Rule-Based Visual Reinforcement Fine-Tuning 20 Mar 2025 · 1 repository · arXiv:2503.16188Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning 18 Mar 2025 · 1 repository · arXiv:2503.15558Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond 13 Mar 2025 · 1 repository · arXiv:2503.10460Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 1 honoured, 1 violated, 6 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 13 harvested samples)
-
AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning 10 Mar 2025 · 1 repository · arXiv:2503.07608Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 3 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model 10 Mar 2025 · 1 repository · arXiv:2503.07703Syntology 11 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples) · 2 pointer-only (licence)
-
R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model 7 Mar 2025 · 1 repository · arXiv:2503.05132Syntology official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
LongWriter-V: Enabling Ultra-Long and High-Fidelity Generation in Vision-Language Models 20 Feb 2025 · 1 repository · arXiv:2502.14834Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
LongPO: Long Context Self-Evolution of Large Language Models through Short-to-Long Preference Optimization 19 Feb 2025 · 1 repository · arXiv:2502.13922Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Robust Adaptation of Large Multimodal Models for Retrieval Augmented Hateful Meme Detection 18 Feb 2025 · 1 repository · arXiv:2502.13061Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples)
-
Large Language Diffusion Models 14 Feb 2025 · 2 repositories · arXiv:2502.09992Syntology 6 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Demystifying Long Chain-of-Thought Reasoning in LLMs 5 Feb 2025 · 1 repository · arXiv:2502.03373Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples)
-
Process Reinforcement through Implicit Rewards 3 Feb 2025 · 5 repositories · arXiv:2502.01456Syntology official (archive's flag): 14 ran · 14 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 17 harvested samples) · 2 pointer-only (licence)
-
GuardReasoner: Towards Reasoning-based LLM Safeguards 30 Jan 2025 · 1 repository · arXiv:2501.18492Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
WILDCHAT-50M: A Deep Dive Into the Role of Synthetic Data in Post-Training 30 Jan 2025 · 1 repository · arXiv:2501.18511Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples)
-
Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate 29 Jan 2025 · 1 repository · arXiv:2501.17703Syntology 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding 22 Jan 2025 · 1 repository · arXiv:2501.13106Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 5 where Syntology's instrument failed) · 6 unverified (of 14 harvested samples) · 2 pointer-only (licence)
-
Confidence v.s. Critique: A Decomposition of Self-Correction Capability for LLMs 27 Dec 2024 · 1 repository · arXiv:2412.19513Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search 24 Dec 2024 · 2 repositories · arXiv:2412.18319Syntology community repositories only · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples)
-
Preference-Oriented Supervised Fine-Tuning: Favoring Target Model Over Aligned Large Language Models 17 Dec 2024 · 1 repository · arXiv:2412.12865Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
ROUTE: Robust Multitask Tuning and Collaboration for Text-to-SQL 13 Dec 2024 · 1 repository · arXiv:2412.10138Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 2 pointer-only (licence)
-
SPRec: Leveraging Self-Play to Debias Preference Alignment for Large Language Model-based Recommendations 12 Dec 2024 · 1 repository · arXiv:2412.09243Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
On the Loss of Context-awareness in General Instruction Fine-tuning 5 Nov 2024 · 1 repository · arXiv:2411.02688Syntology official: no sample here; runs from other or unrecorded repositories · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
MetRex: A Benchmark for Verilog Code Metric Reasoning Using LLMs 5 Nov 2024 · 1 repository · arXiv:2411.03471Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
LongReward: Improving Long-context Large Language Models with AI Feedback 28 Oct 2024 · 1 repository · arXiv:2410.21252Syntology official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 1 pointer-only (licence)
-
Bayesian scaling laws for in-context learning 21 Oct 2024 · 1 repository · arXiv:2410.16531Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
RAG-DDR: Optimizing Retrieval-Augmented Generation Using Differentiable Data Rewards 17 Oct 2024 · 1 repository · arXiv:2410.13509Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples)
-
Rethinking Data Selection at Scale: Random Selection is Almost All You Need 12 Oct 2024 · 1 repository · arXiv:2410.09335Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Packing Analysis: Packing Is More Appropriate for Large Models or Datasets in Supervised Fine-tuning 10 Oct 2024 · 1 repository · arXiv:2410.08081Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
How to Train Long-Context Language Models (Effectively) 3 Oct 2024 · 1 repository · arXiv:2410.02660Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 2 pointer-only (licence)
-
FlashMask: Efficient and Rich Mask Extension of FlashAttention 2 Oct 2024 · 1 repository · arXiv:2410.01359Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data 2 Oct 2024 · 1 repository · arXiv:2410.01560Syntology 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 15 harvested samples)
-
The representation landscape of few-shot learning and fine-tuning in large language models 5 Sep 2024 · 1 repository · arXiv:2409.03662Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-context QA 4 Sep 2024 · 1 repository · arXiv:2409.02897Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
SciLitLLM: How to Adapt LLMs for Scientific Literature Understanding 28 Aug 2024 · 1 repository · arXiv:2408.15545Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 13 harvested samples)
-
Instruct-SkillMix: A Powerful Pipeline for LLM Instruction Tuning 27 Aug 2024 · 1 repository · arXiv:2408.14774Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Does Refusal Training in LLMs Generalize to the Past Tense? 16 Jul 2024 · 1 repository · arXiv:2407.11969Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Model Surgery: Modulating LLM's Behavior Via Simple Parameter Editing 11 Jul 2024 · 1 repository · arXiv:2407.08770Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
Step-Controlled DPO: Leveraging Stepwise Error for Enhanced Mathematical Reasoning 30 Jun 2024 · 1 repository · arXiv:2407.00782Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Understanding and Mitigating Language Confusion in LLMs 28 Jun 2024 · 1 repository · arXiv:2406.20052Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Suri: Multi-constraint Instruction Following for Long-form Text Generation 27 Jun 2024 · 1 repository · arXiv:2406.19371Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 1 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 15 harvested samples) · 15 pointer-only (licence)
-
Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models 19 Jun 2024 · 1 repository · arXiv:2406.13542Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing 12 Jun 2024 · 2 repositories · arXiv:2406.08464Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
A Fine-tuning Dataset and Benchmark for Large Language Models for Protein Understanding 8 Jun 2024 · 1 repository · arXiv:2406.05540Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
Parrot: Multilingual Visual Instruction Tuning 4 Jun 2024 · 2 repositories · arXiv:2406.02539Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 4 unverified (of 7 harvested samples)
-
PediatricsGPT: Large Language Models as Chinese Medical Assistants for Pediatric Applications 29 May 2024 · 1 repository · arXiv:2405.19266Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 1 pointer-only (licence)
-
Getting More Juice Out of the SFT Data: Reward Learning from Human Demonstration Improves SFT for LLM Alignment 28 May 2024 · 1 repository · arXiv:2405.17888Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (of 8 harvested samples) · 1 pointer-only (licence)
-
Triple Preference Optimization: Achieving Better Alignment with Less Data in a Single Step Optimization 26 May 2024 · 1 repository · arXiv:2405.16681Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples)
-
Disperse-Then-Merge: Pushing the Limits of Instruction Tuning via Alignment Tax Reduction 22 May 2024 · 1 repository · arXiv:2405.13432Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Balancing Speciality and Versatility: a Coarse to Fine Framework for Supervised Fine-tuning Large Language Model 16 Apr 2024 · 1 repository · arXiv:2404.10306Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Detoxifying Large Language Models via Knowledge Editing 21 Mar 2024 · 1 repository · arXiv:2403.14472Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 4 harvested samples)
-
SmallToLarge (S2L): Scalable Data Selection for Fine-tuning Large Language Models by Summarizing Training Trajectories of Small Models 12 Mar 2024 · 1 repository · arXiv:2403.07384Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 3 pointer-only (licence)
-
ORPO: Monolithic Preference Optimization without Reference Model 12 Mar 2024 · 4 repositories · arXiv:2403.07691Syntology official (archive's flag): 2 ran · 11 ran (of which 1 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 0 violated, 9 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 14 harvested samples) · 2 pointer-only (licence)
-
Unfamiliar Finetuning Examples Control How Language Models Hallucinate 8 Mar 2024 · 1 repository · arXiv:2403.05612Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Common 7B Language Models Already Possess Strong Math Capabilities 7 Mar 2024 · 2 repositories · arXiv:2403.04706Syntology official (archive's flag): 9 ran · 21 ran (of which 0 constructed an object rather than computing a result; 20 with no instrument failure: 0 honoured, 0 violated, 20 with no contract checked; 1 where Syntology's instrument failed) · 7 unverified (of 28 harvested samples) · 12 pointer-only (licence)
-
DACO: Towards Application-Driven and Comprehensive Data Analysis via Code Generation 4 Mar 2024 · 1 repository · arXiv:2403.02528Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Reflect-RL: Two-Player Online RL Fine-Tuning for LMs 20 Feb 2024 · 1 repository · arXiv:2402.12621Syntology official (archive's flag): 2 ran · 2 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 6 harvested samples)
-
A Critical Evaluation of AI Feedback for Aligning Large Language Models 19 Feb 2024 · 1 repository · arXiv:2402.12366Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples) · 1 pointer-only (licence)
-
UltraLink: An Open-Source Knowledge-Enhanced Multilingual Supervised Fine-tuning Dataset 7 Feb 2024 · 1 repository · arXiv:2402.04588Syntology official: no sample here; runs from other or unrecorded repositories · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Scaling Sparse Fine-Tuning to Large Language Models 29 Jan 2024 · 2 repositories · arXiv:2401.16405Syntology official (archive's flag): 12 ran · 13 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 2 honoured, 0 violated, 7 with no contract checked; 4 where Syntology's instrument failed) · 2 unverified (of 15 harvested samples) · 5 pointer-only (licence)
-
Supervised Fine-tuning in turn Improves Visual Foundation Models 18 Jan 2024 · 1 repository · arXiv:2401.10222Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 12 harvested samples)
-
ReFT: Reasoning with Reinforced Fine-Tuning 17 Jan 2024 · 1 repository · arXiv:2401.08967Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 9 unverified (of 14 harvested samples) · 13 pointer-only (licence)
-
Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation 16 Jan 2024 · 1 repository · arXiv:2401.08417Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples)
-
Extending LLMs' Context Window with 100 Samples 13 Jan 2024 · 1 repository · arXiv:2401.07004Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models 2 Jan 2024 · 2 repositories · arXiv:2401.01335Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning 25 Dec 2023 · 1 repository · arXiv:2312.15685Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 1 pointer-only (licence)
-
On Task Performance and Model Calibration with Supervised and Self-Ensembled In-Context Learning 21 Dec 2023 · 1 repository · arXiv:2312.13772Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
HuRef: HUman-REadable Fingerprint for Large Language Models 8 Dec 2023 · 1 repository · arXiv:2312.04828Syntology official (archive's flag): 3 ran · 3 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Beyond Imitation: Leveraging Fine-grained Quality Signals for Alignment 7 Nov 2023 · 1 repository · arXiv:2311.04072Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 0 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Tailoring Self-Rationalizers with Multi-Reward Distillation 6 Nov 2023 · 1 repository · arXiv:2311.02805Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch 6 Nov 2023 · 3 repositories · arXiv:2311.03099Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Vanishing Gradients in Reinforcement Finetuning of Language Models 31 Oct 2023 · 1 repository · arXiv:2310.20703Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Qilin-Med: Multi-stage Knowledge Injection Advanced Medical Large Language Model 13 Oct 2023 · 1 repository · arXiv:2310.09089Syntology official (archive's flag): 2 ran · 5 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
How Abilities in Large Language Models are Affected by Supervised Fine-tuning Data Composition 9 Oct 2023 · 2 repositories · arXiv:2310.05492Syntology community repositories only · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Reinforcement Learning in the Era of LLMs: What is Essential? What is needed? An RL Perspective on RLHF, Prompting, and Beyond 9 Oct 2023 · 2 repositories · arXiv:2310.06147Syntology 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Chat Vector: A Simple Approach to Equip LLMs with Instruction Following and Model Alignment in New Languages 7 Oct 2023 · 2 repositories · arXiv:2310.04799Syntology 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 2 pointer-only (licence)
-
GPT-Fathom: Benchmarking Large Language Models to Decipher the Evolutionary Path towards GPT-4 and Beyond 28 Sep 2023 · 1 repository · arXiv:2309.16583Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 9 harvested samples)
-
Okapi: Instruction-tuned Large Language Models in Multiple Languages with Reinforcement Learning from Human Feedback 29 Jul 2023 · 2 repositories · arXiv:2307.16039Syntology official: harvested, nothing ran · 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 1 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 11 harvested samples) · 7 pointer-only (licence)