Methods › General › Distillation › SFT › Papers, page 4
Shrink and Fine-Tune
SFT
Papers archive 2025-07-28
archive papers tagged: 415 · with a code link: 204 · where Syntology ran a sample: 103 (86 with a run with no instrument failure, 17 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (103 of 415 tagged: 86 with a run with no instrument failure, 17 where every run was a failure of Syntology's instrument)
Page 4 of 5: papers 301 to 400 of 415, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
A Fine-tuning Dataset and Benchmark for Large Language Models for Protein Understanding 8 Jun 2024 · 1 repository · arXiv:2406.05540Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
Extroversion or Introversion? Controlling The Personality of Your Large Language Models 7 Jun 2024 · 1 repository · arXiv:2406.04583
-
Proofread: Fixes All Errors with One Tap 6 Jun 2024 · 0 repositories · arXiv:2406.04523
-
Parrot: Multilingual Visual Instruction Tuning 4 Jun 2024 · 2 repositories · arXiv:2406.02539Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 4 unverified (of 7 harvested samples)
-
Strengthened Symbol Binding Makes Large Language Models Reliable Multiple-Choice Selectors 3 Jun 2024 · 1 repository · arXiv:2406.01026
-
LongSkywork: A Training Recipe for Efficiently Extending Context Length in Large Language Models 2 Jun 2024 · 0 repositories · arXiv:2406.00605
-
A Practice-Friendly LLM-Enhanced Paradigm with Preference Parsing for Sequential Recommendation 1 Jun 2024 · 0 repositories · arXiv:2406.00333
-
InstructionCP: A fast approach to transfer Large Language Models into target language 30 May 2024 · 0 repositories · arXiv:2405.20175
-
PediatricsGPT: Large Language Models as Chinese Medical Assistants for Pediatric Applications 29 May 2024 · 1 repository · arXiv:2405.19266Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 1 pointer-only (licence)
-
Getting More Juice Out of the SFT Data: Reward Learning from Human Demonstration Improves SFT for LLM Alignment 28 May 2024 · 1 repository · arXiv:2405.17888Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (of 8 harvested samples) · 1 pointer-only (licence)
-
Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment 28 May 2024 · 1 repository · arXiv:2405.17931
-
LeDex: Training LLMs to Better Self-Debug and Explain Code 28 May 2024 · 0 repositories · arXiv:2405.18649
-
Bayesian RG Flow in Neural Network Field Theories 27 May 2024 · 1 repository · arXiv:2405.17538
-
Cross-Modal Safety Alignment: Is textual unlearning all you need? 27 May 2024 · 0 repositories · arXiv:2406.02575
-
Exploring the LLM Journey from Cognition to Expression with Linear Representations 27 May 2024 · 0 repositories · arXiv:2405.16964
-
Automatically Generating Numerous Context-Driven SFT Data for LLMs across Diverse Granularity 26 May 2024 · 1 repository · arXiv:2405.16579
-
Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer 26 May 2024 · 0 repositories · arXiv:2405.16436
-
Triple Preference Optimization: Achieving Better Alignment with Less Data in a Single Step Optimization 26 May 2024 · 1 repository · arXiv:2405.16681Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples)
-
Multilingual Prosody Transfer: Comparing Supervised & Transfer Learning 23 May 2024 · 0 repositories · arXiv:2406.00022
-
360Zhinao Technical Report 22 May 2024 · 1 repository · arXiv:2405.13386
-
Disperse-Then-Merge: Pushing the Limits of Instruction Tuning via Alignment Tax Reduction 22 May 2024 · 1 repository · arXiv:2405.13432Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Intuitive Fine-Tuning: Towards Simplifying Alignment into a Single Process 20 May 2024 · 1 repository · arXiv:2405.11870
-
Rethinking Overlooked Aspects in Vision-Language Models 20 May 2024 · 0 repositories · arXiv:2405.11850
-
Advanced Natural-based interaction for the ITAlian language: LLaMAntino-3-ANITA 11 May 2024 · 0 repositories · arXiv:2405.07101
-
A Lightweight Sparse Focus Transformer for Remote Sensing Image Change Captioning 10 May 2024 · 1 repository · arXiv:2405.06598
-
Optimizing Language Model's Reasoning Abilities with Weak Supervision 7 May 2024 · 0 repositories · arXiv:2405.04086
-
Labeling supervised fine-tuning data with the scaling law 5 May 2024 · 2 repositories · arXiv:2405.02817
-
FLAME: Factuality-Aware Alignment for Large Language Models 2 May 2024 · 0 repositories · arXiv:2405.01525
-
Prefix Text as a Yarn: Eliciting Non-English Alignment in Foundation Language Model 25 Apr 2024 · 0 repositories · arXiv:2404.16766
-
Weak-to-Strong Extrapolation Expedites Alignment 25 Apr 2024 · 1 repository · arXiv:2404.16792
-
Insights into Alignment: Evaluating DPO and its Variants Across Multiple Tasks 23 Apr 2024 · 1 repository · arXiv:2404.14723
-
A Preference-driven Paradigm for Enhanced Translation with Large Language Models 17 Apr 2024 · 0 repositories · arXiv:2404.11288
-
Balancing Speciality and Versatility: a Coarse to Fine Framework for Supervised Fine-tuning Large Language Model 16 Apr 2024 · 1 repository · arXiv:2404.10306Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Learn Your Reference Model for Real Good Alignment 15 Apr 2024 · 0 repositories · arXiv:2404.09656
-
Chinese Tiny LLM: Pretraining a Chinese-Centric Large Language Model 5 Apr 2024 · 0 repositories · arXiv:2404.04167
-
Token-Efficient Leverage Learning in Large Language Models 1 Apr 2024 · 1 repository · arXiv:2404.00914
-
Extensive Self-Contrast Enables Feedback-Free Language Model Alignment 31 Mar 2024 · 2 repositories · arXiv:2404.00604
-
Dialectical Alignment: Resolving the Tension of 3H and Security Threats of LLMs 30 Mar 2024 · 0 repositories · arXiv:2404.00486
-
Injecting New Knowledge into Large Language Models via Supervised Fine-Tuning 30 Mar 2024 · 0 repositories · arXiv:2404.00213
-
Detoxifying Large Language Models via Knowledge Editing 21 Mar 2024 · 1 repository · arXiv:2403.14472Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 4 harvested samples)
-
A Novel Paradigm Boosting Translation Capabilities of Large Language Models 18 Mar 2024 · 0 repositories · arXiv:2403.11430
-
A Probabilistic Approach for Model Alignment with Human Comparisons 16 Mar 2024 · 0 repositories · arXiv:2403.10771
-
ORPO: Monolithic Preference Optimization without Reference Model 12 Mar 2024 · 4 repositories · arXiv:2403.07691Syntology official (archive's flag): 2 ran · 11 ran (of which 1 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 0 violated, 9 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 14 harvested samples) · 2 pointer-only (licence)
-
SmallToLarge (S2L): Scalable Data Selection for Fine-tuning Large Language Models by Summarizing Training Trajectories of Small Models 12 Mar 2024 · 1 repository · arXiv:2403.07384Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 3 pointer-only (licence)
-
Unfamiliar Finetuning Examples Control How Language Models Hallucinate 8 Mar 2024 · 1 repository · arXiv:2403.05612Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Common 7B Language Models Already Possess Strong Math Capabilities 7 Mar 2024 · 2 repositories · arXiv:2403.04706Syntology official (archive's flag): 9 ran · 21 ran (of which 0 constructed an object rather than computing a result; 20 with no instrument failure: 0 honoured, 0 violated, 20 with no contract checked; 1 where Syntology's instrument failed) · 7 unverified (of 28 harvested samples) · 12 pointer-only (licence)
-
Teaching Large Language Models to Reason with Reinforcement Learning 7 Mar 2024 · 0 repositories · arXiv:2403.04642
-
Balancing Enhancement, Harmlessness, and General Capabilities: Enhancing Conversational LLMs with Direct RLHF 4 Mar 2024 · 0 repositories · arXiv:2403.02513
-
DACO: Towards Application-Driven and Comprehensive Data Analysis via Code Generation 4 Mar 2024 · 1 repository · arXiv:2403.02528Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Provably Robust DPO: Aligning Language Models with Noisy Feedback 1 Mar 2024 · 0 repositories · arXiv:2403.00409
-
Reflect-RL: Two-Player Online RL Fine-Tuning for LMs 20 Feb 2024 · 1 repository · arXiv:2402.12621Syntology official (archive's flag): 2 ran · 2 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 6 harvested samples)
-
A Critical Evaluation of AI Feedback for Aligning Large Language Models 19 Feb 2024 · 1 repository · arXiv:2402.12366Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples) · 1 pointer-only (licence)
-
RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language Models 15 Feb 2024 · 0 repositories · arXiv:2402.10038
-
ICDPO: Effectively Borrowing Alignment Capability of Others via In-context Direct Preference Optimization 14 Feb 2024 · 1 repository · arXiv:2402.09320
-
EntGPT: Linking Generative Large Language Models with Knowledge Bases 9 Feb 2024 · 1 repository · arXiv:2402.06738
-
Rethinking Data Selection for Supervised Fine-Tuning 8 Feb 2024 · 0 repositories · arXiv:2402.06094
-
Pedagogical Alignment of Large Language Models 7 Feb 2024 · 1 repository · arXiv:2402.05000
-
UltraLink: An Open-Source Knowledge-Enhanced Multilingual Supervised Fine-tuning Dataset 7 Feb 2024 · 1 repository · arXiv:2402.04588Syntology official: no sample here; runs from other or unrecorded repositories · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
ANLS* -- A Universal Document Processing Metric for Generative Large Language Models 6 Feb 2024 · 2 repositories · arXiv:2402.03848
-
Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback 6 Feb 2024 · 1 repository · arXiv:2402.03746
-
The Political Preferences of LLMs 2 Feb 2024 · 0 repositories · arXiv:2402.01789
-
Scaling Sparse Fine-Tuning to Large Language Models 29 Jan 2024 · 2 repositories · arXiv:2401.16405Syntology official (archive's flag): 12 ran · 13 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 2 honoured, 0 violated, 7 with no contract checked; 4 where Syntology's instrument failed) · 2 unverified (of 15 harvested samples) · 5 pointer-only (licence)
-
YODA: Teacher-Student Progressive Learning for Language Models 28 Jan 2024 · 0 repositories · arXiv:2401.15670
-
Equipping Language Models with Tool Use Capability for Tabular Data Analysis in Finance 27 Jan 2024 · 0 repositories · arXiv:2401.15328
-
Supervised Fine-tuning in turn Improves Visual Foundation Models 18 Jan 2024 · 1 repository · arXiv:2401.10222Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 12 harvested samples)
-
ReFT: Reasoning with Reinforced Fine-Tuning 17 Jan 2024 · 1 repository · arXiv:2401.08967Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 9 unverified (of 14 harvested samples) · 13 pointer-only (licence)
-
Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation 16 Jan 2024 · 1 repository · arXiv:2401.08417Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples)
-
Extending LLMs' Context Window with 100 Samples 13 Jan 2024 · 1 repository · arXiv:2401.07004Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
An Experimental Design Framework for Label-Efficient Supervised Finetuning of Large Language Models 12 Jan 2024 · 0 repositories · arXiv:2401.06692
-
Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models 2 Jan 2024 · 2 repositories · arXiv:2401.01335Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
GeoGalactica: A Scientific Large Language Model in Geoscience 31 Dec 2023 · 1 repository · arXiv:2401.00434
-
LaFFi: Leveraging Hybrid Natural Language Feedback for Fine-tuning Language Models 31 Dec 2023 · 0 repositories · arXiv:2401.00907
-
Advancing TTP Analysis: Harnessing the Power of Large Language Models with Retrieval Augmented Generation 30 Dec 2023 · 1 repository · arXiv:2401.00280
-
What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning 25 Dec 2023 · 1 repository · arXiv:2312.15685Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 1 pointer-only (licence)
-
Carve3D: Improving Multi-view Reconstruction Consistency for Diffusion Models with RL Finetuning 21 Dec 2023 · 0 repositories · arXiv:2312.13980
-
On Task Performance and Model Calibration with Supervised and Self-Ensembled In-Context Learning 21 Dec 2023 · 1 repository · arXiv:2312.13772Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
HuRef: HUman-REadable Fingerprint for Large Language Models 8 Dec 2023 · 1 repository · arXiv:2312.04828Syntology official (archive's flag): 3 ran · 3 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context Learning 4 Dec 2023 · 1 repository · arXiv:2312.01552
-
A Two-Stage Adaptation of Large Language Models for Text Ranking 28 Nov 2023 · 1 repository · arXiv:2311.16720
-
ShareGPT4V: Improving Large Multi-Modal Models with Better Captions 21 Nov 2023 · 1 repository · arXiv:2311.12793
-
Examining Modularity in Multilingual LMs via Language-Specialized Subnetworks 14 Nov 2023 · 0 repositories · arXiv:2311.08273
-
ChiMed-GPT: A Chinese Medical Large Language Model with Full Training Regime and Better Alignment to Human Preferences 10 Nov 2023 · 1 repository · arXiv:2311.06025
-
Beyond Imitation: Leveraging Fine-grained Quality Signals for Alignment 7 Nov 2023 · 1 repository · arXiv:2311.04072Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 0 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch 6 Nov 2023 · 3 repositories · arXiv:2311.03099Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Tailoring Self-Rationalizers with Multi-Reward Distillation 6 Nov 2023 · 1 repository · arXiv:2311.02805Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Vanishing Gradients in Reinforcement Finetuning of Language Models 31 Oct 2023 · 1 repository · arXiv:2310.20703Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
SuperHF: Supervised Iterative Learning from Human Feedback 25 Oct 2023 · 1 repository · arXiv:2310.16763Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
DavIR: Data Selection via Implicit Reward for Large Language Models 16 Oct 2023 · 0 repositories · arXiv:2310.13008
-
Qilin-Med: Multi-stage Knowledge Injection Advanced Medical Large Language Model 13 Oct 2023 · 1 repository · arXiv:2310.09089Syntology official (archive's flag): 2 ran · 5 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Understanding the Effects of RLHF on LLM Generalisation and Diversity 10 Oct 2023 · 1 repository · arXiv:2310.06452Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
How Abilities in Large Language Models are Affected by Supervised Fine-tuning Data Composition 9 Oct 2023 · 2 repositories · arXiv:2310.05492Syntology community repositories only · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Reinforcement Learning in the Era of LLMs: What is Essential? What is needed? An RL Perspective on RLHF, Prompting, and Beyond 9 Oct 2023 · 2 repositories · arXiv:2310.06147Syntology 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Chat Vector: A Simple Approach to Equip LLMs with Instruction Following and Model Alignment in New Languages 7 Oct 2023 · 2 repositories · arXiv:2310.04799Syntology 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 2 pointer-only (licence)
-
CITING: Large Language Models Create Curriculum for Instruction Tuning 4 Oct 2023 · 0 repositories · arXiv:2310.02527
-
Automatic Pair Construction for Contrastive Post-training 3 Oct 2023 · 1 repository · arXiv:2310.02263
-
GPT-Fathom: Benchmarking Large Language Models to Decipher the Evolutionary Path towards GPT-4 and Beyond 28 Sep 2023 · 1 repository · arXiv:2309.16583Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 9 harvested samples)
-
RRHF: Rank Responses to Align Language Models with Human Feedback 21 Sep 2023 · 1 repository
-
Statistical Rejection Sampling Improves Preference Optimization 13 Sep 2023 · 0 repositories · arXiv:2309.06657
-
Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning 5 Sep 2023 · 1 repository · arXiv:2309.02591
-
Efficient RLHF: Reducing the Memory Usage of PPO 1 Sep 2023 · 0 repositories · arXiv:2309.00754