Browse State-of-the-Art › Instruction Following
Instruction Following
609 papers with code · 1 benchmark · 14 datasets archive 2025-07-28
Instruction following is the basic task of the model. This task is dedicated to evaluating the ability of the large model to follow human instructions. It is hoped that the model can generate controllable and safe answers.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| IFEval (4 rows) | AutoIF (Llama3 70B) | Self-play with Execution Feedback: Improving Instruction-following... | code | Syntology ran 1 of 1 samples · 0 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
14 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 609 papers with code (1,135 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
23 May 2023 20 repositories listed Syntology ran 17 of 26 samples · 9 unverified · 17 pointer-only (licence)Our best model family, which we name Guanaco, outperforms all previous openly released models on the Vicuna benchmark, reaching 99.
-
20 Dec 2022 19 repositories listed Syntology ran 7 of 17 samples · 10 unverified · 1 pointer-only (licence)Applying our method to the vanilla GPT3, we demonstrate a 33% absolute improvement over the original model on Super-NaturalInstructions, on par with the performance of InstructGPT-001, which was trained with private…
-
17 Apr 2023 13 repositories listed Syntology ran 16 of 51 samples · 35 unverifiedInstruction tuning large language models (LLMs) using machine-generated instruction-following data has improved zero-shot capabilities on new tasks, but the idea is less explored in the multimodal field.
-
2 Apr 2019 13 repositories listed Syntology ran 3 of 15 samples · 12 unverified · 15 pointer-only (licence)We present Habitat, a platform for research in embodied artificial intelligence (AI).
-
16 Apr 2022 10 repositories listed Syntology ran 7 of 28 samples · 21 unverified · 4 pointer-only (licence)This large and diverse collection of tasks enables rigorous benchmarking of cross-task generalization under instructions -- training models to follow instructions on a subset of tasks and evaluating them on the…
-
18 Jun 2024 7 repositories listed Syntology ran 15 of 29 samples · 14 unverifiedWe introduce ChatGLM, an evolving family of large language models that we have been developing over time.
-
28 Mar 2023 7 repositories listedWe present LLaMA-Adapter, a lightweight adaption method to efficiently fine-tune LLaMA into an instruction-following model.
-
19 Dec 2024 6 repositories listed Syntology ran 1 of 3 samples · 2 unverifiedIn addition, for hosted solutions, the proprietary models currently include two mixture-of-experts (MoE) variants: Qwen2.
-
6 Feb 2024 5 repositories listed Syntology ran 5 of 5 samples · 0 unverified · 4 pointer-only (licence)Knowledge distillation (KD) is widely used for compressing a teacher model to a smaller student model, reducing its inference cost and memory footprint while preserving model capabilities.
-
21 Sep 2023 5 repositories listedStudying how people interact with large language models (LLMs) in real-world scenarios is increasingly important due to their widespread use in various applications.
-
1 Sep 2023 5 repositories listed Syntology ran 13 of 20 samples · 7 unverified · 7 pointer-only (licence)We introduce Point-Bind, a 3D multi-modality model aligning point clouds with 2D image, language, audio, and video.
-
4 Sep 2018 5 repositories listedWe propose to decompose instruction execution to goal prediction and action generation.
-
14 Nov 2023 4 repositories listed Syntology ran 2 of 7 samples · 5 unverifiedOne core capability of Large Language Models (LLMs) is to follow natural language instructions.
-
21 Sep 2023 4 repositories listed Syntology ran 11 of 13 samples · 2 unverifiedFor example, training on the context length of 8192 needs 16x computational costs in self-attention layers as that of 2048.
-
7 Jun 2023 4 repositories listed Syntology ran 6 of 12 samples · 6 unverified · 2 pointer-only (licence)Our evaluations show that the best model in any given evaluation reaches on average 87% of ChatGPT performance, and 73% of GPT-4 performance, suggesting that further investment in building better base models and…
-
24 Apr 2023 4 repositories listed Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)In this paper, we show an avenue for creating large amounts of instruction data with varying levels of complexity using LLM instead of humans.
-
15 Aug 2024 3 repositories listedIn this work, we propose a new framework for the knowledge fusion of chat LLMs through two main stages, resulting in FuseChat.
-
4 Jul 2024 3 repositories listed Syntology ran 10 of 10 samples · 0 unverifiedThis report introduces FunAudioLLM, a model family designed to enhance natural voice interactions between humans and large language models (LLMs).
-
22 Jun 2024 3 repositories listedDespite the remarkable advancement of Large language models (LLMs), they still lack delicate controllability under sophisticated constraints, which is critical to enhancing their response quality and the user experience.
-
22 May 2024 3 repositories listedFinetuning large language models with a variety of instruction-response pairs has enhanced their capability to understand and follow instructions.
-
27 Feb 2024 3 repositories listed Syntology ran 9 of 17 samples · 8 unverifiedThis paper presents ShapeLLM, the first 3D Multimodal Large Language Model (LLM) designed for embodied interaction, exploring a universal 3D object understanding with 3D point clouds and languages.
-
19 Feb 2024 3 repositories listed Syntology ran 9 of 16 samples · 7 unverified · 7 pointer-only (licence)We demonstrate the effectiveness of RESTA in both parameter-efficient and full fine-tuning, covering a wide range of downstream tasks, including instruction following in Chinese, English, and Hindi, as well as…
-
10 Feb 2024 3 repositories listedTrained on massive publicly available data, large language models (LLMs) have demonstrated tremendous success across various fields.
-
18 Jan 2024 3 repositories listed Syntology ran 3 of 11 samples · 8 unverifiedWe posit that to achieve superhuman agents, future models require superhuman feedback in order to provide an adequate training signal.
-
13 Nov 2023 3 repositories listed Syntology ran 4 of 6 samples · 2 unverified · 1 pointer-only (licence)To mitigate the potential misuse of large language models (LLMs), recent research has developed watermarking algorithms, which restrict the generation process to leave an invisible trace for watermark detection.
-
6 Nov 2023 3 repositories listed Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)We experiment with encoder- and decoder-based LMs, showing that: (1) SFT delta parameter value ranges are typically small (within 0.
-
28 Sep 2023 3 repositories listed Syntology ran 5 of 9 samples · 4 unverified · 8 pointer-only (licence)We propose a memory-efficient finetuning algorithm for large language models (LLMs) that supports finetuning LLMs with 65B parameters in 2/3/4-bit precision on as little as one 24GB GPU.
-
23 Aug 2023 3 repositories listed Syntology ran 1 of 2 samples · 1 unverifiedTo achieve this, we first propose several metrics to access the quality of multimodal instruction data.
-
23 Aug 2023 3 repositories listed Syntology ran 7 of 10 samples · 3 unverified · 10 pointer-only (licence)In the realm of Large Language Models (LLMs), the balance between instruction data quality and quantity is a focal point.
-
20 Jul 2023 3 repositories listed Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)Recently, there has been growing interest in extending the context length of large language models (LLMs), aiming to effectively process long inputs of one turn or conversations with more extensive histories.
Syntology lines on 23 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections