Browse State-of-the-Art › Prompt Engineering
Prompt Engineering
454 papers with code · 16 benchmarks · 16 datasets archive 2025-07-28
Prompt engineering is the process of designing and refining the prompts used to generate text from language models, such as GPT-3 or similar models. The goal of prompt engineering is to improve the quality and relevance of the generated text by carefully crafting the prompts to elicit the desired responses from the model.
Prompt engineering involves several steps, including selecting the appropriate model architecture and parameters, designing the prompt format and structure, selecting the appropriate task and training data, and fine-tuning the model using the selected prompt and data.
Prompt engineering is a crucial step in the development of language models, as it can greatly influence the quality and effectiveness of the model's responses. By carefully designing and refining the prompts used to generate text, researchers and developers can improve the accuracy and relevance of the model's output, making it more useful for a wide range of applications, including chatbots, language translation, content creation, and more.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
16 leaderboard tables shown for this task, 16 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 16 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
16 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 454 papers with code (1,236 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
26 Feb 2021 82 repositories listed Syntology ran 16 of 20 samples · 4 unverified · 16 pointer-only (licence)State-of-the-art computer vision systems are trained to predict a fixed set of predetermined object categories.
-
2 Sep 2021 18 repositories listed Syntology ran 3 of 12 samples · 9 unverifiedLarge pre-trained vision-language models like CLIP have shown great potential in learning representations that are transferable across a wide range of downstream tasks.
-
10 Mar 2022 12 repositories listed Syntology ran 4 of 6 samples · 2 unverifiedWith the rise of powerful pre-trained vision-language models like CLIP, it becomes essential to investigate ways to adapt these models to downstream datasets.
-
18 Mar 2021 10 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 2 pointer-only (licence)Prompting a pretrained language model with natural language patterns has been proved effective for natural language understanding (NLU).
-
15 Oct 2021 8 repositories listed Syntology ran 8 of 15 samples · 7 unverifiedLarge language models have recently been shown to attain reasonable zero-shot generalization on a diverse set of tasks (Brown et al., 2020).
-
23 Mar 2022 6 repositories listed Syntology ran 17 of 27 samples · 10 unverified · 15 pointer-only (licence)The current modus operandi in adapting pre-trained models involves updating all the backbone parameters, ie, full fine-tuning.
-
29 Dec 2022 5 repositories listedNearly all jurisdictions in the United States require a professional license exam, commonly referred to as "the Bar Exam," as a precondition for law practice.
-
3 Nov 2022 5 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedBy conditioning on natural language instructions, large language models (LLMs) have displayed impressive capabilities as general-purpose computers.
-
26 Feb 2024 4 repositories listedExperiments illustrate that LangGPT significantly enhances the performance of LLMs.
-
13 Aug 2023 4 repositories listed Syntology ran 3 of 8 samples · 5 unverifiedDespite the simplicity of our method, an IP-Adapter with only 22M parameters can achieve comparable or even better performance to a fully fine-tuned image prompt model.
-
23 Jun 2023 4 repositories listedMultimodal Large Language Model (MLLM) relies on the powerful LLM to perform multimodal tasks, showing amazing emergent abilities in recent studies, such as writing poems based on an image.
-
30 Aug 2021 4 repositories listed Syntology ran 3 of 4 samples · 1 unverifiedLarge-scale pre-trained language models have contributed significantly to natural language processing by demonstrating remarkable abilities as few-shot learners.
-
6 Jun 2024 3 repositories listed Syntology ran 3 of 5 samples · 2 unverifiedGenerative Artificial Intelligence (GenAI) systems are increasingly being deployed across diverse industries and research domains.
-
15 Feb 2024 3 repositories listedOur findings reveal that, intriguingly, CoT reasoning paths can be elicited from pre-trained LLMs by simply altering the \textit{decoding} process.
-
19 Dec 2022 3 repositories listed Syntology ran 1 of 9 samples · 8 unverifiedInstead of laborious human engineering, we propose prompt adaptation, a general framework that automatically adapts original user input to model-preferred prompts.
-
6 Oct 2022 3 repositories listed Syntology ran 4 of 6 samples · 2 unverifiedPre-trained vision-language (V-L) models such as CLIP have shown excellent generalization ability to downstream tasks.
-
5 Oct 2022 3 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedPrompting is a brittle process wherein small modifications to the prompt can cause large variations in the model predictions, and therefore significant effort is dedicated towards designing a painstakingly "perfect…
-
9 Oct 2021 3 repositories listed Syntology ran 0 of 3 samples · 3 unverifiedLarge-scale contrastive vision-language pre-training has shown significant progress in visual representation learning.
-
12 Jun 2025 2 repositories listedParticularly, we design a fine-grained multi-objective reward specifically for time series forecasting, and then introduce GRIP (group-based relative importance for policy optimization), which leverages non-uniform…
-
4 Feb 2025 2 repositories listedEnsuring the safety of autonomous vehicles requires virtual scenario-based testing, which depends on the robust evaluation and generation of safety-critical scenarios.
-
30 Oct 2024 2 repositories listedRecent advances in LLM have been instrumental in autonomous robot control and human-robot interaction by leveraging their vast general knowledge and capabilities to understand and reason across a wide range of tasks and…
-
27 Aug 2024 2 repositories listedSpecifically, we introduce a relationship-guided attention module to capture pair-wise associations among entities and attributes for low-level prompt learning.
-
12 Jul 2024 2 repositories listed Syntology ran 8 of 17 samples · 9 unverifiedThe LAPT framework operates autonomously, requiring only ID class names as input and eliminating the need for manual intervention.
-
28 Jun 2024 2 repositories listedIn particular, changes in morphology and lexicon, i.
-
12 Jun 2024 2 repositories listedLarge audio-language models (LALMs) enhance traditional large language models by integrating audio perception capabilities, allowing them to tackle audio-related tasks.
-
8 Jun 2024 2 repositories listedTo address these limitations, we introduce RobustAlpacaEval, a new benchmark that consists of semantically equivalent case-level queries and emphasizes the importance of using the worst prompt performance to gauge the…
-
24 May 2024 2 repositories listedThis study evaluates the OpenAPI completion performance of GitHub Copilot, a prevalent commercial code completion tool, and proposes a set of task-specific optimizations leveraging Meta's open-source model Code Llama.
-
9 May 2024 2 repositories listedIncomplete relevance judgments limit the re-usability of test collections.
-
24 Apr 2024 2 repositories listed Syntology ran 0 of 3 samples · 3 unverifiedLarge language models (LLMs) are rapidly emerging in Artificial Intelligence (AI) applications, especially in the fields of natural language processing and generative AI.
-
23 Apr 2024 2 repositories listed(3) Corrective learning.
Syntology lines on 16 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections