Browse State-of-the-Art › Protein Design
Protein Design
74 papers with code · 2 benchmarks · 6 datasets archive 2025-07-28
Formally, given the design requirements of users, models are required to generate protein amino acid sequences that align with those requirements.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
2 leaderboard tables shown for this task, 2 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| CATH 4.2 (8 rows) | Knowledge-Design | Knowledge-Design: Pushing the Limit of Protein Design via... | code | — | Compare |
| CATH 4.3 (2 rows) | GVP-large | Knowledge-Design: Pushing the Limit of Protein Design via... | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
6 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 74 papers with code (175 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
11 Feb 2024 4 repositories listedStarting with a set of pre-trained LoRA adapters, our gating strategy uses the hidden states to dynamically mix adapted layers, allowing the resulting X-LoRA model to draw upon different capabilities and create…
-
27 Jun 2022 4 repositories listed Syntology ran 8 of 18 samples · 10 unverified · 2 pointer-only (licence)Attention-based models trained on protein sequences have demonstrated incredible success at classification and generation tasks relevant for artificial intelligence-driven protein design.
-
11 May 2022 4 repositories listed Syntology ran 1 of 2 samples · 1 unverifiedIn this work we introduce RITA: a suite of autoregressive generative models for protein sequences, with up to 1.
-
27 Feb 2024 3 repositories listedIn this work, we propose TaxDiff, a taxonomic-guided diffusion model for controllable protein sequence generation that combines biological species information with the generative capabilities of diffusion models to…
-
9 Feb 2023 3 repositories listed Syntology ran 2 of 3 samples · 1 unverified · 1 pointer-only (licence)Current AI-assisted protein design mainly utilizes protein sequential and structural information.
-
8 Feb 2023 3 repositories listed Syntology ran 3 of 10 samples · 7 unverifiedImportantly, we demonstrate that the geometry-complete denoising process of GCDM learned for 3D molecule generation enables the model to generate a significant proportion of valid and energetically-stable large…
-
3 Sep 2020 3 repositories listed Syntology ran 5 of 6 samples · 1 unverified · 6 pointer-only (licence)Learning on 3D structures of large biomolecules is emerging as a distinct area in machine learning, but there has yet to emerge a unifying network architecture that simultaneously leverages the graph-structured and…
-
8 Jan 2024 2 repositories listed Syntology ran 7 of 9 samples · 2 unverifiedProtein design often begins with the knowledge of a desired function from a motif which motif-scaffolding aims to construct a functional protein around.
-
8 Oct 2023 2 repositories listedWe present FrameFlow, a method for fast protein backbone generation using SE(3) flow matching.
-
30 Jun 2023 2 repositories listed Syntology ran 9 of 15 samples · 6 unverified · 15 pointer-only (licence)Diffusion models have been successful on a range of conditional generation tasks including molecular design and text-to-image generation.
-
20 Aug 2022 2 repositories listed Syntology ran 2 of 12 samples · 10 unverified · 2 pointer-only (licence)Data-driven predictive methods which can efficiently and accurately transform protein sequences into biologically active structures are highly valuable for scientific research and medical development.
-
9 Dec 2017 2 repositories listedHere we present an embedding of natural protein sequences using a Variational Auto-Encoder and use it to predict how mutations affect protein function.
-
9 Jun 2025 1 repository listedLarge language models (LLMs) have become the cornerstone of modern AI.
-
9 Jun 2025 1 repository listedWe also introduce DSM(ppi), a variant fine-tuned to generate protein binders by attending to target sequences.
-
28 May 2025 1 repository listedExisting PLMs generate protein sequences based on a single-condition constraint from a specific modality, struggling to simultaneously satisfy multiple constraints across different modalities.
-
13 May 2025 1 repository listedWhile AI methods like AlphaFold can predict accurate structural models for many protein complexes, reliably estimating the quality of these predicted models (estimation of model accuracy, or EMA) for model ranking and…
-
2 May 2025 1 repository listed Syntology ran 0 of 3 samples · 3 unverifiedAdvances in protein engineering have the potential to revolutionize biotechnology and healthcare by designing proteins with tailored properties.
-
28 Mar 2025 1 repository listedHowever, existing methods in MOQD rely on tessellating the feature space into a grid structure, which prevents their application in domains where feature spaces are unknown or must be learned, such as complex biological…
-
2 Mar 2025 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Here, we develop Proteina, a new large-scale flow-based protein backbone generator that utilizes hierarchical fold class labels for conditioning and relies on a tailored scalable transformer architecture with up to 5x…
-
20 Feb 2025 1 repository listedProtein backbone generation plays a central role in de novo protein design and is significant for many biological and medical applications.
-
18 Feb 2025 1 repository listedThe motif-scaffolding problem is a central task in computational protein design: Given the coordinates of atoms in a geometry chosen to confer a desired biochemical function (a motif), the task is to identify diverse…
-
12 Feb 2025 1 repository listedProtein flexibility, measured by the B-factor or Debye-Waller factor, is essential for protein functions such as structural support, enzyme activity, cellular communication, and molecular transport.
-
25 Jan 2025 1 repository listedDesigning proteins with specific attributes offers an important solution to address biomedical challenges.
-
12 Jan 2025 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)For steering text-to-image models with a human preference reward, we find that FK steering a 0.
-
27 Nov 2024 1 repository listed Syntology ran 0 of 5 samples · 5 unverifiedDeep generative models show promise for de novo protein design, but their effectiveness within specific protein families remains underexplored.
-
17 Nov 2024 1 repository listedThe impressive performance of large language models (LLMs) has led to their consideration as models of human language processing.
-
9 Nov 2024 1 repository listed Syntology ran 0 of 1 samples · 1 unverifiedWe introduce Concept Bottleneck Protein Language Models (CB-pLM), a generative masked language model with a layer where each neuron corresponds to an interpretable concept.
-
4 Nov 2024 1 repository listed Syntology ran 0 of 3 samples · 3 unverifiedTo fill these gaps, we propose Bridge-IF, a generative diffusion bridge model for inverse folding, which is designed to learn the probabilistic dependency between the distributions of backbone structures and protein…
-
25 Oct 2024 1 repository listedIn this work, we introduce PeptideGPT, a protein language model tailored to generate protein sequences with distinct properties: hemolytic activity, solubility, and non-fouling characteristics.
-
22 Oct 2024 1 repository listedTo address this, we present RL-DIF, a categorical diffusion model for inverse folding that is pre-trained on sequence recovery and tuned via reinforcement learning on structural consistency.
Syntology lines on 14 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections