Browse State-of-the-Art › Protein Language Model
Protein Language Model
47 papers with code · 1 benchmark · 5 datasets archive 2025-07-28
Use of protein language models in protein property prediction
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| DAVIS-DTA (1 row) | LEP-AD | LEP-AD: Language Embedding of Proteins and Attention to Drugs... | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
5 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 47 papers with code (73 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
5 Mar 2024 2 repositories listed Syntology ran 1 of 3 samples · 2 unverifiedIn this paper, we propose ESM-AA (ESM All-Atom), a novel approach that enables atom-scale and residue-scale unified molecular modeling.
-
1 Dec 2023 2 repositories listedMeanwhile, the ESM-NBR obtains the MCC values for DNA-binding residues prediction of 0.
-
21 Apr 2020 2 repositories listedMessenger RNA (mRNA) vaccines are being used for COVID-19, but still suffer from the critical issue of mRNA instability and degradation, which is a major obstacle in the storage, distribution, and efficacy of the…
-
21 Jun 2025 1 repository listedAccurate prediction of antibody-antigen (Ab-Ag) binding affinity is essential for therapeutic design and vaccine development, yet the performance of current models is limited by noisy experimental labels, heterogeneous…
-
22 May 2025 1 repository listedProtein language models (pLMs) pre-trained on vast protein sequence databases excel at various downstream tasks but lack the structural knowledge essential for many biological applications.
-
5 May 2025 1 repository listedOverall, this study assesses the state of the field for leveraging PLMs for sequence-to-k_(cat) prediction on a set of diverse ADK orthologs.
-
31 Dec 2024 1 repository listedAmong the generations that we synthesized, we found a bright fluorescent protein at a far distance (58% sequence identity) from known fluorescent proteins, which we estimate is equivalent to simulating five hundred…
-
24 Dec 2024 1 repository listed Syntology ran 9 of 11 samples · 2 unverifiedThird, we introduce a method for fine-tuning a protein inverse folding model to steer it toward desired flexibility in specified regions.
-
24 Dec 2024 1 repository listedAdditionally, to explore the applicability and generalization ability of the models, we constructed a short peptide data set and an external data set to test the retrained models.
-
2 Dec 2024 1 repository listed Syntology ran 1 of 9 samples · 8 unverifiedDesigning novel functional proteins crucially depends on accurately modeling their fitness landscape.
-
22 Nov 2024 1 repository listedResults To overcome the sparsity of experimental training data, we created a dataset of 39 million single amino acid variants from the human proteome applying the multiple sequence alignment-based effect predictor GEMME…
-
29 Oct 2024 1 repository listedSelf-supervised training of language models (LMs) has seen great success for protein sequences in learning meaningful representations and for generative drug design.
-
28 Oct 2024 1 repository listed Syntology ran 1 of 5 samples · 4 unverified · 5 pointer-only (licence)Enzyme engineering enables the modification of wild-type proteins to meet industrial and research demands by enhancing catalytic activity, stability, binding affinities, and other properties.
-
25 Oct 2024 1 repository listedIn this work, we introduce PeptideGPT, a protein language model tailored to generate protein sequences with distinct properties: hemolytic activity, solubility, and non-fouling characteristics.
-
17 Oct 2024 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)In this paper, we introduce DPLM-2, a multimodal protein foundation model that extends discrete diffusion protein language model (DPLM) to accommodate both sequences and structures.
-
Fine-tuning the ESM2 protein language model to understand the functional impact of missense variants14 Oct 2024 1 repository listedElucidating the functional effect of missense variants is of crucial importance, yet challenging.
-
24 Sep 2024 1 repository listedMoreover, the potential of raw sequence and protein structure has not been fully investigated.
-
24 Jun 2024 1 repository listedThese findings demonstrate the potential of tcrLM in predicting TCR-antigen binding specificity, with significant implications for advancing immunotherapy and personalized medicine.
-
29 May 2024 1 repository listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)Proteins are complex molecules responsible for different functions in nature.
-
6 May 2024 1 repository listedAntiFold outperforms existing inverse folding tools on sequence recovery across complementarity-determining regions, with designed sequences showing high structural similarity to their solved counterpart.
-
23 Apr 2024 1 repository listedProtein circular permutations are crucial for understanding protein evolution and functionality.
-
15 Apr 2024 1 repository listedThis study introduces a quantitative definition and benchmarking framework AMPCliff for the AC phenomenon in antimicrobial peptides (AMPs) composed by canonical amino acids.
-
15 Mar 2024 1 repository listedWe demonstrate that ESM2 is dominated by three regions in the sparsity-ruggedness plane, two of which are better suited for sparse Fourier transforms.
-
28 Feb 2024 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedThis paper introduces diffusion protein language model (DPLM), a versatile protein language model that demonstrates strong generative and predictive capabilities for protein sequences.
-
8 Feb 2024 1 repository listedA dominant paradigm is to train a model to jointly generate the antibody sequence and the structure as a candidate.
-
7 Feb 2024 1 repository listedTo address this issue, we introduce the integration of remote homology detection to distill structural information into protein language models without requiring explicit protein structures as input.
-
19 Dec 2023 1 repository listedAccurate identification of protein nucleic-acid-binding residues poses a significant challenge with important implications for various biological processes and drug design.
-
14 Dec 2023 1 repository listedSignal peptide (SP) is a short peptide located in the N-terminus of proteins.
-
26 Oct 2023 1 repository listedMoreover, despite the wealth of benchmarks and studies in the natural language community, there remains a lack of a comprehensive benchmark for systematically evaluating protein language model quality.
-
11 Oct 2023 1 repository listedPhonological reconstruction is one of the central problems in historical linguistics where a proto-word of an ancestral language is determined from the observed cognate words of daughter languages.
Syntology lines on 7 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections