Papers › PS4: a Next-Generation Dataset for Protein Single Sequence Secondary Structure Prediction

PS4: a Next-Generation Dataset for Protein Single Sequence Secondary Structure Prediction

1 Mar 2023biorXiv Preprint 2023 3archive 2025-07-28

Omar Peracha

Protein secondary structure prediction is a subproblem of protein folding. A lightweight algorithm capable of accurately predicting secondary structure from only the protein residue sequence could therefore provide a useful input for tertiary structure prediction, alleviating the reliance on MSA typically seen in today's best-performing models. This in turn could see the development of protein folding algorithms which perform better on orphan proteins, and which are much more accessible for both research and industry adoption due to reducing the necessary computational resources to run. Unfortunately, existing datasets for secondary structure prediction are small, creating a bottleneck in the rate of progress of automatic secondary structure prediction. Furthermore, protein chains in these datasets are often not identified, hampering the ability of researchers to use external domain knowledge when developing new algorithms. We present PS4, a dataset of 18,731 non-redundant protein chains and their respective Q8 secondary structure labels. Each chain is identified by its PDB code, and the dataset is also non-redundant against other secondary structure datasets commonly seen in the literature. We perform ablation studies by training secondary structure prediction algorithms on the PS4 training set, and obtain state-of-the-art Q8 and Q3 accuracy on the CB513 test set in zero shots, without further fine-tuning. Furthermore, we provide a software toolkit for the community to run our evaluation algorithms, train models from scratch and add new samples to the dataset. All code and data required to reproduce our results and make new inferences is available at https://github.com/omarperacha/ps4-dataset

PaperPDFCode

Code

omarperacha/ps4-dataset officialmentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

PredictionProtein FoldingProtein Secondary Structure Prediction

Datasets

Introduced by this paper, per the archive.

PS4

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Protein Secondary Structure Prediction CB513 PS4-Mega Q3 0.868 #1 of 10 Archive leaderboard report
Protein Secondary Structure Prediction CB513 PS4-Mega Q8 0.763 #1 of 10 Archive leaderboard report
Protein Secondary Structure Prediction CB513 PS4-Conv Q3 0.863 #2 of 10 Archive leaderboard report
Protein Secondary Structure Prediction CB513 PS4-Conv Q8 0.756 #2 of 10 Archive leaderboard report
Protein Secondary Structure Prediction PS4 PS4-Mega Q8 0.782 #1 of 2 Archive leaderboard report
Protein Secondary Structure Prediction PS4 PS4-Conv Q8 0.779 #2 of 2 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Test

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections