{"url":"/dataset/flip-aav-designed-vs-mutant","name":"FLIP -- AAV, Designed vs mutant","full_name":"adeno-associated virus","description_markdown":"FLIP includes several benchmark datasets that contain a variety of protein sequences, each with a real-valued label indicating its \"fitness\" (how well the protein performs some particular function). The goal is to predict the fitness of a given protein sequence using the sequence. Different representations of protein sequences (e.g. learned embeddings from large language models) may prove helpful here.\r\n\r\nThis sub-dataset (AAV) is a set of 201,426 training sequences and 82,583 test sequences in which the goal is to predict the fitness of mutants of the capsid protein from the adeno-associated virus (AAV). The training set proteins were designed, while the test set proteins are random mutants. The absolute value of the fitness is not important, but its ranking / relative value is -- protein designers would like to be able to pick a sequence with high fitness relative to those in the training set. Performance is therefore usually assessed using Spearman's r correlation coefficient.","description_withheld":null,"homepage":"https://benchmark.protein.properties/","introduced_date":"2022-01-19","introduced_date_note":null,"introduced_by":{"paper":"/paper/flip-benchmark-tasks-in-fitness-landscape","title":"FLIP: Benchmark tasks in fitness landscape inference for proteins","first_author":"Christian Dallago","url":null},"license":{"name":"Academic Free License v3.0","url":"https://github.com/J-SNACKKB/FLIP/blob/main/LICENSE"},"modalities":[{"name":"Biology","url":"/datasets/modality/biology"}],"tasks":[{"name":"Protein Function Prediction","url":"/task/protein-function-prediction","datasets_with_task":"/datasets/task/protein-function-prediction"},{"name":"regression","url":"/task/regression-1","datasets_with_task":"/datasets/task/regression-1"},{"name":"Protein Design","url":"/task/protein-design","datasets_with_task":"/datasets/task/protein-design"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["FLIP -- AAV, Designed vs mutant"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}