{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2605-13789","title":"ENSEMBITS: an alphabet of protein conformational ensembles","arxiv_id":"2605.13789","date":"2026-05-13","proceeding":null,"authors":["Kaiwen Shi","Carlos Oliver"],"abstract":"Protein structure tokenizers (PSTs) are workhorses in protein language modeling, function prediction, and evolutionary analysis. However, existing PSTs only capture local geometry of static structures, and miss the correlated motions and alternative conformational states revealed by protein ensembles. Here we introduce Ensembits, the first tokenizer of protein conformational ensembles. Ensembits address challenges inherent to tokenizing dynamics: deriving informative geometric descriptors across conformations, permutation-invariance encoding of variable-size ensembles, and conquering sparsity in dynamics data. Trained with a Residual VQ-VAE using a frame distillation objective on a large molecular dynamics corpus, Ensembits outperforms all related methods on RMSF prediction, and is the strongest standalone structural tokenizer on an token-conditioned ANOVA test on per-residue motion amplitude. Ensembits further matches or exceeds static tokenizers on EC, GO, binding site/affinity prediction, and zero-shot mutation-effect prediction despite using far less pretraining data. Notably, the distillation objective enables Ensembits to predict dynamics token from one single predicted structure, which alleviates dynamics data sparsity. As the field moves from static structure prediction toward ensemble generation, Ensembits offer the discrete vocabulary needed to bring dynamics into protein language modeling and design.","url_abs":"https://arxiv.org/abs/2605.13789","url_pdf":"https://arxiv.org/pdf/2605.13789","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2605.13789","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2605.13789"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/OliverLaboratory/Ensembits_release","reach":null}],"summary":{"ran":3,"ran_honours":2,"ran_fixture":1,"unverified":3},"by_repo_kind":{"found_in_text":{"samples":9,"ran":6,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"9ebbd7779c8d7317","entry":"SequenceEncoder","repo":"OliverLaboratory/Ensembits_release","repo_kind":"found_in_text","path":"ensembits/tokenizer.py","file_url":"https://github.com/OliverLaboratory/Ensembits_release/blob/HEAD/ensembits/tokenizer.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"9ebbd7779c8d7317"}},{"code_sha256_prefix":"9912f3a756f8cd6f","entry":"SetDecoder","repo":"OliverLaboratory/Ensembits_release","repo_kind":"found_in_text","path":"ensembits/tokenizer.py","file_url":"https://github.com/OliverLaboratory/Ensembits_release/blob/HEAD/ensembits/tokenizer.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"9912f3a756f8cd6f"}},{"code_sha256_prefix":"8cc7e5afe7bf5be3","entry":"SetTransformerEncoder","repo":"OliverLaboratory/Ensembits_release","repo_kind":"found_in_text","path":"ensembits/tokenizer.py","file_url":"https://github.com/OliverLaboratory/Ensembits_release/blob/HEAD/ensembits/tokenizer.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"8cc7e5afe7bf5be3"}},{"code_sha256_prefix":"42ecd4a2f332e343","entry":"_build_perms","repo":"OliverLaboratory/Ensembits_release","repo_kind":"found_in_text","path":"ensembits/tokenizer.py","file_url":"https://github.com/OliverLaboratory/Ensembits_release/blob/HEAD/ensembits/tokenizer.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"42ecd4a2f332e343"}},{"code_sha256_prefix":"bf8dc35deb832492","entry":"_init_orthogonal_basis","repo":"OliverLaboratory/Ensembits_release","repo_kind":"found_in_text","path":"ensembits/tokenizer.py","file_url":"https://github.com/OliverLaboratory/Ensembits_release/blob/HEAD/ensembits/tokenizer.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"bf8dc35deb832492"}},{"code_sha256_prefix":"b80c7461a3e4e62c","entry":"_permute_frames","repo":"OliverLaboratory/Ensembits_release","repo_kind":"found_in_text","path":"ensembits/tokenizer.py","file_url":"https://github.com/OliverLaboratory/Ensembits_release/blob/HEAD/ensembits/tokenizer.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"b80c7461a3e4e62c"}},{"code_sha256_prefix":"d188fd59dc9f07f4","entry":"RVQVAETokenizer","repo":"OliverLaboratory/Ensembits_release","repo_kind":"found_in_text","path":"ensembits/tokenizer.py","file_url":"https://github.com/OliverLaboratory/Ensembits_release/blob/HEAD/ensembits/tokenizer.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"d188fd59dc9f07f4"}},{"code_sha256_prefix":"49fdbdbba8440b92","entry":"VectorQuantizer","repo":"OliverLaboratory/Ensembits_release","repo_kind":"found_in_text","path":"ensembits/tokenizer.py","file_url":"https://github.com/OliverLaboratory/Ensembits_release/blob/HEAD/ensembits/tokenizer.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"49fdbdbba8440b92"}},{"code_sha256_prefix":"2128f8e4ea6b8171","entry":"hungarian_loss","repo":"OliverLaboratory/Ensembits_release","repo_kind":"found_in_text","path":"ensembits/tokenizer.py","file_url":"https://github.com/OliverLaboratory/Ensembits_release/blob/HEAD/ensembits/tokenizer.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"2128f8e4ea6b8171"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.LG","source":"arxiv_2026.jsonl"},"syntology_extracted_results":null}