{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/biot5-towards-generalized-biological","title":"BioT5+: Towards Generalized Biological Understanding with IUPAC Integration and Multi-task Tuning","arxiv_id":"2402.17810","date":"2024-02-27","proceeding":null,"authors":["Qizhi Pei","Lijun Wu","Kaiyuan Gao","Xiaozhuan Liang","Yin Fang","Jinhua Zhu","Shufang Xie","Tao Qin","Rui Yan"],"abstract":"Recent research trends in computational biology have increasingly focused on integrating text and bio-entity modeling, especially in the context of molecules and proteins. However, previous efforts like BioT5 faced challenges in generalizing across diverse tasks and lacked a nuanced understanding of molecular structures, particularly in their textual representations (e.g., IUPAC). This paper introduces BioT5+, an extension of the BioT5 framework, tailored to enhance biological research and drug discovery. BioT5+ incorporates several novel features: integration of IUPAC names for molecular understanding, inclusion of extensive bio-text and molecule data from sources like bioRxiv and PubChem, the multi-task instruction tuning for generality across tasks, and a numerical tokenization technique for improved processing of numerical data. These enhancements allow BioT5+ to bridge the gap between molecular representations and their textual descriptions, providing a more holistic understanding of biological entities, and largely improving the grounded reasoning of bio-text and bio-sequences. The model is pre-trained and fine-tuned with a large number of experiments, including \\emph{3 types of problems (classification, regression, generation), 15 kinds of tasks, and 21 total benchmark datasets}, demonstrating the remarkable performance and state-of-the-art results in most cases. BioT5+ stands out for its ability to capture intricate relationships in biological data, thereby contributing significantly to bioinformatics and computational biology. Our code is available at \\url{https://github.com/QizhiPei/BioT5}.","url_abs":"https://arxiv.org/abs/2402.17810v2","url_pdf":"https://arxiv.org/pdf/2402.17810v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"biot5-towards-generalized-biological","repo_url":"https://github.com/QizhiPei/BioT5","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"drug-discovery","task_name":"Drug Discovery"},{"task_slug":"forward-reaction-prediction","task_name":"Forward reaction prediction"},{"task_slug":"molecule-captioning","task_name":"Molecule Captioning"},{"task_slug":"reagent-prediction","task_name":"Reagent Prediction"},{"task_slug":"retrosynthesis","task_name":"Retrosynthesis"},{"task_slug":"text-based-de-novo-molecule-generation","task_name":"Text-based de novo Molecule Generation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/forward-reaction-prediction-on-mol","task":"Forward reaction prediction","dataset":"Mol-Instruction","model":"BioT5+","rank_in_archive_order":2,"of":2,"metrics":{"Exact":"0.864","Morgan FTS":"0.935","Validity":"1"},"uses_additional_data":false},{"leaderboard":"/sota/molecule-captioning-on-chebi-20","task":"Molecule Captioning","dataset":"ChEBI-20","model":"BioT5+","rank_in_archive_order":4,"of":33,"metrics":{"BLEU-2":"66.6","BLEU-4":"59.1","METEOR":"68.1","ROUGE-1":"71.0","ROUGE-2":"58.4","ROUGE-L":"65.0"},"uses_additional_data":false},{"leaderboard":"/sota/reagent-prediction-on-mol-instruction","task":"Reagent Prediction","dataset":"Mol-Instruction","model":"BioT5+","rank_in_archive_order":2,"of":2,"metrics":{"Exact":"0.257","Morgan FTS":"0.512","Validity":"1"},"uses_additional_data":false},{"leaderboard":"/sota/retrosynthesis-on-mol-instruction","task":"Retrosynthesis","dataset":"Mol-Instruction","model":"BioT5+","rank_in_archive_order":2,"of":2,"metrics":{"Exact":"0.642","Morgan FTS":"0.866","Validity":"1"},"uses_additional_data":false},{"leaderboard":"/sota/text-based-de-novo-molecule-generation-on","task":"Text-based de novo Molecule Generation","dataset":"ChEBI-20","model":"BioT5+","rank_in_archive_order":3,"of":20,"metrics":{"BLEU":"87.2","Exact Match":"52.2","Frechet ChemNet Distance (FCD)":"0.353","Levenshtein":"12.776","MACCS FTS":"90.7","Morgan FTS":"77.9","Parameter Count":"252000000","RDK FTS":"83.5","Text2Mol":"57.9","Validity":"100"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2402.17810","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}