Papers › BioMedGPT: Open Multimodal Generative Pre-trained Transformer for BioMedicine

BioMedGPT: Open Multimodal Generative Pre-trained Transformer for BioMedicine

18 Aug 2023arXiv:2308.09442archive 2025-07-28

Foundation models (FMs) have exhibited remarkable performance across a wide range of downstream tasks in many domains. Nevertheless, general-purpose FMs often face challenges when confronted with domain-specific problems, due to their limited access to the proprietary training data in a particular domain. In biomedicine, there are various biological modalities, such as molecules, proteins, and cells, which are encoded by the language of life and exhibit significant modality gaps with human natural language. In this paper, we introduce BioMedGPT, an open multimodal generative pre-trained transformer (GPT) for biomedicine, to bridge the gap between the language of life and human natural language. BioMedGPT allows users to easily ``communicate'' with diverse biological modalities through free text, which is the first of its kind. BioMedGPT aligns different biological modalities with natural language via a large generative language model, namely, BioMedGPT-LM. We publish BioMedGPT-10B, which unifies the feature spaces of molecules, proteins, and natural language via encoding and alignment. Through fine-tuning, BioMedGPT-10B outperforms or is on par with human and significantly larger general-purpose foundation models on the biomedical QA task. It also demonstrates promising performance in the molecule QA and protein QA tasks, which could greatly accelerate the discovery of new drugs and therapeutic targets. In addition, BioMedGPT-LM-7B is the first large generative language model based on Llama2 in the biomedical domain, therefore is commercial friendly. Both BioMedGPT-10B and BioMedGPT-LM-7B are open-sourced to the research community. In addition, we publish the datasets that are meticulously curated for the alignment of multi-modalities, i.e., PubChemQA and UniProtQA. All the models, codes, and datasets are available at \url{https://github.com/PharMolix/OpenBioMed}.

PaperPDFCode

In Syntology View this paper on Syntology: its repositories, every harvested function with whether it ran, its licence and the call to fetch it.

Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

pharmolix/openbiomed officialmentioned in papermentioned on GitHubpytorchMIT report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Few-Shot LearningLanguage ModelingLanguage ModellingMultiple Choice Question Answering (MCQA)Question AnsweringZero-Shot Learning

Datasets

Introduced by this paper, per the archive.

PubChemQAUniProtQA

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Few-Shot Learning MedConceptsQA PharMolix/BioMedGPT-LM-7B Accuracy 24.924 #11 of 12 Archive leaderboard report
Multiple Choice Question Answering (MCQA) MMLU (Professional medicine) BioMedGPT-LM-7B Accuracy 51.1 #4 of 6 Archive leaderboard report
Multiple Choice Question Answering (MCQA) MedMCQA BioMedGPT-10B Test Set (Acc-%) 0.514 #6 of 22 Archive leaderboard report
Question Answering MedQA BioMedGPT-10B Accuracy 50.4 #16 of 27 Archive leaderboard report
Question Answering PubChemQA BioMedGPT-10B BLEU-2 0.234 #1 of 2 Archive leaderboard report
Question Answering PubChemQA BioMedGPT-10B BLEU-4 0.141 #1 of 2 Archive leaderboard report
Question Answering PubChemQA BioMedGPT-10B MEATOR 0.308 #1 of 2 Archive leaderboard report
Question Answering PubChemQA BioMedGPT-10B ROUGE-1 0.386 #1 of 2 Archive leaderboard report
Question Answering PubChemQA BioMedGPT-10B ROUGE-2 0.206 #1 of 2 Archive leaderboard report
Question Answering PubChemQA BioMedGPT-10B ROUGE-L 0.332 #1 of 2 Archive leaderboard report
Question Answering PubMedQA BioMedGPT-10B Accuracy 76.1 #14 of 30 Archive leaderboard report
Question Answering UniProtQA BioMedGPT-10B BLEU-2 0.571 #1 of 2 Archive leaderboard report
Question Answering UniProtQA BioMedGPT-10B BLEU-4 0.535 #1 of 2 Archive leaderboard report
Question Answering UniProtQA BioMedGPT-10B MEATOR 0.754 #1 of 2 Archive leaderboard report
Question Answering UniProtQA BioMedGPT-10B ROUGE-1 0.743 #1 of 2 Archive leaderboard report
Question Answering UniProtQA BioMedGPT-10B ROUGE-2 0.759 #1 of 2 Archive leaderboard report
Question Answering UniProtQA BioMedGPT-10B ROUGE-L 0.622 #1 of 2 Archive leaderboard report
Zero-Shot Learning MedConceptsQA PharMolix/BioMedGPT-LM-7B Accuracy 24.747 #11 of 13 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections