{"url":"/dataset/blurb","name":"BLURB","full_name":"Biomedical Language Understanding and Reasoning Benchmark","description_markdown":"**BLURB** is a collection of resources for biomedical natural language processing. In general domains such as newswire and the Web, comprehensive benchmarks and leaderboards such as GLUE have greatly accelerated progress in open-domain NLP. In biomedicine, however, such resources are ostensibly scarce. In the past, there have been a plethora of shared tasks in biomedical NLP, such as BioCreative, BioNLP Shared Tasks, SemEval, and BioASQ, to name just a few. These efforts have played a significant role in fueling interest and progress by the research community, but they typically focus on individual tasks. The advent of neural language models such as BERTs provides a unifying foundation to leverage transfer learning from unlabeled text to support a wide range of NLP applications. To accelerate progress in biomedical pretraining strategies and task-specific methods, it is thus imperative to create a broad-coverage benchmark encompassing diverse biomedical tasks.\r\n\r\nInspired by prior efforts toward this direction (e.g., BLUE), BLURB (short for Biomedical Language Understanding and Reasoning Benchmark) was created. BLURB comprises of a comprehensive benchmark for PubMed-based biomedical NLP applications, as well as a leaderboard for tracking progress by the community. BLURB includes thirteen publicly available datasets in six diverse tasks. To avoid placing undue emphasis on tasks with many available datasets, such as named entity recognition (NER), BLURB reports the macro average across all tasks as the main score. The BLURB leaderboard is model-agnostic. Any system capable of producing the test predictions using the same training and development data can participate. The main goal of BLURB is to lower the entry barrier in biomedical NLP and help accelerate progress in this vitally important field for positive societal and human impact.\r\n\r\nSource: [BLURB](https://microsoft.github.io/BLURB/index.html)\r\n\r\nImage source: [BLURB](https://microsoft.github.io/BLURB/index.html)","description_withheld":null,"homepage":"https://microsoft.github.io/BLURB/index.html","introduced_date":"2020-07-31","introduced_date_note":null,"introduced_by":{"paper":"/paper/domain-specific-language-model-pretraining","title":"Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing","first_author":"Yu Gu","url":null},"license":null,"modalities":[{"name":"Texts","url":"/datasets/modality/texts"},{"name":"Biomedical","url":"/datasets/modality/biomedical"}],"tasks":[{"name":"Question Answering","url":"/task/question-answering","datasets_with_task":"/datasets/task/question-answering"},{"name":"Text Classification","url":"/task/text-classification","datasets_with_task":"/datasets/task/text-classification"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["BLURB"],"data_loaders":[],"num_papers_in_archive":40,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/question-answering-on-blurb","task":"Question Answering","dataset_variant":"BLURB","rows":4,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"BioLinkBERT (large)","paper":"/paper/linkbert-pretraining-language-models-with","metrics":{"Accuracy":"83.5"},"code_links":[{"title":"michiyasunaga/LinkBERT","url":"https://github.com/michiyasunaga/LinkBERT"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/text-classification-on-blurb","task":"Text Classification","dataset_variant":"BLURB","rows":3,"metrics":["F1"],"first_row_in_archive_order":{"model":"BioLinkBERT (large)","paper":"/paper/linkbert-pretraining-language-models-with","metrics":{"F1":"84.88"},"code_links":[{"title":"michiyasunaga/LinkBERT","url":"https://github.com/michiyasunaga/LinkBERT"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/evaluation-of-large-language-model","title":"Evaluation of large language model performance on the Biomedical Language Understanding and Reasoning Benchmark","date":"2024-05-17","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/linkbert-pretraining-language-models-with","title":"LinkBERT: Pretraining Language Models with Document Links","date":"2022-03-29","rows_on_this_dataset":4,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":14,"samples_ran":0,"samples_unverified":14,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/domain-specific-language-model-pretraining","title":"Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing","date":"2020-07-31","rows_on_this_dataset":2,"code_links":2,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":1,"samples_harvested":14,"samples_ran":0,"samples_unverified":14,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":1,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}