Datasets › 2010 i2b2/VA
2010 i2b2/VA
2010 i2b2/VA is a biomedical dataset for relation classification and entity typing.
Benchmarks archive 2025-07-28
All 4 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Clinical Concept Extraction | 2010 i2b2/VA | BERTlarge (MIMIC) Exact Span F1 90.25 | Enhancing Clinical Concept Extraction with Contextual Embeddings | — | 5 | Compare |
| Clinical Assertion Status Detection | 2010 i2b2/VA | BiLSTM (SparkNLP) Micro F1 0.939 | Improving Clinical Document Understanding on COVID-19... | JohnSnowLabs/spark-nlp-workshop | 1 | Compare |
| FG-1-PG-1 | 2010 i2b2/VA | CFNER F1 (macro) 0.3626 | Distilling Causal Effect from Miscellaneous Other-Class... | zzz47zzz/CFNER | 1 | Compare |
| Relation Extraction | 2010 i2b2/VA | Spark NLP Macro F1 69.1 | Deeper Clinical Document Understanding Using Relation Extraction | JohnSnowLabs/spark-nlp-workshop | 1 | Compare |
Papers archive 2025-07-28
8 shown of 8 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 18. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| Distilling Causal Effect from Miscellaneous Other-Class for Continual Named Entity Recognition | 1 | 1 | 8 Oct 2022 | not harvested |
| Deeper Clinical Document Understanding Using Relation Extraction | 1 | 1 | 25 Dec 2021 | not harvested |
| Improving Clinical Document Understanding on COVID-19 Research with Spark NLP | 1 | 1 | 7 Dec 2020 | not harvested |
| CharacterBERT: Reconciling ELMo and BERT for Word-Level Open-Vocabulary Representations From Characters | 2 | 1 | 20 Oct 2020 | ran 5 of 8 samples (3 unverified) |
| Cost-effective Selection of Pretraining Data: A Case Study of Pretraining BERT on Social Media | 0 | 1 | 2 Oct 2020 | not harvested |
| Embedding Strategies for Specialized Domains: Application to Clinical Entity Recognition | 1 | 1 | 1 Jul 2019 | not harvested |
| Enhancing Clinical Concept Extraction with Contextual Embeddings | 0 | 1 | 22 Feb 2019 | not harvested |
| Machine-learned solutions for three stages of clinical information extraction: the state of the art at i2b2 2010 | 0 | 1 | 12 May 2011 | not harvested |
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
No licence recorded in the archive. Absence here is not a statement about the dataset's terms.
Modalities archive 2025-07-28
Languages archive 2025-07-28
No language tagged.
Variants archive 2025-07-28
- 2010 i2b2/VA
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections