Datasets › ChEBI-20

ChEBI-20

Introduced by Carl Edwards et al. in Text2Mol: Cross-Modal Molecule Retrieval with Natural Language Queries1 Nov 2021 archive 2025-07-28

Dataset contains 33,010 molecule-description pairs split into 80\%/10\%/10\% train/val/test splits. The goal of the task is to retrieve the relevant molecule for a natural language description. It is defined as follows:

To push the boundaries of multimodal models, we present a new IR task: \textbf{Text2Mol}.

Given a text query and list of molecules without any reference textual information (represented, for example, as SMILES strings, graphs, or other equivalent representations) retrieve the molecule corresponding to the query. From a text description of a molecule, the model must incorporate the information in the description into a semantic representation which can be used to directly retrieve the molecule. This requires the integration of two very different types of information: the structured knowledge represented by text and the chemical properties present in molecular graphs. We assume there is only one correct (relevant) molecule for each description, so we consider two measures for this task: Hits@1 and mean reciprocal rank (MRR).

80\% of the data is used for training. Retrieval is done against the entire corpus of molecules (train, val, test).

Benchmarks archive 2025-07-28

All 4 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Molecule Captioning ChEBI-20 Mol-LLM (Mistral-Instruct-v0.2) BLEU-2 73.2 — — 33 Compare
Text-based de novo Molecule Generation ChEBI-20 LDMol BLEU 92.6 LDMol: Text-to-Molecule Diffusion Model with... jinhojsk515/ldmol 20 Compare
Cross-Modal Retrieval ChEBI-20 CLASS (ORMA) Hits@1 67.4 CLASS: Enhancing Cross-Modal Text-Molecule Retrieval... — 9 Compare
Image Captioning ChEBI-20 GIT-Mol BLEU 0.924 GIT-Mol: A Multi-modal Large Language Model for... ai-hpc-research-team/git-mol 1 Compare

Papers archive 2025-07-28

26 shown of 26 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 43. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
XMolCap: Advancing Molecular Captioning through Multimodal Fusion and Explainable Graph Neural Networks 1 1 23 May 2025 not harvested
CLASS: Enhancing Cross-Modal Text-Molecule Retrieval Performance and Training Efficiency 0 2 17 Feb 2025 not harvested
Automatic Annotation Augmentation Boosts Translation between Molecules and Natural Language 1 3 10 Feb 2025 not harvested
Mol-LLM: Multimodal Generalist Molecular LLM with Improved Graph Utilization 0 2 5 Feb 2025 not harvested
Property Enhanced Instruction Tuning for Multi-task Molecule Generation with Large Language Models 1 1 24 Dec 2024 not harvested
MolReFlect: Towards Fine-grained In-Context Alignment between Molecules and Texts 0 2 22 Nov 2024 not harvested
Exploring Optimal Transport-Based Multi-Grained Alignments for Text-Molecule Retrieval 0 1 4 Nov 2024 not harvested
Towards Cross-Modal Text-Molecule Retrieval with Better Modality Alignment 1 1 31 Oct 2024 not harvested
Mol2Lang-VLM: Vision- and Text-Guided Generative Pre-trained Language Models for Advancing Molecule Captioning through Multimodal Fusion 1 1 15 Aug 2024 not harvested
Deep Sketched Output Kernel Regression for Structured Prediction 1 1 13 Jun 2024 not harvested
LDMol: Text-to-Molecule Diffusion Model with Structurally Informative Latent Space 1 1 28 May 2024 ran 6 of 9 samples (3 unverified; 1 pointer-only for licence)
BioT5+: Towards Generalized Biological Understanding with IUPAC Integration and Multi-task Tuning 1 2 27 Feb 2024 not harvested
Text-Guided Molecule Generation with Diffusion Language Model 1 2 20 Feb 2024 ran 14 of 21 samples (7 unverified; 21 pointer-only for licence)
InstructMol: Multi-Modal Integration for Building a Versatile and Reliable Molecular Assistant in Drug Discovery 1 2 27 Nov 2023 ran 3 of 3 samples (0 unverified)
MolCA: Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal Adapter 1 2 19 Oct 2023 ran 8 of 16 samples (8 unverified; 16 pointer-only for licence)
BioT5: Enriching Cross-modal Integration in Biology with Chemical Knowledge and Natural Language Associations 1 2 11 Oct 2023 ran 1 of 1 samples (0 unverified)
GIT-Mol: A Multi-modal Large Language Model for Molecular Science with Graph, Image, and Text 1 2 14 Aug 2023 ran 1 of 1 samples (0 unverified)
Empowering Molecule Discovery for Molecule-Caption Translation with Large Language Models: A ChatGPT Perspective 1 4 11 Jun 2023 not harvested
MolFM: A Multimodal Molecular Foundation Model 2 4 6 Jun 2023 not harvested
MolXPT: Wrapping Molecules with Text for Generative Pre-training 1 2 18 May 2023 not harvested
Adversarial Modality Alignment Network for Cross-Modal Molecule Retrieval 1 1 8 Mar 2023 not harvested
Unifying Molecular and Textual Representations via Multi-task Language Modelling 1 8 29 Jan 2023 not harvested
A Molecular Multimodal Foundation Model Associating Molecule Graphs with Natural Language 4 3 12 Sep 2022 ran 2 of 8 samples (6 unverified; 2 pointer-only for licence)
Graph-based Molecular Representation Learning 1 1 8 Jul 2022 not harvested
Translation between Molecules and Natural Language 1 7 25 Apr 2022 ran 0 of 5 samples (5 unverified)
Text2Mol: Cross-Modal Molecule Retrieval with Natural Language Queries 1 3 1 Nov 2021 not harvested

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • ChEBI-20

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections