Papers › Text2Mol: Cross-Modal Molecule Retrieval with Natural Language Queries

Text2Mol: Cross-Modal Molecule Retrieval with Natural Language Queries

1 Nov 2021EMNLP 2021 11archive 2025-07-28

Carl Edwards, ChengXiang Zhai, Heng Ji

We propose a new task, Text2Mol, to retrieve molecules using natural language descriptions as queries. Natural language and molecules encode information in very different ways, which leads to the exciting but challenging problem of integrating these two very different modalities. Although some work has been done on text-based retrieval and structure-based retrieval, this new task requires integrating molecules and natural language more directly. Moreover, this can be viewed as an especially challenging cross-lingual retrieval problem by considering the molecules as a language with a very unique grammar. We construct a paired dataset of molecules and their corresponding text descriptions, which we use to learn an aligned common semantic embedding space for retrieval. We extend this to create a cross-modal attention-based model for explainability and reranking by interpreting the attentions as association rules. We also employ an ensemble approach to integrate our different architectures, which significantly improves results from 0.372 to 0.499 MRR. This new multimodal approach opens a new perspective on solving problems in chemistry literature understanding and molecular machine learning.

PaperPDFCode

Code

cnedwards/text2mol officialmentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Cross-Modal RetrievalNatural Language QueriesRerankingRetrieval

Datasets

Introduced by this paper, per the archive.

ChEBI-20

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Cross-Modal Retrieval ChEBI-20 All-Ensemble Hits@1 34.4 #7 of 9 Archive leaderboard report
Cross-Modal Retrieval ChEBI-20 All-Ensemble Hits@10 81.1 #7 of 9 Archive leaderboard report
Cross-Modal Retrieval ChEBI-20 All-Ensemble Mean Rank 20.21 #7 of 9 Archive leaderboard report
Cross-Modal Retrieval ChEBI-20 All-Ensemble Test MRR 49.9 #7 of 9 Archive leaderboard report
Cross-Modal Retrieval ChEBI-20 MLP1 Hits@1 22.4 #8 of 9 Archive leaderboard report
Cross-Modal Retrieval ChEBI-20 MLP1 Hits@10 68.6 #8 of 9 Archive leaderboard report
Cross-Modal Retrieval ChEBI-20 MLP1 Mean Rank 30.38 #8 of 9 Archive leaderboard report
Cross-Modal Retrieval ChEBI-20 MLP1 Test MRR 37.2 #8 of 9 Archive leaderboard report
Cross-Modal Retrieval ChEBI-20 GCN2 Hits@1 22.3 #9 of 9 Archive leaderboard report
Cross-Modal Retrieval ChEBI-20 GCN2 Hits@10 68.9 #9 of 9 Archive leaderboard report
Cross-Modal Retrieval ChEBI-20 GCN2 Mean Rank 41.90 #9 of 9 Archive leaderboard report
Cross-Modal Retrieval ChEBI-20 GCN2 Test MRR 37.1 #9 of 9 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections