Browse State-of-the-Art › Text-based de novo Molecule Generation
Text-based de novo Molecule Generation
13 papers with code · 1 benchmark · 3 datasets archive 2025-07-28
Text-based de novo molecule generation involves utilizing natural language processing (NLP) techniques and chemical information to generate entirely new molecular structures. In this approach, molecular structures are typically encoded as text strings, resembling chemical formulas or SMILES (Simplified Molecular Input Line Entry System). Subsequently, by applying NLP models such as recurrent neural networks (RNNs) or Transformer models, these text strings are processed to generate novel molecular structures with desired properties.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| ChEBI-20 (20 rows) | LDMol | LDMol: Text-to-Molecule Diffusion Model with Structurally... | code | Syntology ran 6 of 9 samples · 3 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
3 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
13 shown of 13 papers with code (14 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
6 Jun 2023 2 repositories listedIn this study, we introduce MolFM, a multimodal molecular foundation model designed to facilitate joint representation learning from molecular structures, biomedical texts, and knowledge graphs.
-
10 Feb 2025 1 repository listedRecent advancements in AI for biological research focus on integrating molecular data with natural language to accelerate drug discovery.
-
16 Dec 2024 1 repository listedGenerating novel molecules with higher properties than the training space, namely the out-of-distribution generation, is important for de novo drug design.
-
28 Jul 2024 1 repository listedIn this work, we introduce ChemBFN, a language model that handles chemistry tasks based on Bayesian flow networks working on discrete data.
-
28 May 2024 1 repository listed Syntology ran 6 of 9 samples · 3 unverified · 1 pointer-only (licence)With the emergence of diffusion models as the frontline of generative models, many researchers have proposed molecule generation techniques with conditional diffusion models.
-
27 Feb 2024 1 repository listedHowever, previous efforts like BioT5 faced challenges in generalizing across diverse tasks and lacked a nuanced understanding of molecular structures, particularly in their textual representations (e.
-
20 Feb 2024 1 repository listed Syntology ran 14 of 21 samples · 7 unverified · 21 pointer-only (licence)In this work, we propose the Text-Guided Molecule Generation with Diffusion Language Model (TGM-DLM), a novel approach that leverages diffusion models to address the limitations of autoregressive methods.
-
11 Oct 2023 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedRecent advancements in biological research leverage the integration of molecules, proteins, and natural language to enhance drug discovery.
-
14 Aug 2023 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedLarge language models have made significant strides in natural language processing, enabling innovative applications in molecular science by processing textual representations of molecules.
-
11 Jun 2023 1 repository listedIn this work, we propose a novel LLM-based framework (MolReGPT) for molecule-caption translation, where an In-Context Few-Shot Molecule Learning paradigm is introduced to empower molecule discovery with LLMs like…
-
18 May 2023 1 repository listedConsidering that text is the most important record for scientific discovery, in this paper, we propose MolXPT, a unified language model of text and molecules pre-trained on SMILES (a sequence representation of…
-
29 Jan 2023 1 repository listedHere, we propose the first multi-domain, multi-task language model that can solve a wide range of tasks in both the chemical and natural language domains.
-
25 Apr 2022 1 repository listed Syntology ran 0 of 5 samples · 5 unverifiedWe present MolT5 - a self-supervised learning framework for pretraining models on a vast amount of unlabeled natural language text and molecule strings.
Syntology lines on 5 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections