Papers › XMolCap: Advancing Molecular Captioning through Multimodal Fusion and Explainable...

XMolCap: Advancing Molecular Captioning through Multimodal Fusion and Explainable Graph Neural Networks

23 May 2025IEEE Journal of Biomedical and Health Informatics 2025 5archive 2025-07-28

Duong Thanh Tran, Nguyen Doan Hieu Nguyen, Nhat Truong Pham, Rajan Rakkiyappan, Rajendra Karki, Balachandran Manavalan

Large language models (LLMs) have significantly advanced computational biology by enabling the integration of molecular, protein, and natural language data to accelerate drug discovery. However, existing molecular captioning approaches often underutilize diverse molecular modalities and lack interpretability. In this study, we introduce XMolCap, a novel explainable molecular captioning framework that integrates molecular images, SMILES strings, and graph-based structures through a stacked multimodal fusion mechanism. The framework is built upon a BioT5-based encoder-decoder architecture, which serves as the backbone for extracting feature representations from SELFIES. By leveraging specialized models such as SwinOCSR, SciBERT, and GIN-MoMu, XMolCap effectively captures complementary information from each modality. Our model not only achieves state-of-the-art performance on two benchmark datasets (L+M-24 and ChEBI-20), outperforming several strong baselines, but also provides detailed, functional group-aware, and property-specific explanations through graph-based interpretation. XMolCap is publicly available at https://github.com/cbbl-skku-org/XMolCap/ for reproducibility and local deployment. We believe it holds strong potential for clinical and pharmaceutical applications by generating accurate, interpretable molecular descriptions that deepen our understanding of molecular properties and interactions.

PaperPDFCode

Code

cbbl-skku-org/XMolCap mentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Drug DiscoveryMolecule Captioning

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Molecule Captioning ChEBI-20 XMolCap BLEU-2 62.0 #7 of 33 Archive leaderboard report
Molecule Captioning ChEBI-20 XMolCap BLEU-4 53.8 #7 of 33 Archive leaderboard report
Molecule Captioning ChEBI-20 XMolCap METEOR 67.8 #7 of 33 Archive leaderboard report
Molecule Captioning ChEBI-20 XMolCap ROUGE-1 53.9 #7 of 33 Archive leaderboard report
Molecule Captioning ChEBI-20 XMolCap ROUGE-2 61.8 #7 of 33 Archive leaderboard report
Molecule Captioning ChEBI-20 XMolCap ROUGE-L 63.8 #7 of 33 Archive leaderboard report
Molecule Captioning L+M-24 XMolCap BLEU-2 77.4 #2 of 6 Archive leaderboard report
Molecule Captioning L+M-24 XMolCap BLEU-4 56.0 #2 of 6 Archive leaderboard report
Molecule Captioning L+M-24 XMolCap METEOR 73.8 #2 of 6 Archive leaderboard report
Molecule Captioning L+M-24 XMolCap ROUGE-1 78.9 #2 of 6 Archive leaderboard report
Molecule Captioning L+M-24 XMolCap ROUGE-2 59.4 #2 of 6 Archive leaderboard report
Molecule Captioning L+M-24 XMolCap ROUGE-L 56.8 #2 of 6 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections