{"url":"/dataset/pseudomd-1m","name":"PseudoMD-1M","full_name":null,"description_markdown":"Pre-training dataset used in paper \"[From Artificially Real to Real: Leveraging Pseudo Data from Large Language Models for Low-Resource Molecule Discovery](https://arxiv.org/abs/2309.05203)\" (AAAI 2024)\r\n\r\nPseudoMD-1M dataset is the first artificially-real dataset for cross-modal molecule discovery, which consists of 1,020,139 pseudo molecule-description pairs. Every molecule is represented using its Canonical SMILES notation, sourced from PubChem via the PUG View API. On average, each description within PseudoMD-1M contains 5.11 sentences, 106.47 words, and 165.07 tokens.\r\n\r\n### Citation\r\nIf you found the dataset useful, please cite:\r\n```bibtex\r\n@article{chen2023artificially,\r\n  title={From Artificially Real to Real: Leveraging Pseudo Data from Large Language Models for Low-Resource Molecule Discovery},\r\n  author={Chen, Yuhan and Xi, Nuwa and Du, Yanrui and Wang, Haochun and Jianyu, Chen and Zhao, Sendong and Qin, Bing},\r\n  journal={arXiv preprint arXiv:2309.05203},\r\n  year={2023}\r\n}\r\n```","description_withheld":null,"homepage":"https://huggingface.co/datasets/SCIR-HI/PseudoMD-1M","introduced_date":"2023-09-11","introduced_date_note":null,"introduced_by":null,"license":{"name":"apache-2.0","url":null},"modalities":[],"tasks":[],"languages":[],"variants":["PseudoMD-1M"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}