Datasets › USPTO-30K

USPTO-30K

Introduced by Lucas Morin et al. in MolGrapher: Graph-based Visual Recognition of Chemical Structures23 Aug 2023 archive 2025-07-28

We introduce USPTO-30K, a large-scale benchmark dataset of annotated molecule images, which overcomes these limitations. It is created using the pairs of images and MolFiles by the United States Patent and Trademark Office. Each molecule was independently selected among all the available documents from 2001 to 2020. The set consists of three subsets to decouple the study of clean molecules, molecules with abbreviations and large molecules.

USPTO-10K contains 10,000 clean molecules, i.e. without any abbreviated groups. USPTO-10K-abb contains 10,000 molecules with superatom groups. USPTO-10K-L contains 10,000 clean molecules with more than 70 atoms.

Benchmarks archive 2025-07-28

No leaderboard in the archive resolves to this dataset.

Papers archive 2025-07-28

No paper in the archive has a leaderboard row on this dataset; the archive counts 1 paper for it but never published that list.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

No task tagged in the archive.

License archive 2025-07-28

CDLA-Permissive-1.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • USPTO-30K

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections