{"url":"/dataset/ham","name":"HAM","full_name":"Human-annotated Mappings","description_markdown":"**HAM** is a dataset for molecular graph partitioning. This dataset contains coarse-grained (CG) mappings of 1206 organic molecules with less than 25 heavy atoms. Each molecule was downloaded from the PubChem database as SMILES. One molecule was assigned to two annotators to compare the human agreement between CG mappings. Downloaded SMILES were hand-mapped. The completed annotations were reviewed by a third person, to identify and remove unreasonable mappings (eg: one bead mappings) which did not agree with the given guidelines. Hence, there are 1.68 annotations per molecule in the current database (16% removed).\n\nSource: [https://github.com/rochesterxugroup/HAM_dataset](https://github.com/rochesterxugroup/HAM_dataset)","description_withheld":null,"homepage":"https://github.com/rochesterxugroup/HAM_dataset","introduced_date":null,"introduced_date_note":null,"introduced_by":{"paper":"/paper/graph-neural-network-based-coarse-grained","title":"Graph Neural Network Based Coarse-Grained Mapping Prediction","first_author":"Zhiheng Li","url":null},"license":null,"modalities":[{"name":"Graphs","url":"/datasets/modality/graphs"}],"tasks":[{"name":"Metric Learning","url":"/task/metric-learning","datasets_with_task":"/datasets/task/metric-learning"},{"name":"graph partitioning","url":"/task/graph-partitioning","datasets_with_task":"/datasets/task/graph-partitioning"}],"languages":[],"variants":["HAM"],"data_loaders":[{"repo":"https://github.com/rochesterxugroup/HAM_dataset","url":"https://github.com/rochesterxugroup/HAM_dataset","frameworks":[]}],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}