{"url":"/dataset/identity-access-management-dataset","name":"Identity Access Management dataset","full_name":null,"description_markdown":"We release 280 synthetic IAM graphs generated using IAM graphs of commercial companies.\r\nSpecifically, we vary the number of nodes, but keep graph density as is, i.e. in the range of 0.259 ± 0.198 (avg ± std). \r\nTo generate a synthetic graph, we first\r\nsample the number of users and datastores from uniform distributions over the following intervals [10, 150] and [50, 300]\r\nrespectively that cover variations of those parameters across\r\nreal graphs. After fixing node counts we sample with replacement\r\nthe actual nodes from a real world graph, which is chosen\r\nat random. Then we add Gaussian N(0, 0.01) noise to node\r\nembeddings and renormalize them. To match the graph density with the density of the underlying baseline we sample\r\nedges from a multinomial distribution, where each component is proportional to the cosine distance between a user and\r\na datastore embeddings. Also we enforce the invariant that\r\ndynamic edges are always a subset of all permission edges.\r\nA synthetic graph generated in such a way is an ”upsampled”\r\nversion of an underlying real world graph.","description_withheld":null,"homepage":"https://github.com/mikhail247/IAMAX","introduced_date":"2022-05-02","introduced_date_note":null,"introduced_by":{"paper":"/paper/using-constraint-programming-and-graph","title":"Using Constraint Programming and Graph Representation Learning for Generating Interpretable Cloud Security Policies","first_author":"Mikhail Kazdagli","url":null},"license":null,"modalities":[{"name":"Graphs","url":"/datasets/modality/graphs"}],"tasks":[],"languages":[],"variants":["Identity Access Management dataset"],"data_loaders":[{"repo":"https://github.com/mikhail247/iamax","url":"https://github.com/mikhail247/iamax","frameworks":[]}],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}