{"url":"/dataset/jemma","name":"JEMMA","full_name":null,"description_markdown":"**JEMMA** is an Extensible Java Dataset for ML4Code Applications, which is a large-scale dataset targeted at ML4 code. JEMMA comes with a considerable amount of pre-processed information such as metadata, representations (e.g., code tokens, ASTs, graphs), and several properties (e.g., metrics, static analysis results) for 50,000 Java projects from the 50KC dataset, with over 1.2 million classes and over 8 million methods.\r\n\r\nSource: [JEMMA: An Extensible Java Dataset for ML4Code Applications](https://arxiv.org/pdf/2212.09132v1.pdf)\r\n\r\nImage Source: [https://arxiv.org/pdf/2212.09132v1.pdf](https://arxiv.org/pdf/2212.09132v1.pdf)","description_withheld":null,"homepage":"https://github.com/giganticode/jemma","introduced_date":"2022-12-18","introduced_date_note":null,"introduced_by":{"paper":"/paper/jemma-an-extensible-java-dataset-for-ml4code","title":"JEMMA: An Extensible Java Dataset for ML4Code Applications","first_author":"Anjan Karmakar","url":null},"license":null,"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[],"languages":[],"variants":["JEMMA"],"data_loaders":[],"num_papers_in_archive":2,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-25T09:33:49+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}