{"url":"/dataset/tekgen","name":"TekGen","full_name":null,"description_markdown":"The Dataset is part of the KELM corpus\r\n\r\nThis is the Wikipedia text--Wikidata KG aligned corpus used to train the data-to-text generation model. Please note that this is a corpus generated with distant supervision and should not be used as gold standard for evaluation.\r\n\r\nIt consists of 3 files:\r\n\r\n    https://storage.googleapis.com/gresearch/kelm-corpus/updated-2021/quadruples-train.tsv\r\n    https://storage.googleapis.com/gresearch/kelm-corpus/updated-2021/quadruples-validation.tsv\r\n    https://storage.googleapis.com/gresearch/kelm-corpus/updated-2021/quadruples-test.tsv\r\n\r\nEach file contains one example per line. Each example is a json object with three fields:\r\n\r\n    triples: A list of triples of the form (subject, relation, object). eg. (Person X, award received, Award Y). If the triple has a subproperty, then it is quadruple instead. eg. (Person X, Award Y, received on, Date Z).\r\n\r\n    serialized triples: triples concatenated together as used for input to T5. The format is \"<subject> <relation> <object>\" where some subjects have multiple relations, e.g. \"<subject> <relation1> <object1> <relation2> <object2> <relation3> <object3>\". For more details on how these relations are grouped, please refer to the paper.\r\n\r\n    sentence: The wikipedia sentence aligned to these triples.\r\n\r\nThe names, aliases and Wikidata Ids of the entities can be found in https://storage.googleapis.com/gresearch/kelm-corpus/updated-2021/entities.jsonl.","description_withheld":null,"homepage":"https://github.com/google-research-datasets/KELM-corpus#part-1-tekgen-training-corpus","introduced_date":"2020-10-23","introduced_date_note":null,"introduced_by":{"paper":"/paper/large-scale-knowledge-graph-based-synthetic","title":"Knowledge Graph Based Synthetic Corpus Generation for Knowledge-Enhanced Language Model Pre-training","first_author":"Oshin Agarwal","url":null},"license":{"name":"CC BY-SA 2.0 license","url":"https://github.com/google-research-datasets/KELM-corpus"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"},{"name":"Graphs","url":"/datasets/modality/graphs"}],"tasks":[{"name":"Joint Entity and Relation Extraction","url":"/task/joint-entity-and-relation-extraction","datasets_with_task":"/datasets/task/joint-entity-and-relation-extraction"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["TekGen"],"data_loaders":[],"num_papers_in_archive":12,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/joint-entity-and-relation-extraction-on-9","task":"Joint Entity and Relation Extraction","dataset_variant":"TekGen","rows":2,"metrics":["F1"],"first_row_in_archive_order":{"model":"ReGen-SCST","paper":"/paper/regen-reinforcement-learning-for-text-and","metrics":{"F1":"62.3"},"code_links":[{"title":"IBM/regen","url":"https://github.com/IBM/regen"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/regen-reinforcement-learning-for-text-and","title":"ReGen: Reinforcement Learning for Text and Knowledge Base Generation using Pretrained Language Models","date":"2021-08-27","rows_on_this_dataset":2,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":0,"samples_unverified":1,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":1,"samples_harvested":1,"samples_ran":0,"samples_unverified":1,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":1,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}