{"url":"/dataset/uk-key-stage-readability","name":"UK Key Stage Readability","full_name":"UK Key Stage Readability for English Texts","description_markdown":"Education is increasingly data-driven, and the ability to analyse and adapt educational materials quickly and effectively is important for keeping materials contemporary and interesting. These approaches also have the potential to personalise learning experiences. One of the challenges in this domain is aligning new literature with the appropriate educational stages. This dataset aims to contribute to alleviating this knowledge gap. \r\n\r\nThis dataset has been generated through literature in the public domain from Project Gutenberg, and cross-referenced by the UK Key Stage equivalents from the Lexile Reading Framework. \r\n\r\nThe dataset contains a total of 20,000 rows evenly distributed across four educational stages - Key Stage 2 (KS2), Key Stage 3 (KS3), Key Stage 4 (KS4), and Key Stage 5 (KS5).\r\n\r\nThe data has been split into Train (80%, 16,000 objects) and Test (20%, 4,000 objects) sets. \r\n\r\nThe data is multimodal and contains:\r\n- Text - the cropped excerpt of text, which is limited to 512 tokens to the nearest complete sentence. \r\n- Linguistic Features - each extracted from the text excerpt","description_withheld":null,"homepage":"https://www.kaggle.com/datasets/birdy654/uk-key-stage-readability-for-english-texts","introduced_date":"2024-11-26","introduced_date_note":null,"introduced_by":{"paper":"/paper/what-differentiates-educational-literature-a","title":"What Differentiates Educational Literature? A Multimodal Fusion Approach of Transformers and Computational Linguistics","first_author":"Jordan J. Bird","url":null},"license":{"name":"MIT","url":null},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Text Classification","url":"/task/text-classification","datasets_with_task":"/datasets/task/text-classification"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["UK Key Stage Readability"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/text-classification-on-uk-key-stage","task":"Text Classification","dataset_variant":"UK Key Stage Readability","rows":15,"metrics":["F1"],"first_row_in_archive_order":{"model":"ELECTRA + ANN","paper":"/paper/what-differentiates-educational-literature-a","metrics":{"F1":"99.6"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/what-differentiates-educational-literature-a","title":"What Differentiates Educational Literature? A Multimodal Fusion Approach of Transformers and Computational Linguistics","date":"2024-11-26","rows_on_this_dataset":15,"code_links":0,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}