{"url":"/dataset/owt2","name":"OWT2","full_name":"OpenWebtext2","description_markdown":"**OpenWebText2** is an enhanced version of the original **OpenWebTextCorpus**. It encompasses all **Reddit submissions** from **2005** up until **April 2020**, with additional months becoming available after the corresponding **PushShift dump files** are released¹²³. Here are the key details:\r\n\r\n- **Dataset Description**:\r\n    - **OpenWebText2** is part of the **EleutherAi/The Pile dataset**.\r\n    - It covers **Reddit submissions** and is designed to be a high-quality internet dataset.\r\n    - The dataset was created by scraping **URLs** extracted from **Reddit submissions** with a minimum score of **3** as a proxy for quality⁴.\r\n    - The **plug-and-play version** of **OpenWebText2** contains **17,103,059 documents** and is approximately **65.86 GB** when uncompressed².\r\n\r\n- **Features**:\r\n    - Each document in **OpenWebText2** has two main features:\r\n        - **Title**: A string representing the title of the submission.\r\n        - **Text**: A string containing the content of the submission¹.\r\n\r\n- **License and Version**:\r\n    - **License**: No known license.\r\n    - **Version**: 1.0.0¹.\r\n\r\n- **Size**:\r\n    - **Download Size**: Approximately **27.3 GiB**.\r\n    - **Dataset Size**: Approximately **63.8 GiB**³.\r\n\r\n(1) the_pile_openwebtext2 | TensorFlow Datasets. https://www.tensorflow.org/datasets/community_catalog/huggingface/the_pile_openwebtext2.\r\n(2) GitHub - EleutherAI/openwebtext2. https://github.com/EleutherAI/openwebtext2.\r\n(3) the_pile_openwebtext2 · Datasets at Hugging Face. https://huggingface.co/datasets/the_pile_openwebtext2.\r\n(4) Eleuther AI site | OpenWebText2. https://researcher2.eleuther.ai/projects/open-web-text2/.\r\n(5) OpenWebText2 — EleutherAI. https://www.eleuther.ai/artifacts/openwebtext2.","description_withheld":null,"homepage":"https://github.com/EleutherAI/openwebtext2","introduced_date":null,"introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["OWT2"],"data_loaders":[],"num_papers_in_archive":2,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/language-modelling-on-openwebtext2","task":"Language Modelling","dataset_variant":"OpenWebtext2","rows":1,"metrics":["BPB"],"first_row_in_archive_order":{"model":"Gopher","paper":"/paper/scaling-language-models-methods-analysis-1","metrics":{"BPB":"0.677"},"code_links":[{"title":"allenai/dolma","url":"https://github.com/allenai/dolma"},{"title":"rvlopes/gloria","url":"https://github.com/rvlopes/gloria"},{"title":"bramiozo/PubScience","url":"https://github.com/bramiozo/PubScience"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/scaling-language-models-methods-analysis-1","title":"Scaling Language Models: Methods, Analysis & Insights from Training Gopher","date":"2021-12-08","rows_on_this_dataset":1,"code_links":3,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}