{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/the-refinedweb-dataset-for-falcon-llm","title":"The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only","arxiv_id":"2306.01116","date":"2023-06-01","proceeding":null,"authors":["Guilherme Penedo","Quentin Malartic","Daniel Hesslow","Ruxandra Cojocaru","Alessandro Cappelli","Hamza Alobeidli","Baptiste Pannier","Ebtesam Almazrouei","Julien Launay"],"abstract":"Large language models are commonly trained on a mixture of filtered web data and curated high-quality corpora, such as social media conversations, books, or technical papers. This curation process is believed to be necessary to produce performant models with broad zero-shot generalization abilities. However, as larger models requiring pretraining on trillions of tokens are considered, it is unclear how scalable is curation and whether we will run out of unique high-quality data soon. At variance with previous beliefs, we show that properly filtered and deduplicated web data alone can lead to powerful models; even significantly outperforming models from the state-of-the-art trained on The Pile. Despite extensive filtering, the high-quality data we extract from the web is still plentiful, and we are able to obtain five trillion tokens from CommonCrawl. We publicly release an extract of 600 billion tokens from our RefinedWeb dataset, and 1.3/7.5B parameters language models trained on it.","url_abs":"https://arxiv.org/abs/2306.01116v1","url_pdf":"https://arxiv.org/pdf/2306.01116v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"the-refinedweb-dataset-for-falcon-llm","repo_url":"https://github.com/ai21labs/factor","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"the-refinedweb-dataset-for-falcon-llm","repo_url":"https://github.com/MindSpore-scientific/code-14/tree/main/The_RefinedWeb_Dataset_for_Falcon","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null}],"tasks":[{"task_slug":"zero-shot-generalization","task_name":"Zero-shot Generalization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2306.01116","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2306.01116"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/MindSpore-scientific/code-14/tree/main/The_RefinedWeb_Dataset_for_Falcon","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ai21labs/factor","reach":null}],"summary":{"unverified":4},"by_repo_kind":{"listed":{"samples":4,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"fc721854c3bf3762","entry":"labels2cat","repo":"MindSpore-scientific/code-14","repo_kind":"listed","path":"Three-Dimensional-Lip-Motion-Network-for-Text-Independent-Speaker-Recognition-master/3LMNet_ms.py","file_url":"https://github.com/MindSpore-scientific/code-14/blob/HEAD/Three-Dimensional-Lip-Motion-Network-for-Text-Independent-Speaker-Recognition-master/3LMNet_ms.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"fc721854c3bf3762"}},{"code_sha256_prefix":"962937e4c560ebb3","entry":"labels2onehot","repo":"MindSpore-scientific/code-14","repo_kind":"listed","path":"Three-Dimensional-Lip-Motion-Network-for-Text-Independent-Speaker-Recognition-master/3LMNet_ms.py","file_url":"https://github.com/MindSpore-scientific/code-14/blob/HEAD/Three-Dimensional-Lip-Motion-Network-for-Text-Independent-Speaker-Recognition-master/3LMNet_ms.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"962937e4c560ebb3"}},{"code_sha256_prefix":"9c1b2df7fb7a5f6f","entry":"onehot2labels","repo":"MindSpore-scientific/code-14","repo_kind":"listed","path":"Three-Dimensional-Lip-Motion-Network-for-Text-Independent-Speaker-Recognition-master/3LMNet_ms.py","file_url":"https://github.com/MindSpore-scientific/code-14/blob/HEAD/Three-Dimensional-Lip-Motion-Network-for-Text-Independent-Speaker-Recognition-master/3LMNet_ms.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"9c1b2df7fb7a5f6f"}},{"code_sha256_prefix":"8de727761d866421","entry":"pendulum","repo":"MindSpore-scientific/code-14","repo_kind":"listed","path":"SciNet/utils.py","file_url":"https://github.com/MindSpore-scientific/code-14/blob/HEAD/SciNet/utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"8de727761d866421"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}