{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/longwanjuan-towards-systematic-measurement","title":"LongWanjuan: Towards Systematic Measurement for Long Text Quality","arxiv_id":"2402.13583","date":"2024-02-21","proceeding":null,"authors":["Kai Lv","Xiaoran Liu","Qipeng Guo","Hang Yan","Conghui He","Xipeng Qiu","Dahua Lin"],"abstract":"The quality of training data are crucial for enhancing the long-text capabilities of foundation models. Despite existing efforts to refine data quality through heuristic rules and evaluations based on data diversity and difficulty, there's a lack of systematic approaches specifically tailored for assessing long texts. Addressing this gap, our work systematically measures the quality of long texts by evaluating three fundamental linguistic dimensions: coherence, cohesion, and complexity. Drawing inspiration from the aforementioned three dimensions, we introduce a suite of metrics designed to evaluate the quality of long texts, encompassing both statistical and pre-trained language model-based ones. Leveraging these metrics, we present LongWanjuan, a bilingual dataset specifically tailored to enhance the training of language models for long-text tasks with over 160B tokens. In LongWanjuan, we categorize long texts into holistic, aggregated, and chaotic types, enabling a detailed analysis of long-text quality. Furthermore, we devise a data mixture recipe that strategically balances different types of long texts within LongWanjuan, leading to significant improvements in model performance on long-text tasks. The code and dataset are available at https://github.com/OpenLMLab/LongWanjuan.","url_abs":"https://arxiv.org/abs/2402.13583v2","url_pdf":"https://arxiv.org/pdf/2402.13583v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"longwanjuan-towards-systematic-measurement","repo_url":"https://github.com/openlmlab/longwanjuan","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"CC-BY-4.0"}}],"tasks":[{"task_slug":"diversity","task_name":"Diversity"},{"task_slug":"language-modelling","task_name":"Language Modelling"}],"methods":[],"datasets_introduced":[{"slug":"longwanjuan","name":"LongWanjuan","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2402.13583","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2402.13583"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/openlmlab/longwanjuan","reach":{"status":"ok","spdx":"CC-BY-4.0"}}],"summary":{"ran":5,"unverified":2},"by_repo_kind":{"official":{"samples":7,"ran":5,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":7,"samples":[{"code_sha256_prefix":"bfd015bc4320f853","entry":"get_conj_len","repo":"openlmlab/longwanjuan","repo_kind":"official","path":"cohesion/get_pron_conn_score.py","file_url":"https://github.com/openlmlab/longwanjuan/blob/HEAD/cohesion/get_pron_conn_score.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"CC-BY-4.0","inline_ok":false,"mcp_get_code":{"code_sha256":"bfd015bc4320f853"}},{"code_sha256_prefix":"43bbd1db03552247","entry":"get_joint_collate_fn","repo":"openlmlab/longwanjuan","repo_kind":"official","path":"cohesion/dmr/utils.py","file_url":"https://github.com/openlmlab/longwanjuan/blob/HEAD/cohesion/dmr/utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"CC-BY-4.0","inline_ok":false,"mcp_get_code":{"code_sha256":"43bbd1db03552247"}},{"code_sha256_prefix":"cde1c293853ab3dd","entry":"get_pronoun_len_cn","repo":"openlmlab/longwanjuan","repo_kind":"official","path":"cohesion/get_pron_conn_score.py","file_url":"https://github.com/openlmlab/longwanjuan/blob/HEAD/cohesion/get_pron_conn_score.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"CC-BY-4.0","inline_ok":false,"mcp_get_code":{"code_sha256":"cde1c293853ab3dd"}},{"code_sha256_prefix":"2cf3c1d9179a0137","entry":"read_discovery_connectives","repo":"openlmlab/longwanjuan","repo_kind":"official","path":"cohesion/dmr/utils.py","file_url":"https://github.com/openlmlab/longwanjuan/blob/HEAD/cohesion/dmr/utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"CC-BY-4.0","inline_ok":false,"mcp_get_code":{"code_sha256":"2cf3c1d9179a0137"}},{"code_sha256_prefix":"61324f8d787d32d0","entry":"start_with_conj","repo":"openlmlab/longwanjuan","repo_kind":"official","path":"cohesion/dmr/get_data.py","file_url":"https://github.com/openlmlab/longwanjuan/blob/HEAD/cohesion/dmr/get_data.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"CC-BY-4.0","inline_ok":false,"mcp_get_code":{"code_sha256":"61324f8d787d32d0"}},{"code_sha256_prefix":"849a335efbfaf67b","entry":"eval","repo":"openlmlab/longwanjuan","repo_kind":"official","path":"cohesion/dmr/dmv_train.py","file_url":"https://github.com/openlmlab/longwanjuan/blob/HEAD/cohesion/dmr/dmv_train.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"CC-BY-4.0","inline_ok":false,"mcp_get_code":{"code_sha256":"849a335efbfaf67b"}},{"code_sha256_prefix":"11f8272d9bd3e794","entry":"get_pronoun_num_en","repo":"openlmlab/longwanjuan","repo_kind":"official","path":"cohesion/get_pron_conn_score.py","file_url":"https://github.com/openlmlab/longwanjuan/blob/HEAD/cohesion/get_pron_conn_score.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"CC-BY-4.0","inline_ok":false,"mcp_get_code":{"code_sha256":"11f8272d9bd3e794"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}