{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/qg-sms-enhancing-test-item-analysis-via","title":"QG-SMS: Enhancing Test Item Analysis via Student Modeling and Simulation","arxiv_id":"2503.05888","date":"2025-03-07","proceeding":null,"authors":["Bang Nguyen","Tingting Du","Mengxia Yu","Lawrence Angrave","Meng Jiang"],"abstract":"While the Question Generation (QG) task has been increasingly adopted in educational assessments, its evaluation remains limited by approaches that lack a clear connection to the educational values of test items. In this work, we introduce test item analysis, a method frequently used by educators to assess test question quality, into QG evaluation. Specifically, we construct pairs of candidate questions that differ in quality across dimensions such as topic coverage, item difficulty, item discrimination, and distractor efficiency. We then examine whether existing QG evaluation approaches can effectively distinguish these differences. Our findings reveal significant shortcomings in these approaches with respect to accurately assessing test item quality in relation to student performance. To address this gap, we propose a novel QG evaluation framework, QG-SMS, which leverages Large Language Model for Student Modeling and Simulation to perform test item analysis. As demonstrated in our extensive experiments and human evaluation study, the additional perspectives introduced by the simulated student profiles lead to a more effective and robust assessment of test items.","url_abs":"https://arxiv.org/abs/2503.05888v2","url_pdf":"https://arxiv.org/pdf/2503.05888v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"large-language-model","task_name":"Large Language Model"},{"task_slug":"question-generation","task_name":"Question Generation"},{"task_slug":"question-generation","task_name":"Question-Generation"},{"task_slug":"topic-coverage","task_name":"Topic coverage"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2503.05888","atlas_url":"https://app.syntology.ai/?focus=2503.05888","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2503.05888"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/bnguyen5/qg-sms","reach":null}],"summary":{"ran_draft_wrong":1,"unverified":3},"by_repo_kind":{"found_in_text":{"samples":3,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":4,"samples":[{"code_sha256_prefix":"70d139854ce1775e","entry":"prompt_to_chatml","repo":null,"repo_kind":null,"path":null,"file_url":null,"link_basis":"identical_code_first_harvested_elsewhere","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":null,"inline_ok":false,"mcp_get_code":{"code_sha256":"70d139854ce1775e"}},{"code_sha256_prefix":"50e8b5f3decb3ef3","entry":"Complete","repo":"bnguyen5/qg-sms","repo_kind":"found_in_text","path":"qg-sms/evaluate.py","file_url":"https://github.com/bnguyen5/qg-sms/blob/HEAD/qg-sms/evaluate.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"50e8b5f3decb3ef3"}},{"code_sha256_prefix":"03954fe7d0e72ef2","entry":"openai_completion","repo":"bnguyen5/qg-sms","repo_kind":"found_in_text","path":"qg-sms/evaluate.py","file_url":"https://github.com/bnguyen5/qg-sms/blob/HEAD/qg-sms/evaluate.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"03954fe7d0e72ef2"}},{"code_sha256_prefix":"df5e032a7f9614a1","entry":"simulate","repo":"bnguyen5/qg-sms","repo_kind":"found_in_text","path":"qg-sms/evaluate.py","file_url":"https://github.com/bnguyen5/qg-sms/blob/HEAD/qg-sms/evaluate.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"df5e032a7f9614a1"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}