{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2603-17566","title":"KA2L: A Knowledge-Aware Active Learning Framework for LLMs","arxiv_id":"2603.17566","date":"2026-03-18","proceeding":null,"authors":["Haoxuan Yin","Chen Tang","Yangfan Wang","Lian Yan","Jingchi Jiang"],"abstract":"Fine-tuning large language models (LLMs) with high-quality knowledge has been shown to enhance their performance effectively. However, there is a paucity of research on the depth of domain-specific knowledge comprehension by LLMs and the application of targeted active learning to improve their expertise. To address this gap, we introduce the Knowledge-Aware Active Learning (KA2L) framework. This framework assesses LLMs' mastery of specific knowledge points to aid in constructing unanswerable or unknowable questions through latent space analysis. This active learning strategy enhances training efficiency by focusing on knowledge the model has yet to master, thereby minimizing redundancy in learning already acquired information. This study innovatively employs a knowledge distribution probing technique to examine the hidden states of specific Transformer layers and identify the distribution of known and unknown knowledge within the LLM. Additionally, a hidden-state decoding method is proposed to generate numerous unknown questions in natural language from the latent knowledge space. In our experiments, we selected nine open-source LLMs to validate the effectiveness of the proposed framework. Results indicate that KA2L not only significantly reduces 50% annotation and computation costs across two open-domain and one vertical-domain dataset but also achieves better performance, offering valuable insights into active learning strategies for LLMs. The code is available at https://github.com/greenjerry/KA2L.","url_abs":"https://arxiv.org/abs/2603.17566","url_pdf":"https://arxiv.org/pdf/2603.17566","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2603.17566","atlas_url":"https://app.syntology.ai/?focus=2603.17566","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2603.17566"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/greenjerry/KA2L","reach":null}],"summary":{"ran":2,"unverified":3},"by_repo_kind":{"found_in_text":{"samples":5,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":5,"samples":[{"code_sha256_prefix":"88b6ea37874fe00a","entry":"EmbeddingDataset","repo":"greenjerry/KA2L","repo_kind":"found_in_text","path":"KnowledgeDistributionProbe/model.py","file_url":"https://github.com/greenjerry/KA2L/blob/HEAD/KnowledgeDistributionProbe/model.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"88b6ea37874fe00a"}},{"code_sha256_prefix":"443e840cb2e87e2a","entry":"StoppingCriteriaSub","repo":"greenjerry/KA2L","repo_kind":"found_in_text","path":"KnowledgeDistributionProbe/model.py","file_url":"https://github.com/greenjerry/KA2L/blob/HEAD/KnowledgeDistributionProbe/model.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"443e840cb2e87e2a"}},{"code_sha256_prefix":"d046d838ba243351","entry":"ModelAndTokenizer","repo":"greenjerry/KA2L","repo_kind":"found_in_text","path":"KnowledgeDistributionProbe/model.py","file_url":"https://github.com/greenjerry/KA2L/blob/HEAD/KnowledgeDistributionProbe/model.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"d046d838ba243351"}},{"code_sha256_prefix":"53c1da4a27663812","entry":"generate_first_with_hidden_states","repo":"greenjerry/KA2L","repo_kind":"found_in_text","path":"KnowledgeDistributionProbe/model.py","file_url":"https://github.com/greenjerry/KA2L/blob/HEAD/KnowledgeDistributionProbe/model.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"53c1da4a27663812"}},{"code_sha256_prefix":"9bb44264c4b6cd78","entry":"save_ckpt","repo":"greenjerry/KA2L","repo_kind":"found_in_text","path":"KnowledgeDistributionProbe/model.py","file_url":"https://github.com/greenjerry/KA2L/blob/HEAD/KnowledgeDistributionProbe/model.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"9bb44264c4b6cd78"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.CL","source":"arxiv_2026.jsonl"},"syntology_extracted_results":null}