{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/fineval-a-chinese-financial-domain-knowledge","title":"FinEval: A Chinese Financial Domain Knowledge Evaluation Benchmark for Large Language Models","arxiv_id":"2308.09975","date":"2023-08-19","proceeding":null,"authors":["Xin Guo","Haotian Xia","Zhaowei Liu","Hanyang Cao","Zhi Yang","Zhiqiang Liu","Sizhe Wang","Jinyi Niu","Chuqi Wang","Yanhui Wang","Xiaolong Liang","Xiaoming Huang","Bing Zhu","Zhongyu Wei","Yun Chen","Weining Shen","Liwen Zhang"],"abstract":"Large language models have demonstrated outstanding performance in various natural language processing tasks, but their security capabilities in the financial domain have not been explored, and their performance on complex tasks like financial agent remains unknown. This paper presents FinEval, a benchmark designed to evaluate LLMs' financial domain knowledge and practical abilities. The dataset contains 8,351 questions categorized into four different key areas: Financial Academic Knowledge, Financial Industry Knowledge, Financial Security Knowledge, and Financial Agent. Financial Academic Knowledge comprises 4,661 multiple-choice questions spanning 34 subjects such as finance and economics. Financial Industry Knowledge contains 1,434 questions covering practical scenarios like investment research. Financial Security Knowledge assesses models through 1,640 questions on topics like application security and cryptography. Financial Agent evaluates tool usage and complex reasoning with 616 questions. FinEval has multiple evaluation settings, including zero-shot, five-shot with chain-of-thought, and assesses model performance using objective and subjective criteria. Our results show that Claude 3.5-Sonnet achieves the highest weighted average score of 72.9 across all financial domain categories under zero-shot setting. Our work provides a comprehensive benchmark closely aligned with Chinese financial domain.","url_abs":"https://arxiv.org/abs/2308.09975v2","url_pdf":"https://arxiv.org/pdf/2308.09975v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"fineval-a-chinese-financial-domain-knowledge","repo_url":"https://github.com/sufe-aiflm-lab/fineval","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"multiple-choice","task_name":"Multiple-choice"}],"methods":[{"method_slug":"absolute-position-encodings","method_name":"Absolute Position Encodings"},{"method_slug":"adam","method_name":"Adam"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"bpe","method_name":"BPE"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"gpt-4","method_name":"GPT-4"},{"method_slug":"label-smoothing","method_name":"Label Smoothing"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"position-wise-feed-forward-layer","method_name":"Position-Wise Feed-Forward Layer"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"transformer","method_name":"Transformer"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2308.09975","atlas_url":"https://app.syntology.ai/?focus=2308.09975","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2308.09975"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/sufe-aiflm-lab/fineval","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"ran":5},"by_repo_kind":{"official":{"samples":5,"ran":5,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"4074f3ebedccb3b5","entry":"extract_cotanswer","repo":"sufe-aiflm-lab/fineval","repo_kind":"official","path":"code/opensource_eval/2industry_eval/utils.py","file_url":"https://github.com/sufe-aiflm-lab/fineval/blob/HEAD/code/opensource_eval/2industry_eval/utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"4074f3ebedccb3b5"}},{"code_sha256_prefix":"949511fc0783abf6","entry":"extract_cotsuggestion","repo":"sufe-aiflm-lab/fineval","repo_kind":"official","path":"code/opensource_eval/2industry_eval/utils.py","file_url":"https://github.com/sufe-aiflm-lab/fineval/blob/HEAD/code/opensource_eval/2industry_eval/utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"949511fc0783abf6"}},{"code_sha256_prefix":"30aecd96194fddce","entry":"extract_questions_and_text","repo":"sufe-aiflm-lab/fineval","repo_kind":"official","path":"code/opensource_eval/34security+agenteval/utils.py","file_url":"https://github.com/sufe-aiflm-lab/fineval/blob/HEAD/code/opensource_eval/34security%2Bagenteval/utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"30aecd96194fddce"}},{"code_sha256_prefix":"f97998403d38cffc","entry":"load_json","repo":"sufe-aiflm-lab/fineval","repo_kind":"official","path":"code/opensource_eval/2industry_eval/utils.py","file_url":"https://github.com/sufe-aiflm-lab/fineval/blob/HEAD/code/opensource_eval/2industry_eval/utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"f97998403d38cffc"}},{"code_sha256_prefix":"48952da28c749fea","entry":"recognize_and_convert","repo":"sufe-aiflm-lab/fineval","repo_kind":"official","path":"code/opensource_eval/34security+agenteval/evaluate.py","file_url":"https://github.com/sufe-aiflm-lab/fineval/blob/HEAD/code/opensource_eval/34security%2Bagenteval/evaluate.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"48952da28c749fea"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}