{"url":"/dataset/financebench","name":"FinanceBench","full_name":null,"description_markdown":"Certainly! **FinanceBench** is a groundbreaking benchmark designed for evaluating the performance of large language models (LLMs) in the domain of financial question answering (QA). Here are the key details about FinanceBench:\r\n\r\n1. **What is FinanceBench?**\r\n   - **FinanceBench** is a **first-of-its-kind test suite** specifically tailored for assessing LLMs' capabilities in answering financial questions.\r\n   - It focuses on **open book financial QA** and comprises a collection of **10,231 questions** related to publicly traded companies.\r\n   - Each question comes with corresponding answers and evidence strings.\r\n\r\n2. **Why FinanceBench Matters:**\r\n   - The questions in FinanceBench are **ecologically valid**, covering a diverse set of scenarios.\r\n   - They are intentionally designed to be **clear-cut and straightforward**, serving as a minimum performance standard.\r\n   - FinanceBench aims to evaluate how well LLMs handle financial queries, especially those related to publicly traded companies.\r\n\r\n3. **Model Evaluation:**\r\n   - Researchers tested **16 state-of-the-art model configurations**, including **GPT-4-Turbo**, **Llama2**, and **Claude2**.\r\n   - The evaluation involved a sample of **150 cases** from FinanceBench, with **manual review of answers** (totaling **2,400**).\r\n   - Notably, **existing LLMs have limitations** for financial QA. For instance:\r\n     - **GPT-4-Turbo**, when used with a retrieval system, **incorrectly answered or refused to answer 81% of questions**.\r\n     - Augmentation techniques (such as longer context windows) improved performance but are **unrealistic for enterprise settings** due to increased latency and inability to handle larger financial documents.\r\n   - All examined models exhibit weaknesses, such as **hallucinations**, which limit their suitability for enterprise use.\r\n\r\n4. **Availability:**\r\n   - The **FinanceBench cases** are available **open-source** for further exploration and research.\r\n\r\n¹: Islam, P., Kannappan, A., Kiela, D., Qian, R., Scherrer, N., & Vidgen, B. (2023). FinanceBench: A New Benchmark for Financial Question Answering. arXiv preprint arXiv:2311.11944.\r\n²: [Link to the official paper](https://arxiv.org/abs/2311.11944)\r\n³: [Papers with Code - FinanceBench](https://paperswithcode.com/paper/financebench-a-new-benchmark-for-financial)\r\n\r\nSource: Conversation with Bing, 3/16/2024\r\n(1) Papers with Code - FinanceBench: A New Benchmark for Financial Question .... https://paperswithcode.com/paper/financebench-a-new-benchmark-for-financial.\r\n(2) FinanceBench: A New Benchmark for Financial Question Answering. https://arxiv.org/abs/2311.11944.\r\n(3) Papers with Code - Paper tables with annotated results for FinanceBench .... https://paperswithcode.com/paper/financebench-a-new-benchmark-for-financial/review/.","description_withheld":null,"homepage":"https://github.com/patronus-ai/financebench","introduced_date":"2023-11-20","introduced_date_note":null,"introduced_by":{"paper":"/paper/financebench-a-new-benchmark-for-financial","title":"FinanceBench: A New Benchmark for Financial Question Answering","first_author":"Pranab Islam","url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["FinanceBench"],"data_loaders":[],"num_papers_in_archive":27,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}