{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/eq-bench-an-emotional-intelligence-benchmark","title":"EQ-Bench: An Emotional Intelligence Benchmark for Large Language Models","arxiv_id":"2312.06281","date":"2023-12-11","proceeding":null,"authors":["Samuel J. Paech"],"abstract":"We introduce EQ-Bench, a novel benchmark designed to evaluate aspects of emotional intelligence in Large Language Models (LLMs). We assess the ability of LLMs to understand complex emotions and social interactions by asking them to predict the intensity of emotional states of characters in a dialogue. The benchmark is able to discriminate effectively between a wide range of models. We find that EQ-Bench correlates strongly with comprehensive multi-domain benchmarks like MMLU (Hendrycks et al., 2020) (r=0.97), indicating that we may be capturing similar aspects of broad intelligence. Our benchmark produces highly repeatable results using a set of 60 English-language questions. We also provide open-source code for an automated benchmarking pipeline at https://github.com/EQ-bench/EQ-Bench and a leaderboard at https://eqbench.com","url_abs":"https://arxiv.org/abs/2312.06281v2","url_pdf":"https://arxiv.org/pdf/2312.06281v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"eq-bench-an-emotional-intelligence-benchmark","repo_url":"https://github.com/eq-bench/eq-bench","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"benchmarking","task_name":"Benchmarking"},{"task_slug":"emotional-intelligence","task_name":"Emotional Intelligence"},{"task_slug":"mmlu","task_name":"MMLU"}],"methods":[{"method_slug":"set","method_name":"SET"}],"datasets_introduced":[{"slug":"emotional-intelligence","name":"EQ-Bench","full_name":""}],"methods_introduced":[],"results":[{"leaderboard":"/sota/emotional-intelligence-on-emotional","task":"Emotional Intelligence","dataset":"EQ-Bench","model":"OpenAI gpt-4-0613","rank_in_archive_order":1,"of":24,"metrics":{"EQ-Bench Score":"62.52"},"uses_additional_data":false},{"leaderboard":"/sota/emotional-intelligence-on-emotional","task":"Emotional Intelligence","dataset":"EQ-Bench","model":"migtissera/SynthIA-70B-v1.5","rank_in_archive_order":2,"of":24,"metrics":{"EQ-Bench Score":"54.83"},"uses_additional_data":false},{"leaderboard":"/sota/emotional-intelligence-on-emotional","task":"Emotional Intelligence","dataset":"EQ-Bench","model":"OpenAI gpt-4-0314","rank_in_archive_order":3,"of":24,"metrics":{"EQ-Bench Score":"53.39"},"uses_additional_data":false},{"leaderboard":"/sota/emotional-intelligence-on-emotional","task":"Emotional Intelligence","dataset":"EQ-Bench","model":"Qwen/Qwen-72B-Chat","rank_in_archive_order":4,"of":24,"metrics":{"EQ-Bench Score":"52.44"},"uses_additional_data":false},{"leaderboard":"/sota/emotional-intelligence-on-emotional","task":"Emotional Intelligence","dataset":"EQ-Bench","model":"Anthropic Claude2","rank_in_archive_order":5,"of":24,"metrics":{"EQ-Bench Score":"52.14"},"uses_additional_data":false},{"leaderboard":"/sota/emotional-intelligence-on-emotional","task":"Emotional Intelligence","dataset":"EQ-Bench","model":"meta-llama/Llama-2-70b-chat-hf","rank_in_archive_order":6,"of":24,"metrics":{"EQ-Bench Score":"51.56"},"uses_additional_data":false},{"leaderboard":"/sota/emotional-intelligence-on-emotional","task":"Emotional Intelligence","dataset":"EQ-Bench","model":"01-ai/Yi-34B-Chat","rank_in_archive_order":7,"of":24,"metrics":{"EQ-Bench Score":"51.03"},"uses_additional_data":false},{"leaderboard":"/sota/emotional-intelligence-on-emotional","task":"Emotional Intelligence","dataset":"EQ-Bench","model":"OpenAI gpt-3.5-0613","rank_in_archive_order":8,"of":24,"metrics":{"EQ-Bench Score":"49.17"},"uses_additional_data":false},{"leaderboard":"/sota/emotional-intelligence-on-emotional","task":"Emotional Intelligence","dataset":"EQ-Bench","model":"OpenAI gpt-3.5-turbo-0301","rank_in_archive_order":9,"of":24,"metrics":{"EQ-Bench Score":"47.61"},"uses_additional_data":false},{"leaderboard":"/sota/emotional-intelligence-on-emotional","task":"Emotional Intelligence","dataset":"EQ-Bench","model":"Open-Orca/Mistral-7B-OpenOrca","rank_in_archive_order":10,"of":24,"metrics":{"EQ-Bench Score":"44.40"},"uses_additional_data":false},{"leaderboard":"/sota/emotional-intelligence-on-emotional","task":"Emotional Intelligence","dataset":"EQ-Bench","model":"Qwen/Qwen-14B-Chat","rank_in_archive_order":11,"of":24,"metrics":{"EQ-Bench Score":"43.76"},"uses_additional_data":false},{"leaderboard":"/sota/emotional-intelligence-on-emotional","task":"Emotional Intelligence","dataset":"EQ-Bench","model":"OpenAI text-davinci-003","rank_in_archive_order":12,"of":24,"metrics":{"EQ-Bench Score":"43.73"},"uses_additional_data":false},{"leaderboard":"/sota/emotional-intelligence-on-emotional","task":"Emotional Intelligence","dataset":"EQ-Bench","model":"Intel/neural-chat-7b-v3-1","rank_in_archive_order":13,"of":24,"metrics":{"EQ-Bench Score":"43.61"},"uses_additional_data":false},{"leaderboard":"/sota/emotional-intelligence-on-emotional","task":"Emotional Intelligence","dataset":"EQ-Bench","model":"OpenAI text-davinci-002","rank_in_archive_order":14,"of":24,"metrics":{"EQ-Bench Score":"39.44"},"uses_additional_data":false},{"leaderboard":"/sota/emotional-intelligence-on-emotional","task":"Emotional Intelligence","dataset":"EQ-Bench","model":"openchat/openchat 3.5","rank_in_archive_order":15,"of":24,"metrics":{"EQ-Bench Score":"37.08"},"uses_additional_data":false},{"leaderboard":"/sota/emotional-intelligence-on-emotional","task":"Emotional Intelligence","dataset":"EQ-Bench","model":"lmsys/vicuna-33b-v1.3","rank_in_archive_order":16,"of":24,"metrics":{"EQ-Bench Score":"36.52"},"uses_additional_data":false},{"leaderboard":"/sota/emotional-intelligence-on-emotional","task":"Emotional Intelligence","dataset":"EQ-Bench","model":"meta-llama/Llama-2-13b-chat-hf","rank_in_archive_order":17,"of":24,"metrics":{"EQ-Bench Score":"33.02"},"uses_additional_data":false},{"leaderboard":"/sota/emotional-intelligence-on-emotional","task":"Emotional Intelligence","dataset":"EQ-Bench","model":"lmsys/vicuna-13b-v1.1","rank_in_archive_order":18,"of":24,"metrics":{"EQ-Bench Score":" 32.85"},"uses_additional_data":false},{"leaderboard":"/sota/emotional-intelligence-on-emotional","task":"Emotional Intelligence","dataset":"EQ-Bench","model":"meta-llama/Llama-2-7b-chat-hf","rank_in_archive_order":19,"of":24,"metrics":{"EQ-Bench Score":"25.43"},"uses_additional_data":false},{"leaderboard":"/sota/emotional-intelligence-on-emotional","task":"Emotional Intelligence","dataset":"EQ-Bench","model":"Koala 13B","rank_in_archive_order":20,"of":24,"metrics":{"EQ-Bench Score":"24.92"},"uses_additional_data":false},{"leaderboard":"/sota/emotional-intelligence-on-emotional","task":"Emotional Intelligence","dataset":"EQ-Bench","model":"lmsys/vicuna-7b-v1.1","rank_in_archive_order":21,"of":24,"metrics":{"EQ-Bench Score":"22.24"},"uses_additional_data":false},{"leaderboard":"/sota/emotional-intelligence-on-emotional","task":"Emotional Intelligence","dataset":"EQ-Bench","model":"OpenAI text-davinci-001","rank_in_archive_order":22,"of":24,"metrics":{"EQ-Bench Score":"15.19"},"uses_additional_data":false},{"leaderboard":"/sota/emotional-intelligence-on-emotional","task":"Emotional Intelligence","dataset":"EQ-Bench","model":"OpenAI ADA","rank_in_archive_order":23,"of":24,"metrics":{"EQ-Bench Score":"2.25"},"uses_additional_data":false},{"leaderboard":"/sota/emotional-intelligence-on-emotional","task":"Emotional Intelligence","dataset":"EQ-Bench","model":"OpenAI ADA","rank_in_archive_order":24,"of":24,"metrics":{"EQ-Bench Score":"2.25"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2312.06281","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2312.06281"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/eq-bench/eq-bench","reach":null}],"summary":{"ran_violates":1},"by_repo_kind":{"official":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"0d1d1dc8419387fd","entry":"str2bool","repo":"eq-bench/eq-bench","repo_kind":"official","path":"eq-bench.py","file_url":"https://github.com/eq-bench/eq-bench/blob/HEAD/eq-bench.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"0d1d1dc8419387fd"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}