{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2603-06636","title":"SmartBench: Evaluating LLMs in Smart Homes with Anomalous Device States and Behavioral Contexts","arxiv_id":"2603.06636","date":"2026-02-24","proceeding":null,"authors":["Qingsong Zou","Zhi Yan","Zhiyao Xu","Kuofeng Gao","Jingyu Xiao","Yong Jiang"],"abstract":"Due to the strong context-awareness capabilities demonstrated by large language models (LLMs), recent research has begun exploring their integration into smart home assistants to help users manage and adjust their living environments. While LLMs have been shown to effectively understand user needs and provide appropriate responses, most existing studies primarily focus on interpreting and executing user behaviors or instructions. However, a critical function of smart home assistants is the ability to detect when the home environment is in an anomalous state. This involves two key requirements: the LLM must accurately determine whether an anomalous condition is present, and provide either a clear explanation or actionable suggestions. To enhance the anomaly detection capabilities of next-generation LLM-based smart home assistants, we introduce SmartBench, which is the first smart home dataset designed for LLMs, containing both normal and anomalous device states as well as normal and anomalous device state transition contexts. We evaluate 13 mainstream LLMs on this benchmark. The experimental results show that most state-of-the-art models cannot achieve good anomaly detection performance. For example, Claude-Sonnet-4.5 achieves only 66.1% detection accuracy on context-independent anomaly categories, and performs even worse on context-dependent anomalies, with an accuracy of only 57.8%. More experimental results suggest that next-generation LLM-based smart home assistants are still far from being able to effectively detect and handle anomalous conditions in the smart home environment. Our dataset is publicly available at https://github.com/horizonsinzqs/SmartBench.","url_abs":"https://arxiv.org/abs/2603.06636","url_pdf":"https://arxiv.org/pdf/2603.06636","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2603.06636","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2603.06636"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/horizonsinzqs/SmartBench","reach":null}],"summary":{"ran_draft_wrong":5,"ran":3},"by_repo_kind":{"found_in_text":{"samples":8,"ran":8,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":8,"samples":[{"code_sha256_prefix":"e074228f42036d31","entry":"_walk_samples","repo":"horizonsinzqs/SmartBench","repo_kind":"found_in_text","path":"experiments/common/public_benchmark.py","file_url":"https://github.com/horizonsinzqs/SmartBench/blob/HEAD/experiments/common/public_benchmark.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"e074228f42036d31"}},{"code_sha256_prefix":"e65e775464399e47","entry":"build_public_views","repo":"horizonsinzqs/SmartBench","repo_kind":"found_in_text","path":"experiments/common/public_benchmark.py","file_url":"https://github.com/horizonsinzqs/SmartBench/blob/HEAD/experiments/common/public_benchmark.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"e65e775464399e47"}},{"code_sha256_prefix":"cf8c28b2c3471bf9","entry":"load_argus","repo":"horizonsinzqs/SmartBench","repo_kind":"found_in_text","path":"experiments/common/public_benchmark.py","file_url":"https://github.com/horizonsinzqs/SmartBench/blob/HEAD/experiments/common/public_benchmark.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"cf8c28b2c3471bf9"}},{"code_sha256_prefix":"206dda03d87eb3a3","entry":"load_context_independent","repo":"horizonsinzqs/SmartBench","repo_kind":"found_in_text","path":"experiments/common/public_benchmark.py","file_url":"https://github.com/horizonsinzqs/SmartBench/blob/HEAD/experiments/common/public_benchmark.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"206dda03d87eb3a3"}},{"code_sha256_prefix":"f10deba8c8561bee","entry":"load_json","repo":"horizonsinzqs/SmartBench","repo_kind":"found_in_text","path":"experiments/common/public_benchmark.py","file_url":"https://github.com/horizonsinzqs/SmartBench/blob/HEAD/experiments/common/public_benchmark.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"f10deba8c8561bee"}},{"code_sha256_prefix":"2fe0c2560b5e650d","entry":"load_tdcs","repo":"horizonsinzqs/SmartBench","repo_kind":"found_in_text","path":"experiments/common/public_benchmark.py","file_url":"https://github.com/horizonsinzqs/SmartBench/blob/HEAD/experiments/common/public_benchmark.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"2fe0c2560b5e650d"}},{"code_sha256_prefix":"6b59f009c0388bc0","entry":"sample_id","repo":"horizonsinzqs/SmartBench","repo_kind":"found_in_text","path":"experiments/common/public_benchmark.py","file_url":"https://github.com/horizonsinzqs/SmartBench/blob/HEAD/experiments/common/public_benchmark.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"6b59f009c0388bc0"}},{"code_sha256_prefix":"8002a68400850f80","entry":"write_json","repo":"horizonsinzqs/SmartBench","repo_kind":"found_in_text","path":"experiments/common/public_benchmark.py","file_url":"https://github.com/horizonsinzqs/SmartBench/blob/HEAD/experiments/common/public_benchmark.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"8002a68400850f80"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.LG","source":"arxiv_2026.jsonl"},"syntology_extracted_results":null}