{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2604-04323","title":"How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings","arxiv_id":"2604.04323","date":"2026-04-06","proceeding":null,"authors":["Yujian Liu","Jiabao Ji","Li An","Tommi Jaakkola","Yang Zhang","Shiyu Chang"],"abstract":"Agent skills, which are reusable, domain-specific knowledge artifacts, have become a popular mechanism for extending LLM-based agents, yet formally benchmarking skill usage performance remains scarce. Existing skill benchmarking efforts focus on overly idealized conditions, where LLMs are directly provided with hand-crafted, narrowly-tailored task-specific skills for each task, whereas in many realistic settings, the LLM agent may have to search for and select relevant skills on its own, and even the closest matching skills may not be well-tailored for the task. In this paper, we conduct the first comprehensive study of skill utility under progressively challenging realistic settings, where agents must retrieve skills from a large collection of 34k real-world skills and may not have access to any hand-curated skills. Our findings reveal that the benefits of skills are fragile: performance gains degrade consistently as settings become more realistic, with pass rates approaching no-skill baselines in the most challenging scenarios. To narrow this gap, we study skill refinement strategies, including query-specific and query-agnostic approaches, and we show that query-specific refinement substantially recovers lost performance when the initial skills are of reasonable relevance and quality. We further demonstrate the generality of retrieval and refinement on Terminal-Bench 2.0, where they improve the pass rate of Claude Opus 4.6 from 57.7% to 65.5%. Our results, consistent across multiple models, highlight both the promise and the current limitations of skills for LLM-based agents. Our code is available at https://github.com/UCSB-NLP-Chang/Skill-Usage.","url_abs":"https://arxiv.org/abs/2604.04323","url_pdf":"https://arxiv.org/pdf/2604.04323","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2604.04323","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2604.04323"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/UCSB-NLP-Chang/Skill-Usage","reach":null}],"summary":{"unverified":8},"by_repo_kind":{"found_in_text":{"samples":8,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":8,"samples":[{"code_sha256_prefix":"0a76eae7fc90cc9c","entry":"build_run","repo":"UCSB-NLP-Chang/Skill-Usage","repo_kind":"found_in_text","path":".claude/skills/skill-creator/eval-viewer/generate_review.py","file_url":"https://github.com/UCSB-NLP-Chang/Skill-Usage/blob/HEAD/.claude/skills/skill-creator/eval-viewer/generate_review.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"0a76eae7fc90cc9c"}},{"code_sha256_prefix":"7f2d229dbd83f844","entry":"extract_skill_md_content","repo":"UCSB-NLP-Chang/Skill-Usage","repo_kind":"found_in_text","path":"search_server/index_builder.py","file_url":"https://github.com/UCSB-NLP-Chang/Skill-Usage/blob/HEAD/search_server/index_builder.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"7f2d229dbd83f844"}},{"code_sha256_prefix":"3eece029ae84aeed","entry":"extract_skill_md_snippet","repo":"UCSB-NLP-Chang/Skill-Usage","repo_kind":"found_in_text","path":"search_server/index_builder.py","file_url":"https://github.com/UCSB-NLP-Chang/Skill-Usage/blob/HEAD/search_server/index_builder.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"3eece029ae84aeed"}},{"code_sha256_prefix":"a7448f082f36f9b2","entry":"find_runs","repo":"UCSB-NLP-Chang/Skill-Usage","repo_kind":"found_in_text","path":".claude/skills/skill-creator/eval-viewer/generate_review.py","file_url":"https://github.com/UCSB-NLP-Chang/Skill-Usage/blob/HEAD/.claude/skills/skill-creator/eval-viewer/generate_review.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"a7448f082f36f9b2"}},{"code_sha256_prefix":"355cc1877144cdab","entry":"forward","repo":"UCSB-NLP-Chang/Skill-Usage","repo_kind":"found_in_text","path":"terminal-bench-2/model-extraction-relu-logits/environment/forward.py","file_url":"https://github.com/UCSB-NLP-Chang/Skill-Usage/blob/HEAD/terminal-bench-2/model-extraction-relu-logits/environment/forward.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"355cc1877144cdab"}},{"code_sha256_prefix":"5790d7a1b8e143d1","entry":"get_mime_type","repo":"UCSB-NLP-Chang/Skill-Usage","repo_kind":"found_in_text","path":".claude/skills/skill-creator/eval-viewer/generate_review.py","file_url":"https://github.com/UCSB-NLP-Chang/Skill-Usage/blob/HEAD/.claude/skills/skill-creator/eval-viewer/generate_review.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"5790d7a1b8e143d1"}},{"code_sha256_prefix":"8815dd447edad001","entry":"load_records","repo":"UCSB-NLP-Chang/Skill-Usage","repo_kind":"found_in_text","path":"search_server/index_builder.py","file_url":"https://github.com/UCSB-NLP-Chang/Skill-Usage/blob/HEAD/search_server/index_builder.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"8815dd447edad001"}},{"code_sha256_prefix":"68ee2675b183b56f","entry":"relu","repo":"UCSB-NLP-Chang/Skill-Usage","repo_kind":"found_in_text","path":"terminal-bench-2/model-extraction-relu-logits/environment/forward.py","file_url":"https://github.com/UCSB-NLP-Chang/Skill-Usage/blob/HEAD/terminal-bench-2/model-extraction-relu-logits/environment/forward.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"68ee2675b183b56f"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.CL","source":"arxiv_2026.jsonl"},"syntology_extracted_results":{"kind":"leaderboard_placements","source":"Syntology's leaderboard-shaped extractor over the paper's own arXiv-HTML tables: a model pointed at a cell, the number was read from that cell and checked against the board's metric, dataset, split and scale, and an independent check accepted the entry; not reviewed by the paper's authors or the archive's editors","extractor_model":"global.anthropic.claude-sonnet-4-5-20250929-v1:0","verifier_model":null,"prompt_sha":"fa63d4bb9d755694","coverage":{"sentence":"Syntology has checked 6,264 of the 9,581 papers on this site that are newer than the archive; results from the others appear after they are checked.","papers_newer_than_archive":9581,"papers_checked":6264},"entries":[],"not_placed":{"boards":0,"rejected_by_independent_check":0,"refused_by_a_rule":0,"check_did_not_answer":0,"proposed_without_a_cell":0,"declined_by_site":0}}}