{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2509-03897","title":"SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation","arxiv_id":"2509.03897","date":"2025-09-04","proceeding":"EMNLP","authors":["Xiaofu Chen","Israfel Salazar","Yova Kementchedjhieva"],"abstract":"As interest grows in generating long, detailed image captions, standard evaluation metrics become increasingly unreliable. N-gram-based metrics though efficient, fail to capture semantic correctness. Representational Similarity (RS) metrics, designed to address this, initially saw limited use due to high computational costs, while today, despite advances in hardware, they remain unpopular due to low correlation to human judgments. Meanwhile, metrics based on large language models (LLMs) show strong correlation with human judgments, but remain too expensive for iterative use during model development. We introduce SPECS (Specificity-Enhanced CLIPScore), a reference-free RS metric tailored to long image captioning. SPECS modifies CLIP with a new objective that emphasizes specificity: rewarding correct details and penalizing incorrect ones. We show that SPECS matches the performance of open-source LLM-based metrics in correlation to human judgments, while being far more efficient. This makes it a practical alternative for iterative checkpoint evaluation during image captioning model development.Our code can be found at https://github.com/mbzuai-nlp/SPECS.","url_abs":"https://arxiv.org/abs/2509.03897","url_pdf":"https://arxiv.org/pdf/2509.03897","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2509.03897","atlas_url":"https://app.syntology.ai/?focus=2509.03897","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2509.03897"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/mbzuai-nlp/SPECS","reach":null}],"summary":{"ran_draft_wrong":3,"unverified":1},"by_repo_kind":{"found_in_text":{"samples":4,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"11ea502213a91170","entry":"clip_loss","repo":"mbzuai-nlp/SPECS","repo_kind":"found_in_text","path":"model/model_spec.py","file_url":"https://github.com/mbzuai-nlp/SPECS/blob/HEAD/model/model_spec.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"11ea502213a91170"}},{"code_sha256_prefix":"24fc39254981d28c","entry":"hinge_loss","repo":"mbzuai-nlp/SPECS","repo_kind":"found_in_text","path":"model/model_spec.py","file_url":"https://github.com/mbzuai-nlp/SPECS/blob/HEAD/model/model_spec.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"24fc39254981d28c"}},{"code_sha256_prefix":"70d2400aa0f80709","entry":"increment_loss","repo":"mbzuai-nlp/SPECS","repo_kind":"found_in_text","path":"model/model_spec.py","file_url":"https://github.com/mbzuai-nlp/SPECS/blob/HEAD/model/model_spec.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"70d2400aa0f80709"}},{"code_sha256_prefix":"a6e7c621c17939bf","entry":"DetailLossCalculator","repo":"mbzuai-nlp/SPECS","repo_kind":"found_in_text","path":"model/model_spec.py","file_url":"https://github.com/mbzuai-nlp/SPECS/blob/HEAD/model/model_spec.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"a6e7c621c17939bf"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.CV","source":"arxiv_api"},"syntology_extracted_results":null}