{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-to-evaluate-image-captioning","title":"Learning to Evaluate Image Captioning","arxiv_id":"1806.06422","date":"2018-06-17","proceeding":"CVPR 2018 6","authors":["Yin Cui","Guandao Yang","Andreas Veit","Xun Huang","Serge Belongie"],"abstract":"Evaluation metrics for image captioning face two challenges. Firstly,\ncommonly used metrics such as CIDEr, METEOR, ROUGE and BLEU often do not\ncorrelate well with human judgments. Secondly, each metric has well known blind\nspots to pathological caption constructions, and rule-based metrics lack\nprovisions to repair such blind spots once identified. For example, the newly\nproposed SPICE correlates well with human judgments, but fails to capture the\nsyntactic structure of a sentence. To address these two challenges, we propose\na novel learning based discriminative evaluation metric that is directly\ntrained to distinguish between human and machine-generated captions. In\naddition, we further propose a data augmentation scheme to explicitly\nincorporate pathological transformations as negative examples during training.\nThe proposed metric is evaluated with three kinds of robustness tests and its\ncorrelation with human judgments. Extensive experiments show that the proposed\ndata augmentation scheme not only makes our metric more robust toward several\npathological transformations, but also improves its correlation with human\njudgments. Our metric outperforms other metrics on both caption level human\ncorrelation in Flickr 8k and system level human correlation in COCO. The\nproposed approach could be served as a learning based evaluation metric that is\ncomplementary to existing rule-based metrics.","url_abs":"http://arxiv.org/abs/1806.06422v1","url_pdf":"http://arxiv.org/pdf/1806.06422v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-to-evaluate-image-captioning","repo_url":"https://github.com/richardaecn/cvpr18-caption-eval","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":null,"task_name":"8k"},{"task_slug":"data-augmentation","task_name":"Data Augmentation"},{"task_slug":"image-captioning","task_name":"Image Captioning"},{"task_slug":"sentence","task_name":"Sentence"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1806.06422","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1806.06422"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/richardaecn/cvpr18-caption-eval","reach":null}],"summary":{"unverified":1},"by_repo_kind":{"official":{"samples":1,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"125c42eb3ba458ac","entry":"data_loader","repo":"richardaecn/cvpr18-caption-eval","repo_kind":"official","path":"score.py","file_url":"https://github.com/richardaecn/cvpr18-caption-eval/blob/HEAD/score.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"125c42eb3ba458ac"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}