{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/image-textualization-an-automatic-framework","title":"Image Textualization: An Automatic Framework for Creating Accurate and Detailed Image Descriptions","arxiv_id":"2406.07502","date":"2024-06-11","proceeding":null,"authors":["Renjie Pi","Jianshu Zhang","Jipeng Zhang","Rui Pan","Zhekai Chen","Tong Zhang"],"abstract":"Image description datasets play a crucial role in the advancement of various applications such as image understanding, text-to-image generation, and text-image retrieval. Currently, image description datasets primarily originate from two sources. One source is the scraping of image-text pairs from the web. Despite their abundance, these descriptions are often of low quality and noisy. Another is through human labeling. Datasets such as COCO are generally very short and lack details. Although detailed image descriptions can be annotated by humans, the high annotation cost limits the feasibility. These limitations underscore the need for more efficient and scalable methods to generate accurate and detailed image descriptions. In this paper, we propose an innovative framework termed Image Textualization (IT), which automatically produces high-quality image descriptions by leveraging existing multi-modal large language models (MLLMs) and multiple vision expert models in a collaborative manner, which maximally convert the visual information into text. To address the current lack of benchmarks for detailed descriptions, we propose several benchmarks for comprehensive evaluation, which verifies the quality of image descriptions created by our framework. Furthermore, we show that LLaVA-7B, benefiting from training on IT-curated descriptions, acquire improved capability to generate richer image descriptions, substantially increasing the length and detail of their output with less hallucination.","url_abs":"https://arxiv.org/abs/2406.07502v1","url_pdf":"https://arxiv.org/pdf/2406.07502v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"image-textualization-an-automatic-framework","repo_url":"https://github.com/sterzhang/image-textualization","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"hallucination","task_name":"Hallucination"},{"task_slug":null,"task_name":"Image Description"},{"task_slug":"image-generation","task_name":"Image Generation"},{"task_slug":"image-retrieval","task_name":"Image Retrieval"},{"task_slug":"text-to-image-generation-1","task_name":"Text to Image Generation"},{"task_slug":"text-to-image-generation","task_name":"Text-to-Image Generation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2406.07502","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2406.07502"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/sterzhang/image-textualization","reach":{"status":"ok"}}],"summary":{"ran":7,"unverified":2},"by_repo_kind":{"official":{"samples":9,"ran":7,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":9,"samples":[{"code_sha256_prefix":"f8c47ec45e9f630c","entry":"binary_mask_to_polygon","repo":"sterzhang/image-textualization","repo_kind":"official","path":"fg_annotation/mask_depth.py","file_url":"https://github.com/sterzhang/image-textualization/blob/HEAD/fg_annotation/mask_depth.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"f8c47ec45e9f630c"}},{"code_sha256_prefix":"0e4d90447aeef0e5","entry":"chars","repo":"sterzhang/image-textualization","repo_kind":"official","path":"benchmark/Linguistic/readability.py","file_url":"https://github.com/sterzhang/image-textualization/blob/HEAD/benchmark/Linguistic/readability.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"0e4d90447aeef0e5"}},{"code_sha256_prefix":"f0f26534728a1b2e","entry":"load_cache","repo":"sterzhang/image-textualization","repo_kind":"official","path":"benchmark/Count/cnt.py","file_url":"https://github.com/sterzhang/image-textualization/blob/HEAD/benchmark/Count/cnt.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"f0f26534728a1b2e"}},{"code_sha256_prefix":"db73ba37bbba8a74","entry":"polygons_to_binary_mask","repo":"sterzhang/image-textualization","repo_kind":"official","path":"fg_annotation/mask_depth.py","file_url":"https://github.com/sterzhang/image-textualization/blob/HEAD/fg_annotation/mask_depth.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"db73ba37bbba8a74"}},{"code_sha256_prefix":"f4b4c6887ee82c05","entry":"segment","repo":"sterzhang/image-textualization","repo_kind":"official","path":"fg_annotation/mask_depth.py","file_url":"https://github.com/sterzhang/image-textualization/blob/HEAD/fg_annotation/mask_depth.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"f4b4c6887ee82c05"}},{"code_sha256_prefix":"c95c89681f57eed9","entry":"sentences","repo":"sterzhang/image-textualization","repo_kind":"official","path":"benchmark/Linguistic/readability.py","file_url":"https://github.com/sterzhang/image-textualization/blob/HEAD/benchmark/Linguistic/readability.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"c95c89681f57eed9"}},{"code_sha256_prefix":"622fd1f9e0909e9d","entry":"words","repo":"sterzhang/image-textualization","repo_kind":"official","path":"benchmark/Linguistic/readability.py","file_url":"https://github.com/sterzhang/image-textualization/blob/HEAD/benchmark/Linguistic/readability.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"622fd1f9e0909e9d"}},{"code_sha256_prefix":"be67c2fa0015e178","entry":"count_pos","repo":"sterzhang/image-textualization","repo_kind":"official","path":"benchmark/Count/cnt.py","file_url":"https://github.com/sterzhang/image-textualization/blob/HEAD/benchmark/Count/cnt.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"be67c2fa0015e178"}},{"code_sha256_prefix":"0c55cb0492a5714a","entry":"query_ChatGPT","repo":"sterzhang/image-textualization","repo_kind":"official","path":"extract/extract_fr_desc-gpt.py","file_url":"https://github.com/sterzhang/image-textualization/blob/HEAD/extract/extract_fr_desc-gpt.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"0c55cb0492a5714a"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}