{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/text-sketch-image-compression-at-ultra-low","title":"Text + Sketch: Image Compression at Ultra Low Rates","arxiv_id":"2307.01944","date":"2023-07-04","proceeding":null,"authors":["Eric Lei","Yiğit Berkay Uslu","Hamed Hassani","Shirin Saeedi Bidokhti"],"abstract":"Recent advances in text-to-image generative models provide the ability to generate high-quality images from short text descriptions. These foundation models, when pre-trained on billion-scale datasets, are effective for various downstream tasks with little or no further training. A natural question to ask is how such models may be adapted for image compression. We investigate several techniques in which the pre-trained models can be directly used to implement compression schemes targeting novel low rate regimes. We show how text descriptions can be used in conjunction with side information to generate high-fidelity reconstructions that preserve both semantics and spatial structure of the original. We demonstrate that at very low bit-rates, our method can significantly improve upon learned compressors in terms of perceptual and semantic fidelity, despite no end-to-end training.","url_abs":"https://arxiv.org/abs/2307.01944v1","url_pdf":"https://arxiv.org/pdf/2307.01944v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"text-sketch-image-compression-at-ultra-low","repo_url":"https://github.com/leieric/text-sketch","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"image-compression","task_name":"Image Compression"},{"task_slug":"text-to-speech","task_name":"Text to Speech"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/text-to-speech-on-1","task":"Text to Speech","dataset":"^(#$!@#$)(()))******","model":"Reply","rank_in_archive_order":1,"of":1,"metrics":{"0-shot MRR":"شێوازی نووسینی خاڵبەندی بەدەنگ:  سێخاڵ: ... خاڵ: . نیشانەی سەرسووڕمان: ! نیشانەی پرسیار: ؟ فاریزە: ، خاڵبۆر: ؛ جووتخاڵ: : تەقەڵ: - کردنەوەی کەوانە: ( داخستنی کەوانە: ) کردنەوەی کەوانەی گۆشەدا: [ داخستنی کەوانەی گۆشەدار: [ کردنەوەی جووتکەوانە: « داخستنی جووتکەوانە: » هێڵی لار: /"},"uses_additional_data":true}],"syntology":{"syntology_url":"https://syntology.ai/paper/2307.01944","atlas_url":"https://app.syntology.ai/?focus=2307.01944","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2307.01944"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/leieric/Text-Sketch","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/leieric/text-sketch","reach":null}],"summary":{"ran_draft_wrong":1,"unverified":1},"by_repo_kind":{"official":{"samples":2,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":2,"samples":[{"code_sha256_prefix":"4d4553d8d2895e33","entry":"test_epoch","repo":"leieric/text-sketch","repo_kind":"official","path":"train_compressai.py","file_url":"https://github.com/leieric/text-sketch/blob/HEAD/train_compressai.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"4d4553d8d2895e33"}},{"code_sha256_prefix":"89f31bd65cefbe30","entry":"recon_rcc","repo":"leieric/Text-Sketch","repo_kind":"official","path":"eval_PIC.py","file_url":"https://github.com/leieric/Text-Sketch/blob/HEAD/eval_PIC.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"89f31bd65cefbe30"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}