{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/pairaug-what-can-augmented-image-text-pairs","title":"PairAug: What Can Augmented Image-Text Pairs Do for Radiology?","arxiv_id":"2404.04960","date":"2024-04-07","proceeding":"CVPR 2024 1","authors":["Yutong Xie","Qi Chen","Sinuo Wang","Minh-Son To","Iris Lee","Ee Win Khoo","Kerolos Hendy","Daniel Koh","Yong Xia","Qi Wu"],"abstract":"Current vision-language pre-training (VLP) methodologies predominantly depend on paired image-text datasets, a resource that is challenging to acquire in radiology due to privacy considerations and labelling complexities. Data augmentation provides a practical solution to overcome the issue of data scarcity, however, most augmentation methods exhibit a limited focus, prioritising either image or text augmentation exclusively. Acknowledging this limitation, our objective is to devise a framework capable of concurrently augmenting medical image and text data. We design a Pairwise Augmentation (PairAug) approach that contains an Inter-patient Augmentation (InterAug) branch and an Intra-patient Augmentation (IntraAug) branch. Specifically, the InterAug branch of our approach generates radiology images using synthesised yet plausible reports derived from a Large Language Model (LLM). The generated pairs can be considered a collection of new patient cases since they are artificially created and may not exist in the original dataset. In contrast, the IntraAug branch uses newly generated reports to manipulate images. This process allows us to create new paired data for each individual with diverse medical conditions. Our extensive experiments on various downstream tasks covering medical image classification zero-shot and fine-tuning analysis demonstrate that our PairAug, concurrently expanding both image and text data, substantially outperforms image-/text-only expansion baselines and advanced medical VLP baselines. Our code is released at \\url{https://github.com/YtongXie/PairAug}.","url_abs":"https://arxiv.org/abs/2404.04960v1","url_pdf":"https://arxiv.org/pdf/2404.04960v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"pairaug-what-can-augmented-image-text-pairs","repo_url":"https://github.com/ytongxie/pairaug","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"data-augmentation","task_name":"Data Augmentation"},{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"large-language-model","task_name":"Large Language Model"},{"task_slug":"medical-image-classification","task_name":"Medical Image Classification"},{"task_slug":"text-augmentation","task_name":"Text Augmentation"},{"task_slug":"image-classification","task_name":"image-classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2404.04960","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2404.04960"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/YtongXie/PairAug","reach":{"status":"ok"}}],"summary":{"ran":3,"ran_draft_wrong":1,"unverified":5},"by_repo_kind":{"official":{"samples":9,"ran":4,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":9,"samples":[{"code_sha256_prefix":"d655f2b78b965b4a","entry":"chatgpt_completion","repo":"YtongXie/PairAug","repo_kind":"official","path":"InterAug_Step1.py","file_url":"https://github.com/YtongXie/PairAug/blob/HEAD/InterAug_Step1.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"d655f2b78b965b4a"}},{"code_sha256_prefix":"a042902caede448b","entry":"diffusion_step","repo":"YtongXie/PairAug","repo_kind":"official","path":"utils/ptp_utils.py","file_url":"https://github.com/YtongXie/PairAug/blob/HEAD/utils/ptp_utils.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"a042902caede448b"}},{"code_sha256_prefix":"e3c3bb87bbe9b4a7","entry":"latent2image","repo":"YtongXie/PairAug","repo_kind":"official","path":"utils/ptp_utils.py","file_url":"https://github.com/YtongXie/PairAug/blob/HEAD/utils/ptp_utils.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"e3c3bb87bbe9b4a7"}},{"code_sha256_prefix":"85d5eb123674597d","entry":"text_under_image","repo":"YtongXie/PairAug","repo_kind":"official","path":"utils/ptp_utils.py","file_url":"https://github.com/YtongXie/PairAug/blob/HEAD/utils/ptp_utils.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"85d5eb123674597d"}},{"code_sha256_prefix":"23ca7e055397296f","entry":"aggregate_attention","repo":"YtongXie/PairAug","repo_kind":"official","path":"IntraAug_Step2_T2I.py","file_url":"https://github.com/YtongXie/PairAug/blob/HEAD/IntraAug_Step2_T2I.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"23ca7e055397296f"}},{"code_sha256_prefix":"d167441d997c52ce","entry":"get_matrix","repo":"YtongXie/PairAug","repo_kind":"official","path":"utils/seq_aligner.py","file_url":"https://github.com/YtongXie/PairAug/blob/HEAD/utils/seq_aligner.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"d167441d997c52ce"}},{"code_sha256_prefix":"81db21bf982ca0fe","entry":"get_matrix","repo":"YtongXie/PairAug","repo_kind":"official","path":"utils/seq_aligner.py","file_url":"https://github.com/YtongXie/PairAug/blob/HEAD/utils/seq_aligner.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"81db21bf982ca0fe"}},{"code_sha256_prefix":"beb8f6a30678a4d9","entry":"get_traceback_matrix","repo":"YtongXie/PairAug","repo_kind":"official","path":"utils/seq_aligner.py","file_url":"https://github.com/YtongXie/PairAug/blob/HEAD/utils/seq_aligner.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"beb8f6a30678a4d9"}},{"code_sha256_prefix":"4511b96bc04402fb","entry":"num_tokens_from_string","repo":"YtongXie/PairAug","repo_kind":"official","path":"InterAug_Step1.py","file_url":"https://github.com/YtongXie/PairAug/blob/HEAD/InterAug_Step1.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"4511b96bc04402fb"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}