{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2602-00393","title":"Brazilian Portuguese Image Captioning with Transformers: A Study on Cross-Native-Translated Dataset","arxiv_id":"2602.00393","date":"2026-01-30","proceeding":null,"authors":["Gabriel Bromonschenkel","Alessandro L. Koerich","Thiago M. Paixão","Hilário Tomaz Alves de Oliveira"],"abstract":"Image captioning (IC) refers to the automatic generation of natural language descriptions for images, with applications ranging from social media content generation to assisting individuals with visual impairments. While most research has been focused on English-based models, low-resource languages such as Brazilian Portuguese face significant challenges due to the lack of specialized datasets and models. Several studies create datasets by automatically translating existing ones to mitigate resource scarcity. This work addresses this gap by proposing a cross-native-translated evaluation of Transformer-based vision and language models for Brazilian Portuguese IC. We use a version of Flickr30K comprised of captions manually created by native Brazilian Portuguese speakers and compare it to a version with captions automatically translated from English to Portuguese. The experiments include a cross-context approach, where models trained on one dataset are tested on the other to assess the translation impact. Additionally, we incorporate attention maps for model inference interpretation and use the CLIP-Score metric to evaluate the image-description alignment. Our findings show that Swin-DistilBERTimbau consistently outperforms other models, demonstrating strong generalization across datasets. ViTucano, a Brazilian Portuguese pre-trained VLM, surpasses larger multilingual models (GPT-4o, LLaMa 3.2 Vision) in traditional text-based evaluation metrics, while GPT-4 models achieve the highest CLIP-Score, highlighting improved image-text alignment. Attention analysis reveals systematic biases, including gender misclassification, object enumeration errors, and spatial inconsistencies. The datasets and the models generated and analyzed during the current study are available in: https://github.com/laicsiifes/transformer-caption-ptbr.","url_abs":"https://arxiv.org/abs/2602.00393","url_pdf":"https://arxiv.org/pdf/2602.00393","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2602.00393","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2602.00393"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/laicsiifes/transformer-caption-ptbr","reach":null}],"summary":{"unverified":4},"by_repo_kind":{"found_in_text":{"samples":4,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":4,"samples":[{"code_sha256_prefix":"9433e74de602c013","entry":"config_vars","repo":"laicsiifes/transformer-caption-ptbr","repo_kind":"found_in_text","path":"metrics_analysis/utils/config.py","file_url":"https://github.com/laicsiifes/transformer-caption-ptbr/blob/HEAD/metrics_analysis/utils/config.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"9433e74de602c013"}},{"code_sha256_prefix":"9303356715a06b3c","entry":"encode_image","repo":"laicsiifes/transformer-caption-ptbr","repo_kind":"found_in_text","path":"vlm_zero_shot/src/inference_sambanova_openai.py","file_url":"https://github.com/laicsiifes/transformer-caption-ptbr/blob/HEAD/vlm_zero_shot/src/inference_sambanova_openai.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"9303356715a06b3c"}},{"code_sha256_prefix":"cd5c340cf2b35248","entry":"join_datasets","repo":"laicsiifes/transformer-caption-ptbr","repo_kind":"found_in_text","path":"metrics_analysis/utils/data_processing.py","file_url":"https://github.com/laicsiifes/transformer-caption-ptbr/blob/HEAD/metrics_analysis/utils/data_processing.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"cd5c340cf2b35248"}},{"code_sha256_prefix":"ec88284c1848a453","entry":"select_incorrect_data","repo":"laicsiifes/transformer-caption-ptbr","repo_kind":"found_in_text","path":"metrics_analysis/utils/data_processing.py","file_url":"https://github.com/laicsiifes/transformer-caption-ptbr/blob/HEAD/metrics_analysis/utils/data_processing.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"ec88284c1848a453"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.LG","source":"arxiv_2026.jsonl"},"syntology_extracted_results":null}