{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/wukong-100-million-large-scale-chinese-cross","title":"Wukong: A 100 Million Large-scale Chinese Cross-modal Pre-training Benchmark","arxiv_id":"2202.06767","date":"2022-02-14","proceeding":null,"authors":["Jiaxi Gu","Xiaojun Meng","Guansong Lu","Lu Hou","Minzhe Niu","Xiaodan Liang","Lewei Yao","Runhui Huang","Wei zhang","Xin Jiang","Chunjing Xu","Hang Xu"],"abstract":"Vision-Language Pre-training (VLP) models have shown remarkable performance on various downstream tasks. Their success heavily relies on the scale of pre-trained cross-modal datasets. However, the lack of large-scale datasets and benchmarks in Chinese hinders the development of Chinese VLP models and broader multilingual applications. In this work, we release a large-scale Chinese cross-modal dataset named Wukong, which contains 100 million Chinese image-text pairs collected from the web. Wukong aims to benchmark different multi-modal pre-training methods to facilitate the VLP research and community development. Furthermore, we release a group of models pre-trained with various image encoders (ViT-B/ViT-L/SwinT) and also apply advanced pre-training techniques into VLP such as locked-image text tuning, token-wise similarity in contrastive learning, and reduced-token interaction. Extensive experiments and a benchmarking of different downstream tasks including a new largest human-verified image-text test dataset are also provided. Experiments show that Wukong can serve as a promising Chinese pre-training dataset and benchmark for different cross-modal learning methods. For the zero-shot image classification task on 10 datasets, $Wukong_{ViT-L}$ achieves an average accuracy of 73.03%. For the image-text retrieval task, it achieves a mean recall of 71.6% on AIC-ICC which is 12.9% higher than WenLan 2.0. Also, our Wukong models are benchmarked on downstream tasks with other variants on multiple datasets, e.g., Flickr8K-CN, Flickr-30K-CN, COCO-CN, et al. More information can be referred to: https://wukong-dataset.github.io/wukong-dataset/.","url_abs":"https://arxiv.org/abs/2202.06767v4","url_pdf":"https://arxiv.org/pdf/2202.06767v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"wukong-100-million-large-scale-chinese-cross","repo_url":"https://github.com/0jason000/wukong","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"benchmarking","task_name":"Benchmarking"},{"task_slug":"contrastive-learning","task_name":"Contrastive Learning"},{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"image-retrieval","task_name":"Image Retrieval"},{"task_slug":"image-text-retrieval","task_name":"Image-text Retrieval"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"text-retrieval","task_name":"Text Retrieval"},{"task_slug":"zero-shot-image-classification","task_name":"Zero-Shot Image Classification"},{"task_slug":"zero-shot-image-retrieval","task_name":"Zero-shot Image Retrieval"},{"task_slug":"image-classification","task_name":"image-classification"}],"methods":[{"method_slug":"clip","method_name":"CLIP"},{"method_slug":"wenlan","method_name":"WenLan"}],"datasets_introduced":[{"slug":"wukong","name":"Wukong","full_name":""}],"methods_introduced":[],"results":[{"leaderboard":"/sota/image-retrieval-on-coco-cn","task":"Image Retrieval","dataset":"COCO-CN","model":"Wukong (ViT-L/14)","rank_in_archive_order":7,"of":9,"metrics":{"R@1":"74.0","R@10":"98.1","R@5":"94.4"},"uses_additional_data":false},{"leaderboard":"/sota/image-retrieval-on-coco-cn","task":"Image Retrieval","dataset":"COCO-CN","model":"Wukong (ViT-B/32)","rank_in_archive_order":8,"of":9,"metrics":{"R@1":"67.0","R@10":"96.7","R@5":"91.4"},"uses_additional_data":false},{"leaderboard":"/sota/image-retrieval-on-flickr30k-cn","task":"Image Retrieval","dataset":"Flickr30k-CN","model":"Wukong (ViT-L/14)","rank_in_archive_order":9,"of":11,"metrics":{"R@1":"77.4","R@10":"97.0","R@5":"94.5"},"uses_additional_data":false},{"leaderboard":"/sota/image-retrieval-on-flickr30k-cn","task":"Image Retrieval","dataset":"Flickr30k-CN","model":"Wukong (ViT-B/32)","rank_in_archive_order":10,"of":11,"metrics":{"R@1":"67.6","R@10":"94.2","R@5":"89.6"},"uses_additional_data":false},{"leaderboard":"/sota/image-retrieval-on-muge-retrieval","task":"Image Retrieval","dataset":"MUGE Retrieval","model":"Wukong (ViT-L/14)","rank_in_archive_order":6,"of":9,"metrics":{"Mean Recall":"72.1","R@1":"52.7","R@10":"85.6","R@5":"77.9"},"uses_additional_data":false},{"leaderboard":"/sota/image-retrieval-on-muge-retrieval","task":"Image Retrieval","dataset":"MUGE Retrieval","model":"Wukong (ViT-B/32)","rank_in_archive_order":9,"of":9,"metrics":{"Mean Recall":"61.2","R@1":"39.2","R@10":"77.4","R@5":"66.9"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2202.06767","atlas_url":"https://app.syntology.ai/?focus=2202.06767","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2202.06767"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/0jason000/wukong","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"ran":1,"unverified":1},"by_repo_kind":{"listed":{"samples":2,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"8693d1523eb76457","entry":"get_file_id","repo":"0jason000/wukong","repo_kind":"listed","path":"src/dataset/generate_dataset.py","file_url":"https://github.com/0jason000/wukong/blob/HEAD/src/dataset/generate_dataset.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"8693d1523eb76457"}},{"code_sha256_prefix":"e73c16378cc36858","entry":"process_one_data","repo":"0jason000/wukong","repo_kind":"listed","path":"src/dataset/generate_dataset.py","file_url":"https://github.com/0jason000/wukong/blob/HEAD/src/dataset/generate_dataset.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"e73c16378cc36858"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}