{"url":"/method/k3m","slug":"k3m","name":"K3M","full_name":"K3M","full_name_withheld":false,"description_markdown":"**K3M** is a multi-modal pretraining method for e-commerce product data that introduces knowledge modality to correct the noise and supplement the missing of image and text modalities. The modal-encoding layer extracts the features of each modality. The modal-interaction layer is capable of effectively modeling the interaction of multiple modalities, where an initial-interactive feature fusion model is designed to maintain the independence of image modality and text modality, and a structure aggregation module is designed to fuse the information of image, text, and knowledge modalities. K3M is pre-trained with three pretraining tasks, including masked object modeling (MOM), masked language modeling (MLM), and link prediction modeling ([LPM](https://paperswithcode.com/method/local-prior-matching)).","description_state":"present","introduced_year":null,"introduced_by":{"title":"Knowledge Perceived Multi-modal Pretraining in E-commerce","paper":"/paper/knowledge-perceived-multi-modal-pretraining","first_author":"Yushan Zhu","n_authors":7,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/knowledge-perceived-multi-modal-pretraining"},"source":{"url":"https://arxiv.org/abs/2109.00895v1","title":"Knowledge Perceived Multi-modal Pretraining in E-commerce","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Language Model Pre-Training","url":"/methods/category/language-model-pre-training","pwc_aliases":[]}],"n_papers_tagged":1,"archive_num_papers":1,"papers_newest_first":[{"paper":"/paper/knowledge-perceived-multi-modal-pretraining","title":"Knowledge Perceived Multi-modal Pretraining in E-commerce","date":"2021-08-20","arxiv_id":"2109.00895","n_code_links":1,"syntology":{"ran":1,"of":1,"unverified":0,"pointer_only":0}}],"papers_shown":1,"tasks":[{"task":"/task/language-modeling","name":"Language Modeling","papers":1},{"task":"/task/language-modelling","name":"Language Modelling","papers":1},{"task":"/task/link-prediction","name":"Link Prediction","papers":1},{"task":"/task/masked-language-modeling","name":"Masked Language Modeling","papers":1}],"tasks_shown":4,"n_tasks":4,"usage_by_year":[{"year":"2021","papers":1}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/k3m"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}