{"url":"/method/pixel-bert","slug":"pixel-bert","name":"Pixel-BERT","full_name":"Pixel-BERT","full_name_withheld":false,"description_markdown":"Pixel-BERT is a pre-trained model trained to align image pixels with text. The end-to-end framework includes a CNN-based visual encoder and cross-modal transformers for visual and language embedding learning.\r\nThis model has three parts: one fully convolutional neural network that takes pixels of an image as input, one word-level token embedding based on BERT, and a multimodal transformer for jointly learning visual and language embedding.\r\n\r\nFor language, it uses other pretraining works to use Masked Language Modeling (MLM) to predict masked tokens with surrounding text and images. For vision, it uses the random pixel sampling mechanism that makes up for the challenge of predicting pixel-level features. This mechanism is also suitable for solving overfitting issues and improving the robustness of visual features. \r\n\r\nIt applies Image-Text Matching (ITM) to classify whether an image and a sentence pair match for vision and language interaction. \r\n\r\nImage captioning is required to understand language and visual semantics for cross-modality tasks like VQA. Region-based visual features extracted from object detection models like Faster RCNN are used for better performance in the newer version of the model.","description_state":"present","introduced_year":null,"introduced_by":{"title":"Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers","paper":"/paper/pixel-bert-aligning-image-pixels-with-text-by","first_author":"Zhicheng Huang","n_authors":5,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/pixel-bert-aligning-image-pixels-with-text-by"},"source":{"url":"https://arxiv.org/abs/2004.00849v2","title":"Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Vision and Language Pre-Trained Models","url":"/methods/category/vision-and-language-pre-trained-models","pwc_aliases":[]}],"n_papers_tagged":2,"archive_num_papers":2,"papers_newest_first":[{"paper":null,"title":"A survey on knowledge-enhanced multimodal learning","date":"2022-11-19","arxiv_id":"2211.12328","n_code_links":0,"syntology":null},{"paper":"/paper/pixel-bert-aligning-image-pixels-with-text-by","title":"Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers","date":"2020-04-02","arxiv_id":"2004.00849","n_code_links":1,"syntology":null}],"papers_shown":2,"tasks":[{"task":"/task/visual-question-answering","name":"Visual Question Answering (VQA)","papers":2},{"task":"/task/visual-reasoning","name":"Visual Reasoning","papers":2},{"task":"/task/conditional-image-generation","name":"Conditional Image Generation","papers":1},{"task":"/task/factual-visual-question-answering","name":"Factual Visual Question Answering","papers":1},{"task":"/task/fairness","name":"Fairness","papers":1},{"task":"/task/image-captioning","name":"Image Captioning","papers":1},{"task":"/task/image-text-retrieval","name":"Image-text Retrieval","papers":1},{"task":"/task/image-text-matching","name":"Image-text matching","papers":1},{"task":"/task/image-to-text-retrieval","name":"Image-to-Text Retrieval","papers":1},{"task":"/task/knowledge-graphs","name":"Knowledge Graphs","papers":1},{"task":"/task/language-modelling","name":"Language Modelling","papers":1},{"task":"/task/multimodal-deep-learning","name":"Multimodal Deep Learning","papers":1},{"task":"/task/question-answering","name":"Question Answering","papers":1},{"task":"/task/representation-learning","name":"Representation Learning","papers":1},{"task":"/task/retrieval","name":"Retrieval","papers":1},{"task":"/task/sentence","name":"Sentence","papers":1},{"task":"/task/survey","name":"Survey","papers":1},{"task":"/task/text-matching","name":"Text Matching","papers":1},{"task":"/task/text-retrieval","name":"Text Retrieval","papers":1},{"task":"/task/vision-language-navigation","name":"Vision-Language Navigation","papers":1}],"tasks_shown":20,"n_tasks":26,"usage_by_year":[{"year":"2020","papers":1},{"year":"2022","papers":1}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/pixel-bert"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}