Methods › Computer Vision › Vision and Language Pre-Trained Models › Pixel-BERT

Pixel-BERT

2 papers tagged archive 2025-07-28

Introduced by Zhicheng Huang et al. in Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

archive 2025-07-28 Description, source and code snippet are the archive's method entry.

Pixel-BERT is a pre-trained model trained to align image pixels with text. The end-to-end framework includes a CNN-based visual encoder and cross-modal transformers for visual and language embedding learning. This model has three parts: one fully convolutional neural network that takes pixels of an image as input, one word-level token embedding based on BERT, and a multimodal transformer for jointly learning visual and language embedding.

For language, it uses other pretraining works to use Masked Language Modeling (MLM) to predict masked tokens with surrounding text and images. For vision, it uses the random pixel sampling mechanism that makes up for the challenge of predicting pixel-level features. This mechanism is also suitable for solving overfitting issues and improving the robustness of visual features.

It applies Image-Text Matching (ITM) to classify whether an image and a sentence pair match for vision and language interaction.

Image captioning is required to understand language and visual semantics for cross-modality tasks like VQA. Region-based visual features extracted from object detection models like Faster RCNN are used for better performance in the newer version of the model.

PaperSource

Papers archive 2025-07-28

2 shown of 2, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.

Tasks archive 2025-07-28

20 shown of 26 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.

TaskPapers
Visual Question Answering (VQA)2
Visual Reasoning2
Conditional Image Generation1
Factual Visual Question Answering1
Fairness1
Image Captioning1
Image-text Retrieval1
Image-text matching1
Image-to-Text Retrieval1
Knowledge Graphs1
Language Modelling1
Multimodal Deep Learning1
Question Answering1
Representation Learning1
Retrieval1
Sentence1
Survey1
Text Matching1
Text Retrieval1
Vision-Language Navigation1

Usage over time archive 2025-07-28

Papers per year tagged with Pixel-BERT: 2020 to 2022, peak 1 1 0 2020: 1 paper 2020 2021: 0 papers 2021 2022: 1 paper 2022
Papers per year the archive tags with this method, by the paper's archive date (2 dated). Bars are counts, not a trend claim.

Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).

Categories archive 2025-07-28

Vision and Language Pre-Trained Models

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections