Papers › Convolutional Xformers for Vision

Convolutional Xformers for Vision

25 Jan 2022arXiv:2201.10271archive 2025-07-28

Pranav Jeevan, Amit Sethi

Vision transformers (ViTs) have found only limited practical use in processing images, in spite of their state-of-the-art accuracy on certain benchmarks. The reason for their limited use include their need for larger training datasets and more computational resources compared to convolutional neural networks (CNNs), owing to the quadratic complexity of their self-attention mechanism. We propose a linear attention-convolution hybrid architecture -- Convolutional X-formers for Vision (CXV) -- to overcome these limitations. We replace the quadratic attention with linear attention mechanisms, such as Performer, Nystr\"omformer, and Linear Transformer, to reduce its GPU usage. Inductive prior for image data is provided by convolutional sub-layers, thereby eliminating the need for class token and positional embeddings used by the ViTs. We also propose a new training method where we use two different optimizers during different phases of training and show that it improves the top-1 image classification accuracy across different architectures. CXV outperforms other architectures, token mixers (e.g. ConvMixer, FNet and MLP Mixer), transformer models (e.g. ViT, CCT, CvT and hybrid Xformers), and ResNets for image classification in scenarios with limited data and GPU resources (cores, RAM, power).

PaperPDFCode

Code

pranavphoenix/cxv officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Image Classificationimage-classification

1 archive task tag without a task page not shown.

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Image Classification CIFAR-10 Convolutional Performer for Vision (CPV) Percentage correct 94.46 #155 of 265 Archive leaderboard report
Image Classification CIFAR-100 Convolutional Linear Transformer for Vision (CLTV) Percentage correct 60.11 #198 of 211 Archive leaderboard report
Image Classification Tiny ImageNet Classification Convolutional Nystromformer for Vision (CNV) Validation Acc 49.56 #23 of 23 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Absolute Position EncodingsAdamAttentionAverage PoolingBPEBatch NormalizationCCTConvolutionCvTDense ConnectionsDepthwise ConvolutionDepthwise Separable ConvolutionDropoutFAVOR+Label SmoothingLayer NormalizationLinear LayerMulti-Head AttentionPerformerPointwise ConvolutionPosition-Wise Feed-Forward LayerResidual ConnectionSoftmaxTransformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections