Papers › HyenaPixel: Global Image Context with Convolutions

HyenaPixel: Global Image Context with Convolutions

29 Feb 2024arXiv:2402.19305archive 2025-07-28

Julian Spravil, Sebastian Houben, Sven Behnke

In computer vision, a larger effective receptive field (ERF) is associated with better performance. While attention natively supports global context, its quadratic complexity limits its applicability to tasks that benefit from high-resolution input. In this work, we extend Hyena, a convolution-based attention replacement, from causal sequences to bidirectional data and two-dimensional image space. We scale Hyena's convolution kernels beyond the feature map size, up to 191×191, to maximize ERF while maintaining sub-quadratic complexity in the number of pixels. We integrate our two-dimensional Hyena, HyenaPixel, and bidirectional Hyena into the MetaFormer framework. For image categorization, HyenaPixel and bidirectional Hyena achieve a competitive ImageNet-1k top-1 accuracy of 84.9% and 85.2%, respectively, with no additional training data, while outperforming other convolutional and large-kernel networks. Combining HyenaPixel with attention further improves accuracy. We attribute the success of bidirectional Hyena to learning the data-dependent geometric arrangement of pixels without a fixed neighborhood definition. Experimental results on downstream tasks suggest that HyenaPixel with large filters and a fixed neighborhood leads to better localization performance.

PaperPDFCode

Code

spravil/HyenaPixel officialmentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Image ClassificationObject DetectionSemantic Segmentation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Image Classification ImageNet HyenaPixel-Bidirectional-Former-B36 Top 1 Accuracy 85.2% #246 of 1060 Archive leaderboard report
Image Classification ImageNet HyenaPixel-Former-B36 Top 1 Accuracy 84.9% #277 of 1060 Archive leaderboard report
Image Classification ImageNet HyenaPixel-Attention-Former-S18 Top 1 Accuracy 83.6% #410 of 1060 Archive leaderboard report
Image Classification ImageNet HyenaPixel-Bidirectional-Former-S18 Top 1 Accuracy 83.5% #422 of 1060 Archive leaderboard report
Image Classification ImageNet HyenaPixel-Former-S18 Top 1 Accuracy 83.2% #450 of 1060 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

ConvolutionMetaFormer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections