Papers › PE-former: Pose Estimation Transformer

PE-former: Pose Estimation Transformer

9 Dec 2021arXiv:2112.04981archive 2025-07-28

Paschalis Panteleris, Antonis Argyros

Vision transformer architectures have been demonstrated to work very effectively for image classification tasks. Efforts to solve more challenging vision tasks with transformers rely on convolutional backbones for feature extraction. In this paper we investigate the use of a pure transformer architecture (i.e., one with no CNN backbone) for the problem of 2D body pose estimation. We evaluate two ViT architectures on the COCO dataset. We demonstrate that using an encoder-decoder transformer architecture yields state of the art results on this estimation problem.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

padeler/pe-former officialmentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

DecoderImage ClassificationPose Estimationimage-classification

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Pose Estimation COCO (Common Objects in Context) PEFORMER-Xcit-dino-p8 AP 72.6 #9 of 10 Archive leaderboard report
Pose Estimation COCO (Common Objects in Context) PEFORMER-Xcit-dino-p8 AR 79.4 #9 of 10 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Absolute Position EncodingsAdamAttentionBPEBatch NormalizationConvolutionCross-Covariance AttentionDense ConnectionsDepthwise ConvolutionDetrDropoutFeedforward NetworkLabel SmoothingLayer NormalizationLinear LayerLocal Patch InteractionMulti-Head AttentionPosition-Wise Feed-Forward LayerResidual ConnectionSoftmaxTransformerXCiTXCiT Layer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections