Papers › ConvFormer: Parameter Reduction in Transformer Models for 3D Human Pose Estimation by...

ConvFormer: Parameter Reduction in Transformer Models for 3D Human Pose Estimation by Leveraging Dynamic Multi-Headed Convolutional Attention

4 Apr 2023arXiv:2304.02147archive 2025-07-28

Alec Diaz-Arias, Dmitriy Shin

Recently, fully-transformer architectures have replaced the defacto convolutional architecture for the 3D human pose estimation task. In this paper we propose \textbf{\textit{ConvFormer}}, a novel convolutional transformer that leverages a new \textbf{\textit{dynamic multi-headed convolutional self-attention}} mechanism for monocular 3D human pose estimation. We designed a spatial and temporal convolutional transformer to comprehensively model human joint relations within individual frames and globally across the motion sequence. Moreover, we introduce a novel notion of \textbf{\textit{temporal joints profile}} for our temporal ConvFormer that fuses complete temporal information immediately for a local neighborhood of joint features. We have quantitatively and qualitatively validated our method on three common benchmark datasets: Human3.6M, MPI-INF-3DHP, and HumanEva. Extensive experiments have been conducted to identify the optimal hyper-parameter set. These experiments demonstrated that we achieved a \textbf{significant parameter reduction relative to prior transformer models} while attaining State-of-the-Art (SOTA) or near SOTA on all three datasets. Additionally, we achieved SOTA for Protocol III on H36M for both GT and CPN detection inputs. Finally, we obtained SOTA on all three metrics for the MPI-INF-3DHP dataset and for all three subjects on HumanEva under Protocol II.

PaperPDFCode

Code

ajda1992/convformer mentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

3D Human Pose EstimationMonocular 3D Human Pose EstimationPose Estimation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
3D Human Pose Estimation Human3.6M ConvFormer (T=243, CPN) Average MPJPE (mm) 43.2 #28 of 88 Archive leaderboard report
3D Human Pose Estimation Human3.6M ConvFormer (T=243, CPN) Multi-View or Monocular Monocular #28 of 88 Archive leaderboard report
3D Human Pose Estimation Human3.6M ConvFormer (T=243, CPN) Using 2D ground-truth joints No #28 of 88 Archive leaderboard report
3D Human Pose Estimation HumanEva-I ConvFormer (T=43) Mean Reconstruction Error (mm) 24.3 #19 of 31 Archive leaderboard report
3D Human Pose Estimation MPI-INF-3DHP ConvFormer AUC 69.8 #21 of 108 Archive leaderboard report
3D Human Pose Estimation MPI-INF-3DHP ConvFormer MPJPE 53.6 #21 of 108 Archive leaderboard report
3D Human Pose Estimation MPI-INF-3DHP ConvFormer PCK 96.4 #21 of 108 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Absolute Position EncodingsAdamAttentionBPECPNConvolutionDense ConnectionsDropoutLabel SmoothingLayer NormalizationLinear LayerMulti-Head AttentionPosition-Wise Feed-Forward LayerResidual ConnectionSoftmaxTransformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections