{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/efficientpose-scalable-single-person-pose","title":"EfficientPose: Scalable single-person pose estimation","arxiv_id":"2004.12186","date":"2020-04-25","proceeding":null,"authors":["Daniel Groos","Heri Ramampiaro","Espen A. F. Ihlen"],"abstract":"Single-person human pose estimation facilitates markerless movement analysis in sports, as well as in clinical applications. Still, state-of-the-art models for human pose estimation generally do not meet the requirements of real-life applications. The proliferation of deep learning techniques has resulted in the development of many advanced approaches. However, with the progresses in the field, more complex and inefficient models have also been introduced, which have caused tremendous increases in computational demands. To cope with these complexity and inefficiency challenges, we propose a novel convolutional neural network architecture, called EfficientPose, which exploits recently proposed EfficientNets in order to deliver efficient and scalable single-person pose estimation. EfficientPose is a family of models harnessing an effective multi-scale feature extractor and computationally efficient detection blocks using mobile inverted bottleneck convolutions, while at the same time ensuring that the precision of the pose configurations is still improved. Due to its low complexity and efficiency, EfficientPose enables real-world applications on edge devices by limiting the memory footprint and computational cost. The results from our experiments, using the challenging MPII single-person benchmark, show that the proposed EfficientPose models substantially outperform the widely-used OpenPose model both in terms of accuracy and computational efficiency. In particular, our top-performing model achieves state-of-the-art accuracy on single-person MPII, with low-complexity ConvNets.","url_abs":"https://arxiv.org/abs/2004.12186v2","url_pdf":"https://arxiv.org/pdf/2004.12186v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"efficientpose-scalable-single-person-pose","repo_url":"https://github.com/daniegr/EfficientPose","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"2d-human-pose-estimation","task_name":"2D Human Pose Estimation"},{"task_slug":"computational-efficiency","task_name":"Computational Efficiency"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"}],"methods":[{"method_slug":"1x1-convolution","method_name":"1x1 Convolution"},{"method_slug":"average-pooling","method_name":"Average Pooling"},{"method_slug":"batch-normalization","method_name":"Batch Normalization"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"cross-resolution-features","method_name":"Cross-resolution features"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"depthwise-convolution","method_name":"Depthwise Convolution"},{"method_slug":"depthwise-separable-convolution","method_name":"Depthwise Separable Convolution"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"e-mbconv","method_name":"E-MBConv"},{"method_slug":"e-swish","method_name":"E-swish"},{"method_slug":"efficientnet","method_name":"EfficientNet"},{"method_slug":"heatmap","method_name":"Heatmap"},{"method_slug":"high-level-backbone","method_name":"High-level backbone"},{"method_slug":"high-resolution-input","method_name":"High-resolution input"},{"method_slug":"inverted-residual-block","method_name":"Inverted Residual Block"},{"method_slug":"low-level-backbone","method_name":"Low-level backbone"},{"method_slug":"low-resolution-input","method_name":"Low-resolution input"},{"method_slug":"mobile-densenet","method_name":"Mobile DenseNet"},{"method_slug":"openpose","method_name":"OpenPose"},{"method_slug":"pafs","method_name":"PAFs"},{"method_slug":"pointwise-convolution","method_name":"Pointwise Convolution"},{"method_slug":"rmsprop","method_name":"RMSProp"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"squeeze-and-excitation-block","method_name":"Squeeze-and-Excitation Block"},{"method_slug":"transposed-convolution","method_name":"Transposed convolution"}],"datasets_introduced":[],"methods_introduced":[{"slug":"cross-resolution-features","name":"Cross-resolution features","full_name":"Cross-resolution features"},{"slug":"e-mbconv","name":"E-MBConv","full_name":"E-MBConv"},{"slug":"high-level-backbone","name":"High-level backbone","full_name":"High-level backbone"},{"slug":"high-resolution-input","name":"High-resolution input","full_name":"High-resolution input"},{"slug":"low-level-backbone","name":"Low-level backbone","full_name":"Low-level backbone"},{"slug":"low-resolution-input","name":"Low-resolution input","full_name":"Low-resolution input"},{"slug":"mobile-densenet","name":"Mobile DenseNet","full_name":"Mobile DenseNet"}],"results":[{"leaderboard":"/sota/pose-estimation-on-mpii-human-pose","task":"Pose Estimation","dataset":"MPII Human Pose","model":"EfficientPose IV","rank_in_archive_order":21,"of":46,"metrics":{"PCKh-0.5":"91.2"},"uses_additional_data":false},{"leaderboard":"/sota/pose-estimation-on-mpii-human-pose","task":"Pose Estimation","dataset":"MPII Human Pose","model":"OpenPose","rank_in_archive_order":31,"of":46,"metrics":{"PCKh-0.5":"88.8"},"uses_additional_data":false},{"leaderboard":"/sota/pose-estimation-on-mpii-human-pose","task":"Pose Estimation","dataset":"MPII Human Pose","model":"EfficientPose RT","rank_in_archive_order":38,"of":46,"metrics":{"PCKh-0.5":"84.8"},"uses_additional_data":false},{"leaderboard":"/sota/pose-estimation-on-mpii-single-person","task":"Pose Estimation","dataset":"MPII Single Person","model":"EfficientPose IV","rank_in_archive_order":3,"of":5,"metrics":{"PCKh@0.1":"36.0","PCKh@0.5":"91.2"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}