{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/panoptic-animal-pose-estimators-are-zero-shot","title":"SuperAnimal pretrained pose estimation models for behavioral analysis","arxiv_id":"2203.07436","date":"2022-03-14","proceeding":null,"authors":["Shaokai Ye","Anastasiia Filippova","Jessy Lauer","Steffen Schneider","Maxime Vidal","Tian Qiu","Alexander Mathis","Mackenzie Weygandt Mathis"],"abstract":"Quantification of behavior is critical in applications ranging from neuroscience, veterinary medicine and animal conservation efforts. A common key step for behavioral analysis is first extracting relevant keypoints on animals, known as pose estimation. However, reliable inference of poses currently requires domain knowledge and manual labeling effort to build supervised models. We present a series of technical innovations that enable a new method, collectively called SuperAnimal, to develop unified foundation models that can be used on over 45 species, without additional human labels. Concretely, we introduce a method to unify the keypoint space across differently labeled datasets (via our generalized data converter) and for training these diverse datasets in a manner such that they don't catastrophically forget keypoints given the unbalanced inputs (via our keypoint gradient masking and memory replay approaches). These models show excellent performance across six pose benchmarks. Then, to ensure maximal usability for end-users, we demonstrate how to fine-tune the models on differently labeled data and provide tooling for unsupervised video adaptation to boost performance and decrease jitter across frames. If the models are fine-tuned, we show SuperAnimal models are 10-100$\\times$ more data efficient than prior transfer-learning-based approaches. We illustrate the utility of our models in behavioral classification in mice and gait analysis in horses. Collectively, this presents a data-efficient solution for animal pose estimation.","url_abs":"https://arxiv.org/abs/2203.07436v4","url_pdf":"https://arxiv.org/pdf/2203.07436v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"panoptic-animal-pose-estimators-are-zero-shot","repo_url":"https://github.com/DeepLabCut/DeepLabCut","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"LGPL-3.0"}},{"paper_slug":"panoptic-animal-pose-estimators-are-zero-shot","repo_url":"https://github.com/adaptivemotorcontrollab/modelzoo-figures","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"GPL-3.0"}},{"paper_slug":"panoptic-animal-pose-estimators-are-zero-shot","repo_url":"https://github.com/AlexEMG/DeepLabCut","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"LGPL-3.0"}}],"tasks":[{"task_slug":"2d-pose-estimation","task_name":"2D Pose Estimation"},{"task_slug":"animal-pose-estimation","task_name":"Animal Pose Estimation"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"}],"methods":[],"datasets_introduced":[{"slug":"superanimal-quadruped","name":"SuperAnimal-Quadruped","full_name":""},{"slug":"superanimal-topviewmouse","name":"SuperAnimal-TopViewMouse","full_name":""},{"slug":"irodent","name":"iRodent","full_name":"iRodent Animal Pose Estimation"}],"methods_introduced":[],"results":[{"leaderboard":"/sota/2d-pose-estimation-on-irodent","task":"2D Pose Estimation","dataset":"iRodent","model":"fine-tuned HRNetw32 pretrained on SuperAnimal (1 fac of data)","rank_in_archive_order":1,"of":8,"metrics":{"Average mAP":"72.971"},"uses_additional_data":true},{"leaderboard":"/sota/2d-pose-estimation-on-irodent","task":"2D Pose Estimation","dataset":"iRodent","model":"fine-tuned HRNetw32 pretrained on AP-10K (1 fac of data)","rank_in_archive_order":2,"of":8,"metrics":{"Average mAP":"61.635"},"uses_additional_data":true},{"leaderboard":"/sota/2d-pose-estimation-on-irodent","task":"2D Pose Estimation","dataset":"iRodent","model":"fine-tuned HRNetw32 pretrained on SuperAnimal (0.01 fac of data)","rank_in_archive_order":3,"of":8,"metrics":{"Average mAP":"60.853"},"uses_additional_data":true},{"leaderboard":"/sota/2d-pose-estimation-on-irodent","task":"2D Pose Estimation","dataset":"iRodent","model":"fine-tuned HRNetw32 pretrained on ImageNet","rank_in_archive_order":4,"of":8,"metrics":{"Average mAP":"58.857"},"uses_additional_data":true},{"leaderboard":"/sota/2d-pose-estimation-on-irodent","task":"2D Pose Estimation","dataset":"iRodent","model":"zero-shot HRNet-w32 pretrained on SuperAnimal-Quadruped","rank_in_archive_order":5,"of":8,"metrics":{"Average mAP":"58.557"},"uses_additional_data":true},{"leaderboard":"/sota/2d-pose-estimation-on-irodent","task":"2D Pose Estimation","dataset":"iRodent","model":"zero-shot AnimalTokenPose pretrained on AP-10K","rank_in_archive_order":6,"of":8,"metrics":{"Average mAP":"55.415"},"uses_additional_data":true},{"leaderboard":"/sota/2d-pose-estimation-on-irodent","task":"2D Pose Estimation","dataset":"iRodent","model":"fine-tuned HRNetw32 pretrained on AP-10K (0.01 fac of data)","rank_in_archive_order":7,"of":8,"metrics":{"Average mAP":"43.144"},"uses_additional_data":true},{"leaderboard":"/sota/2d-pose-estimation-on-irodent","task":"2D Pose Estimation","dataset":"iRodent","model":"zero-shot HRNet-w32 pretrained on AP-10K","rank_in_archive_order":8,"of":8,"metrics":{"Average mAP":"40.389"},"uses_additional_data":true},{"leaderboard":"/sota/animal-pose-estimation-on-ap-10k","task":"Animal Pose Estimation","dataset":"AP-10K","model":"SuperAnimal-HRNetw32","rank_in_archive_order":3,"of":10,"metrics":{"AP":"80.113"},"uses_additional_data":true},{"leaderboard":"/sota/animal-pose-estimation-on-ap-10k","task":"Animal Pose Estimation","dataset":"AP-10K","model":"zero-shot SuperAnimal-HRNetw32","rank_in_archive_order":10,"of":10,"metrics":{"AP":"68.038"},"uses_additional_data":false},{"leaderboard":"/sota/animal-pose-estimation-on-animal-pose-dataset","task":"Animal Pose Estimation","dataset":"Animal-Pose Dataset","model":"SuperAnimal-AnimalTokenPose","rank_in_archive_order":1,"of":1,"metrics":{"AP":"86"},"uses_additional_data":true},{"leaderboard":"/sota/animal-pose-estimation-on-horse-10","task":"Animal Pose Estimation","dataset":"Horse-10","model":"SuperAnimal-Quadruped HRNet-w32","rank_in_archive_order":7,"of":8,"metrics":{"Normalized Error (OOD)":"0.1091"},"uses_additional_data":true},{"leaderboard":"/sota/animal-pose-estimation-on-horse-10","task":"Animal Pose Estimation","dataset":"Horse-10","model":"mmpose HRNet-w32 (w/ImageNet pretrained weights)","rank_in_archive_order":8,"of":8,"metrics":{"Normalized Error (OOD)":"0.179"},"uses_additional_data":true},{"leaderboard":"/sota/animal-pose-estimation-on-trimouse-161","task":"Animal Pose Estimation","dataset":"TriMouse-161","model":"SuperAnimal HRNetw32","rank_in_archive_order":2,"of":7,"metrics":{"mAP":"98.547"},"uses_additional_data":false},{"leaderboard":"/sota/animal-pose-estimation-on-trimouse-161","task":"Animal Pose Estimation","dataset":"TriMouse-161","model":"zero-shot SuperAnimal HRNetw32","rank_in_archive_order":7,"of":7,"metrics":{"mAP":"76.139"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2203.07436","atlas_url":"https://app.syntology.ai/?focus=2203.07436","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}