{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/signbert-pre-training-of-hand-model-aware-1","title":"SignBERT: Pre-Training of Hand-Model-Aware Representation for Sign Language Recognition","arxiv_id":"2110.05382","date":"2021-10-11","proceeding":"ICCV 2021 10","authors":["Hezhen Hu","Weichao Zhao","Wengang Zhou","Yuechen Wang","Houqiang Li"],"abstract":"Hand gesture serves as a critical role in sign language. Current deep-learning-based sign language recognition (SLR) methods may suffer insufficient interpretability and overfitting due to limited sign data sources. In this paper, we introduce the first self-supervised pre-trainable SignBERT with incorporated hand prior for SLR. SignBERT views the hand pose as a visual token, which is derived from an off-the-shelf pose extractor. The visual tokens are then embedded with gesture state, temporal and hand chirality information. To take full advantage of available sign data sources, SignBERT first performs self-supervised pre-training by masking and reconstructing visual tokens. Jointly with several mask modeling strategies, we attempt to incorporate hand prior in a model-aware method to better model hierarchical context over the hand sequence. Then with the prediction head added, SignBERT is fine-tuned to perform the downstream SLR task. To validate the effectiveness of our method on SLR, we perform extensive experiments on four public benchmark datasets, i.e., NMFs-CSL, SLR500, MSASL and WLASL. Experiment results demonstrate the effectiveness of both self-supervised learning and imported hand prior. Furthermore, we achieve state-of-the-art performance on all benchmarks with a notable gain.","url_abs":"https://arxiv.org/abs/2110.05382v1","url_pdf":"https://arxiv.org/pdf/2110.05382v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"self-supervised-learning","task_name":"Self-Supervised Learning"},{"task_slug":"sign-language-recognition","task_name":"Sign Language Recognition"}],"methods":[{"method_slug":"slr","method_name":"SLR"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/sign-language-recognition-on-wlasl100","task":"Sign Language Recognition","dataset":"WLASL100","model":"SignBERT","rank_in_archive_order":3,"of":7,"metrics":{"Official Test Split":"true","Top-1 Accuracy":"83.30"},"uses_additional_data":true}],"syntology":{"syntology_url":"https://syntology.ai/paper/2110.05382","atlas_url":"https://app.syntology.ai/?focus=2110.05382","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}