{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-deep-neural-framework-for-continuous-sign","title":"A Deep Neural Framework for Continuous Sign Language Recognition by Iterative Training","arxiv_id":null,"date":"2019-07-01","proceeding":"IEEE Transactions on Multimedia 2019 7","authors":["Runpeng Cui","Hu Liu","ChangShui Zhang"],"abstract":"This work develops a continuous sign language (SL)\r\nrecognition framework with deep neural networks, which directly\r\ntranscribes videos of SL sentences to sequences of ordered gloss\r\nlabels. Previous methods dealing with continuous SL recognition\r\nusually employ hidden Markov models with limited capacity to\r\ncapture the temporal information. In contrast, our proposed\r\narchitecture adopts deep convolutional neural networks with\r\nstacked temporal fusion layers as the feature extraction module,\r\nand bi-directional recurrent neural networks as the sequence\r\nlearning module. We propose an iterative optimization process\r\nfor our architecture to fully exploit the representation capability\r\nof deep neural networks with limited data. We first train the\r\nend-to-end recognition model for alignment proposal, and then\r\nuse the alignment proposal as strong supervisory information\r\nto directly tune the feature extraction module. This training\r\nprocess can run iteratively to achieve improvements on the\r\nrecognition performance. We further contribute by exploring\r\nthe multimodal fusion of RGB images and optical flow in\r\nsign language. Our method is evaluated on two challenging SL\r\nrecognition benchmarks, and outperforms the state-of-the-art by\r\na relative improvement of more than 15% on both databases.","url_abs":"http://www.kresttechnology.com/krest-academic-projects/krest-mtech-projects/CSE/M.Tech%20Computer%20Science%202020/Artificial%20Intelligence/Basepaper-AI/12.%20A%20deep%20neural%20framework.pdf","url_pdf":"http://www.kresttechnology.com/krest-academic-projects/krest-mtech-projects/CSE/M.Tech%20Computer%20Science%202020/Artificial%20Intelligence/Basepaper-AI/12.%20A%20deep%20neural%20framework.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-deep-neural-framework-for-continuous-sign","repo_url":"https://github.com/iliasprc/slrzoo","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"optical-flow-estimation","task_name":"Optical Flow Estimation"},{"task_slug":"sign-language-recognition","task_name":"Sign Language Recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/sign-language-recognition-on-rwth-phoenix","task":"Sign Language Recognition","dataset":"RWTH-PHOENIX-Weather 2014","model":"DNF","rank_in_archive_order":15,"of":22,"metrics":{"Word Error Rate (WER)":"22.86"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}