{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/script-identification-in-natural-scene-image","title":"Script Identification in Natural Scene Image and Video Frame using Attention based Convolutional-LSTM Network","arxiv_id":"1801.00470","date":"2018-01-01","proceeding":null,"authors":["Ankan Kumar Bhunia","Aishik Konwer","Ayan Kumar Bhunia","Abir Bhowmick","Partha P. Roy","Umapada Pal"],"abstract":"Script identification plays a significant role in analysing documents and\nvideos. In this paper, we focus on the problem of script identification in\nscene text images and video scripts. Because of low image quality, complex\nbackground and similar layout of characters shared by some scripts like Greek,\nLatin, etc., text recognition in those cases become challenging. In this paper,\nwe propose a novel method that involves extraction of local and global features\nusing CNN-LSTM framework and weighting them dynamically for script\nidentification. First, we convert the images into patches and feed them into a\nCNN-LSTM framework. Attention-based patch weights are calculated applying\nsoftmax layer after LSTM. Next, we do patch-wise multiplication of these\nweights with corresponding CNN to yield local features. Global features are\nalso extracted from last cell state of LSTM. We employ a fusion technique which\ndynamically weights the local and global features for an individual patch.\nExperiments have been done in four public script identification datasets:\nSIW-13, CVSI2015, ICDAR-17 and MLe2e. The proposed framework achieves superior\nresults in comparison to conventional methods.","url_abs":"http://arxiv.org/abs/1801.00470v4","url_pdf":"http://arxiv.org/pdf/1801.00470v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"script-identification-in-natural-scene-image","repo_url":"https://github.com/ankanbhunia/AttenScriptNetPR","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[],"methods":[{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}