{"url":"/method/semantic-reasoning-network","slug":"semantic-reasoning-network","name":"Semantic Reasoning Network","full_name":"Semantic Reasoning Network","full_name_withheld":false,"description_markdown":"**Semantic reasoning network**, or **SRN**, is an end-to-end trainable framework for scene text recognition that consists of four parts: backbone network, parallel [visual attention](https://paperswithcode.com/method/visual-attention) module (PVAM), global semantic reasoning module (GSRM), and visual-semantic fusion decoder (VSFD). Given an input image, the backbone network is first used to extract 2D features $V$. Then, the PVAM is used to generate $N$ aligned 1-D features $G$, where each feature corresponds to a character in the text and captures the aligned visual information. These $N$ 1-D features $G$ are then fed into a GSRM to capture the semantic information $S$. Finally, the aligned visual features $G$ and the semantic information $S$ are fused by the VSFD to predict $N$ characters. For text string shorter than $N$, ’EOS’ are padded.","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"https://arxiv.org/abs/2003.12294v1","title":"Towards Accurate Scene Text Recognition with Semantic Reasoning Networks","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Scene Text Models","url":"/methods/category/scene-text-models","pwc_aliases":[]}],"n_papers_tagged":3,"archive_num_papers":null,"papers_newest_first":[{"paper":null,"title":"Exploring Optical-Flow-Guided Motion and Detection-Based Appearance for Temporal Sentence Grounding","date":"2022-03-06","arxiv_id":"2203.02966","n_code_links":0,"syntology":null},{"paper":"/paper/bi-temporal-semantic-reasoning-for-the","title":"Bi-Temporal Semantic Reasoning for the Semantic Change Detection in HR Remote Sensing Images","date":"2021-08-13","arxiv_id":"2108.06103","n_code_links":1,"syntology":null},{"paper":"/paper/towards-accurate-scene-text-recognition-with","title":"Towards Accurate Scene Text Recognition with Semantic Reasoning Networks","date":"2020-03-27","arxiv_id":"2003.12294","n_code_links":3,"syntology":null}],"papers_shown":3,"tasks":[{"task":"/task/change-detection","name":"Change Detection","papers":1},{"task":"/task/object","name":"Object","papers":1},{"task":"/task/object-detection","name":"Object Detection","papers":1},{"task":"/task/optical-character-recognition","name":"Optical Character Recognition (OCR)","papers":1},{"task":"/task/optical-flow-estimation","name":"Optical Flow Estimation","papers":1},{"task":"/task/scene-text-recognition","name":"Scene Text Recognition","papers":1},{"task":"/task/sentence","name":"Sentence","papers":1},{"task":"/task/temporal-sentence-grounding","name":"Temporal Sentence Grounding","papers":1},{"task":"/task/object-detection-1","name":"object-detection","papers":1}],"tasks_shown":9,"n_tasks":9,"usage_by_year":[{"year":"2020","papers":1},{"year":"2021","papers":1},{"year":"2022","papers":1}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/semantic-reasoning-network"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}