{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/rotation-sensitive-regression-for-oriented","title":"Rotation-Sensitive Regression for Oriented Scene Text Detection","arxiv_id":"1803.05265","date":"2018-03-14","proceeding":"CVPR 2018 6","authors":["Minghui Liao","Zhen Zhu","Baoguang Shi","Gui-Song Xia","Xiang Bai"],"abstract":"Text in natural images is of arbitrary orientations, requiring detection in\nterms of oriented bounding boxes. Normally, a multi-oriented text detector\noften involves two key tasks: 1) text presence detection, which is a\nclassification problem disregarding text orientation; 2) oriented bounding box\nregression, which concerns about text orientation. Previous methods rely on\nshared features for both tasks, resulting in degraded performance due to the\nincompatibility of the two tasks. To address this issue, we propose to perform\nclassification and regression on features of different characteristics,\nextracted by two network branches of different designs. Concretely, the\nregression branch extracts rotation-sensitive features by actively rotating the\nconvolutional filters, while the classification branch extracts\nrotation-invariant features by pooling the rotation-sensitive features. The\nproposed method named Rotation-sensitive Regression Detector (RRD) achieves\nstate-of-the-art performance on three oriented scene text benchmark datasets,\nincluding ICDAR 2015, MSRA-TD500, RCTW-17 and COCO-Text. Furthermore, RRD\nachieves a significant improvement on a ship collection dataset, demonstrating\nits generality on oriented object detection.","url_abs":"http://arxiv.org/abs/1803.05265v1","url_pdf":"http://arxiv.org/pdf/1803.05265v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"oriented-object-detection","task_name":"Oriented Object Detection"},{"task_slug":"scene-text-detection","task_name":"Scene Text Detection"},{"task_slug":"text-detection","task_name":"Text Detection"},{"task_slug":"object-detection-1","task_name":"object-detection"},{"task_slug":"regression-1","task_name":"regression"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/scene-text-detection-on-msra-td500","task":"Scene Text Detection","dataset":"MSRA-TD500","model":"RRD∗","rank_in_archive_order":14,"of":18,"metrics":{"F-Measure":"79","Precision":"87","Recall":"73"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1803.05265","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}