{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/abcnet-v2-adaptive-bezier-curve-network-for","title":"ABCNet v2: Adaptive Bezier-Curve Network for Real-time End-to-end Text Spotting","arxiv_id":"2105.03620","date":"2021-05-08","proceeding":null,"authors":["Yuliang Liu","Chunhua Shen","Lianwen Jin","Tong He","Peng Chen","Chongyu Liu","Hao Chen"],"abstract":"End-to-end text-spotting, which aims to integrate detection and recognition in a unified framework, has attracted increasing attention due to its simplicity of the two complimentary tasks. It remains an open problem especially when processing arbitrarily-shaped text instances. Previous methods can be roughly categorized into two groups: character-based and segmentation-based, which often require character-level annotations and/or complex post-processing due to the unstructured output. Here, we tackle end-to-end text spotting by presenting Adaptive Bezier Curve Network v2 (ABCNet v2). Our main contributions are four-fold: 1) For the first time, we adaptively fit arbitrarily-shaped text by a parameterized Bezier curve, which, compared with segmentation-based methods, can not only provide structured output but also controllable representation. 2) We design a novel BezierAlign layer for extracting accurate convolution features of a text instance of arbitrary shapes, significantly improving the precision of recognition over previous methods. 3) Different from previous methods, which often suffer from complex post-processing and sensitive hyper-parameters, our ABCNet v2 maintains a simple pipeline with the only post-processing non-maximum suppression (NMS). 4) As the performance of text recognition closely depends on feature alignment, ABCNet v2 further adopts a simple yet effective coordinate convolution to encode the position of the convolutional filters, which leads to a considerable improvement with negligible computation overhead. Comprehensive experiments conducted on various bilingual (English and Chinese) benchmark datasets demonstrate that ABCNet v2 can achieve state-of-the-art performance while maintaining very high efficiency.","url_abs":"https://arxiv.org/abs/2105.03620v3","url_pdf":"https://arxiv.org/pdf/2105.03620v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"abcnet-v2-adaptive-bezier-curve-network-for","repo_url":"https://github.com/Yuliang-Liu/ABCNet_Chinese","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"text-spotting","task_name":"Text Spotting"}],"methods":[{"method_slug":"abcnet","method_name":"ABCNet"},{"method_slug":"bezieralign","method_name":"BezierAlign"},{"method_slug":"convolution","method_name":"Convolution"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/text-spotting-on-icdar-2015","task":"Text Spotting","dataset":"ICDAR 2015","model":"ABCNet v2","rank_in_archive_order":13,"of":18,"metrics":{"F-measure (%) - Generic Lexicon":"73.0","F-measure (%) - Strong Lexicon":"82.7","F-measure (%) - Weak Lexicon":"78.5"},"uses_additional_data":false},{"leaderboard":"/sota/text-spotting-on-inverse-text","task":"Text Spotting","dataset":"Inverse-Text","model":"ABCNet v2","rank_in_archive_order":7,"of":9,"metrics":{"F-measure (%) - Full Lexicon":"47.4","F-measure (%) - No Lexicon":"34.5"},"uses_additional_data":false},{"leaderboard":"/sota/text-spotting-on-scut-ctw1500","task":"Text Spotting","dataset":"SCUT-CTW1500","model":"ABCNet v2","rank_in_archive_order":7,"of":11,"metrics":{"F-Measure (%) - Full Lexicon":"77.2","F-measure (%) - No Lexicon":"57.5"},"uses_additional_data":false},{"leaderboard":"/sota/text-spotting-on-total-text","task":"Text Spotting","dataset":"Total-Text","model":"ABCNet v2","rank_in_archive_order":12,"of":12,"metrics":{"F-measure (%) - Full Lexicon":"78.1","F-measure (%) - No Lexicon":"70.4"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2105.03620","atlas_url":"https://app.syntology.ai/?focus=2105.03620","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}