{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/few-shot-pixel-precise-document-layout","title":"Few-shot pixel-precise document layout segmentation via dynamic instance generation and local thresholding","arxiv_id":null,"date":"2023-08-10","proceeding":"International Journal of Neural Systems (IJNS) 2023 8","authors":["Axel De Nardin","Silvia Zottin","Claudio Piciarelli","Emanuela Colombi","Gian Luca Foresti"],"abstract":"Over the years, the humanities community has increasingly requested the creation of artificial intelligence\r\nframeworks to help the study of cultural heritage. Document Layout segmentation, which aims at\r\nidentifying the different structural components of a document page, is a particularly interesting task\r\nconnected to this trend, specifically when it comes to handwritten texts. While there are many effective\r\napproaches to this problem, they all rely on large amounts of data for the training of the underlying\r\nmodels, which is rarely possible in a real-world scenario, as the process of producing the ground truth\r\nsegmentation task with the required precision to the pixel level is a very time-consuming task and often\r\nrequires a certain degree of domain knowledge regarding the documents at hand. For this reason, in the\r\npresent paper, we propose an effective few-shot learning framework for document layout segmentation\r\nrelying on two novel components, namely a dynamic instance generation and a segmentation refinement\r\nmodule. This approach is able of achieving performances comparable to the current state of the art on\r\nthe popular Diva-HisDB dataset, while relying on just a fraction of the available data.","url_abs":"https://www.worldscientific.com/doi/10.1142/S0129065723500521?srsltid=AfmBOooj60Tfvwgncg7FpFMhlbZbZy5GqpffiGLkilRa6fOi4tKg-KQ-","url_pdf":"https://air.uniud.it/handle/11390/1259086","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"few-shot-learning","task_name":"Few-Shot Learning"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/semantic-segmentation-on-diva-hisdb","task":"Semantic Segmentation","dataset":"DIVA-HisDB","model":"IJNS '23 (few-shot)","rank_in_archive_order":2,"of":3,"metrics":{"Mean IoU (class)":"97.23"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}