{"url":"/dataset/roor","name":"ROOR","full_name":null,"description_markdown":"ROOR is a reading order prediction (ROP) benchmark which annotates layout reading order as ordering relations. \r\n\r\nLayout reading order is typically formulated as a permutation of layout elements, i.e. a sequence containing all the layout elements. \r\nHowever, multiple cases have reflected that this formulation does not adequately convey the complete reading order information in the layout, which may negatively affect the utilization of this signal in downstream VrD tasks. \r\n\r\nTherefore, this work investigate the properties of layout reading order, conceptualizing it with terms `Immediate Succession During Reading`(ISDR) and `Generalized Succession During Reading`(GSDR), and formulate each of which as an ordering relation over layout elements. \r\nThen, ROOR provides the annotation of the ISDR relationship over layout segments, based on the layout annotation of [EC-FUNSD](https://paperswithcode.com/dataset/ec-funsd). \r\nOverall, ROOR comprises 199 samples including 10,662 segments, 31,297 words and 10,967 annotated reading order linking pairs.\r\nWe hope the construction of this benchmark could facilitate the development of automated ROP methods of the improved task form.","description_withheld":null,"homepage":"https://github.com/chongzhangFDU/ROOR-Datasets","introduced_date":"2024-09-29","introduced_date_note":null,"introduced_by":{"paper":"/paper/modeling-layout-reading-order-as-ordering","title":"Modeling Layout Reading Order as Ordering Relations for Visually-rich Document Understanding","first_author":"Chong Zhang","url":null},"license":{"name":"CC-BY-4.0","url":"https://github.com/chongzhangFDU/ROOR-Datasets/blob/main/LICENSE"},"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Relation Extraction","url":"/task/relation-extraction","datasets_with_task":"/datasets/task/relation-extraction"},{"name":"Reading Order Detection","url":"/task/reading-order-detection","datasets_with_task":"/datasets/task/reading-order-detection"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["ROOR"],"data_loaders":[],"num_papers_in_archive":3,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/reading-order-detection-on-roor","task":"Reading Order Detection","dataset_variant":"ROOR","rows":4,"metrics":["Segment-level F1"],"first_row_in_archive_order":{"model":"LayoutLMv3-GlobalPointer (large)","paper":"/paper/modeling-layout-reading-order-as-ordering","metrics":{"Segment-level F1":"82.38"},"code_links":[{"title":"chongzhangFDU/ROOR","url":"https://github.com/chongzhangFDU/ROOR"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/modeling-layout-reading-order-as-ordering","title":"Modeling Layout Reading Order as Ordering Relations for Visually-rich Document Understanding","date":"2024-09-29","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/reading-order-matters-information-extraction","title":"Reading Order Matters: Information Extraction from Visually-rich Documents by Token Path Prediction","date":"2023-10-17","rows_on_this_dataset":1,"code_links":2,"syntology":null},{"paper":"/paper/layoutreader-pre-training-of-text-and-layout","title":"LayoutReader: Pre-training of Text and Layout for Reading Order Detection","date":"2021-08-26","rows_on_this_dataset":1,"code_links":1,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}