{"url":"/dataset/cord-r","name":"CORD-r","full_name":null,"description_markdown":"We introduce FUNSD-r and CORD-r in [Token Path Prediction](https://arxiv.org/abs/2310.11016), the revised VrD-NER datasets to reflect the real-world scenarios of NER on scanned VrDs.\r\n\r\nIn FUNSD and CORD, segment layout annotations are aligned with labeled entities, which makes them not reflect the reading order issue of NER on scanned VrDs, and thus are unsuitable for evaluating current methods. In FUNSD-r and CORD-r, we automatically reannotate the layouts using PP-OCRv3 OCR system, and manually reannotate the named entities as word sequences based on the new layout annotations. Their segment layout annotations are aligned with real-world situations and entity mentions are labeled on words.\r\n\r\nThe proposed CORD-r consists of 999 document samples including the image, layout annotation of segments and words, and labeled entities of 30 categories. For the detailed summary statistics, please refer to the original paper.","description_withheld":null,"homepage":"https://github.com/chongzhangFDU/Token-Path-Prediction-Datasets","introduced_date":null,"introduced_date_note":null,"introduced_by":null,"license":{"name":"CC-BY-4.0","url":null},"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Named Entity Recognition (NER)","url":"/task/named-entity-recognition-ner","datasets_with_task":"/datasets/task/named-entity-recognition-ner"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["CORD-r"],"data_loaders":[],"num_papers_in_archive":3,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/named-entity-recognition-ner-on-cord-r","task":"Named Entity Recognition (NER)","dataset_variant":"CORD-r","rows":4,"metrics":["F1"],"first_row_in_archive_order":{"model":"TPP (LayoutLMv3)","paper":"/paper/reading-order-matters-information-extraction","metrics":{"F1":"91.85"},"code_links":[{"title":"chongzhangfdu/tpp","url":"https://github.com/chongzhangfdu/tpp"},{"title":"WinterShiver/Token-Path-Prediction","url":"https://github.com/WinterShiver/Token-Path-Prediction"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/reading-order-matters-information-extraction","title":"Reading Order Matters: Information Extraction from Visually-rich Documents by Token Path Prediction","date":"2023-10-17","rows_on_this_dataset":2,"code_links":2,"syntology":null},{"paper":"/paper/layoutmask-enhance-text-layout-interaction-in","title":"LayoutMask: Enhance Text-Layout Interaction in Multi-modal Pre-training for Document Understanding","date":"2023-05-30","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/layoutlmv3-pre-training-for-document-ai-with","title":"LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking","date":"2022-04-18","rows_on_this_dataset":1,"code_links":4,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}