Datasets › CORD
CORD (Consolidated Receipt Dataset for Post-OCR Parsing)
OCR is inevitably linked to NLP since its final output is in text. Advances in document intelligence are driving the need for a unified technology that integrates OCR with various NLP tasks, especially semantic parsing. Since OCR and semantic parsing have been studied as separate tasks so far, the datasets for each task on their own are rich, while those for the integrated post-OCR parsing tasks are relatively insufficient. In this study, we publish a consolidated dataset for receipt parsing as the first step towards post-OCR parsing tasks. The dataset consists of thousands of Indonesian receipts, which contains images and box/text annotations for OCR, and multi-level semantic labels for parsing. The proposed dataset can be used to address various OCR and parsing tasks.
Benchmarks archive 2025-07-28
All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Key Information Extraction | CORD | RORE (GeoLayoutLM) F1 98.52 | Modeling Layout Reading Order as Ordering Relations for... | chongzhangFDU/ROOR | 9 | Compare |
Papers archive 2025-07-28
7 shown of 7 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 100. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| Modeling Layout Reading Order as Ordering Relations for Visually-rich Document Understanding | 1 | 1 | 29 Sep 2024 | not harvested |
| Reading Order Matters: Information Extraction from Visually-rich Documents by Token Path Prediction | 2 | 1 | 17 Oct 2023 | not harvested |
| LayoutMask: Enhance Text-Layout Interaction in Multi-modal Pre-training for Document Understanding | 0 | 2 | 30 May 2023 | not harvested |
| GeoLayoutLM: Geometric Pre-training for Visual Information Extraction | 1 | 1 | 21 Apr 2023 | not harvested |
| LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking | 4 | 1 | 18 Apr 2022 | not harvested |
| LiLT: A Simple yet Effective Language-Independent Layout Transformer for Structured Document Understanding | 5 | 1 | 28 Feb 2022 | ran 2 of 3 samples (1 unverified) |
| LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding | 9 | 2 | 29 Dec 2020 | not harvested |
Dataset loaders archive 2025-07-28
1 loader as listed in the archive; links are outbound and not re-checked here.
Tasks archive 2025-07-28
License archive 2025-07-28
Creative Commons Attribution 4.0 International License
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- CORD
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections