{"url":"/dataset/cova","name":"CoVA","full_name":"CoVA dataset for Webpage Object Detection / Information Extraction","description_markdown":"We labeled _7,740_ webpage screenshots spanning _408_ domains (Amazon, Walmart, Target, etc.). Each of these webpages contains exactly one labeled price, title, and image. All other web elements are labeled as background. On average, there are _90_ web elements in a webpage.\r\n\r\nWebpage screenshots and bounding boxes can be obtained [here](https://drive.google.com/drive/folders/1LQPXGhDVh40bIT2-LZfo498M93tidABe?usp=sharing)\r\n\r\n### Train-Val-Test split\r\nWe create a cross-domain split which ensures that each of the train, val and test sets contains webpages from different domains. Specifically, we construct a 3 : 1 : 1 split based on the number of distinct domains. We observed that the top-5 domains (based on number of samples) were Amazon, EBay, Walmart, Etsy, and Target. So, we created 5 different splits for 5-Fold Cross Validation such that each of the major domains is present in one of the 5 splits for test data.","description_withheld":null,"homepage":"https://github.com/kevalmorabia97/CoVA-Web-Object-Detection/blob/master/README.md","introduced_date":"2021-10-24","introduced_date_note":null,"introduced_by":{"paper":"/paper/cova-context-aware-visual-attention-for","title":"CoVA: Context-aware Visual Attention for Webpage Information Extraction","first_author":"Anurendra Kumar","url":null},"license":null,"modalities":[{"name":"Images","url":"/datasets/modality/images"}],"tasks":[{"name":"Webpage Object Detection","url":"/task/webpage-object-detection","datasets_with_task":"/datasets/task/webpage-object-detection"}],"languages":[],"variants":["CoVA"],"data_loaders":[{"repo":"https://github.com/kevalmorabia97/cova-web-object-detection","url":"https://github.com/kevalmorabia97/cova-web-object-detection","frameworks":["pytorch"]}],"num_papers_in_archive":3,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}