{"url":"/dataset/impact-patent","name":"IMPACT Patent","full_name":"A Large-scale Integrated Multimodal Patent Analysis and Creation Dataset for Design Patents","description_markdown":"It is a large-scale multimodal patent dataset with detailed captions for design patent figures.\r\n\r\n💥 Our dataset includes half a million design patents comprising 3.61 million figures along with captions from patents granted by the United States Patent and Trademark Office USPTO over a 16-year period from 2007 to 2022.\r\n\r\n* Multimodality: We introduce a multimodal patent dataset that includes patent images, metadata, and detailed captions to support a variety of NLP, vision, and multimodal tasks. This dataset is also valuable for patent analysis tasks such as classification, retrieval, prior art searches, and design trend analysis.\r\n* Comprehensive dataset: We have compiled a collection of 435,101 patents spanning 16 years from U.S. design patent documents. This extensive collection includes a total of 3,609,805 drawing figures. Additionally, our dataset consists of eleven fields such as the title, patent ID, claims, date of publication, classification code, and extensive image-related information, including the number of images per patent and descriptions of the viewpoints.\r\n* Descriptive captions: To address the absence of descriptions about the designs, such as features and shapes, we generate elaborated captions by employing a vision-language model. It generates descriptive captions for the design figures, capturing details from the sketch. These captions, coupled with the images, enrich our dataset and becomes a valuable resource for advanced patent analysis and multimodal research applications.","description_withheld":null,"homepage":"https://github.com/AI4Patents/IMPACT","introduced_date":"2024-06-04","introduced_date_note":null,"introduced_by":null,"license":{"name":"Creative Commons Attribution Share Alike 4.0","url":"https://choosealicense.com/licenses/cc-by-sa-4.0/"},"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Image Classification","url":"/task/image-classification","datasets_with_task":"/datasets/task/image-classification"},{"name":"Image Retrieval","url":"/task/image-retrieval","datasets_with_task":"/datasets/task/image-retrieval"},{"name":"Cross-Modal Retrieval","url":"/task/cross-modal-retrieval","datasets_with_task":"/datasets/task/cross-modal-retrieval"},{"name":"Zero-Shot Cross-Modal Retrieval","url":"/task/zero-shot-cross-modal-retrieval","datasets_with_task":"/datasets/task/zero-shot-cross-modal-retrieval"},{"name":"Image-text Retrieval","url":"/task/image-text-retrieval","datasets_with_task":"/datasets/task/image-text-retrieval"},{"name":"Patent classification","url":"/task/patent-classification","datasets_with_task":"/datasets/task/patent-classification"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["IMPACT Patent"],"data_loaders":[{"repo":"https://github.com/AI4Patents/IMPACT","url":"https://github.com/AI4Patents/IMPACT","frameworks":["pytorch"]}],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/zero-shot-cross-modal-retrieval-on-impact","task":"Zero-Shot Cross-Modal Retrieval","dataset_variant":"IMPACT Patent","rows":1,"metrics":["Image to Text Recall@1"],"first_row_in_archive_order":{"model":"PatentCLIP","paper":"/paper/impact-a-large-scale-integrated-multimodal","metrics":{"Image to Text Recall@1":"3.42"},"code_links":[{"title":"AI4Patents/IMPACT","url":"https://github.com/AI4Patents/IMPACT"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/impact-a-large-scale-integrated-multimodal","title":"IMPACT: A Large-scale Integrated Multimodal Patent Analysis and Creation Dataset for Design Patents","date":"2024-12-10","rows_on_this_dataset":1,"code_links":1,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}