{"url":"/dataset/unarxive","name":"unarXive","full_name":null,"description_markdown":"A scholarly data set with publications’ full-text, annotated in-text citations, and links to metadata.\r\n\r\nThe unarXive data set contains\r\n\r\n* One million papers in plain text\r\n* 63 million citation contexts\r\n* 39 million reference strings\r\n* A citation network of 16 million connections\r\n\r\nThe data is generated from all LaTeX sources on [arXiv](https://arxiv.org/) from 1991–2020/07 and therefore of higher quality than data generated from PDF files.\r\nFurthermore, as all citing papers are available in full text, citation contexts of arbitrary size can be extracted.\r\n\r\nTypical uses of the data set are approaches in\r\n\r\n* Citation recommendation\r\n* Citation context analysis\r\n* Reference string parsing\r\n\r\nThe code for generating the data set is [publicly available](https://github.com/IllDepence/unarXive).","description_withheld":null,"homepage":"https://github.com/IllDepence/unarXive","introduced_date":"2020-03-02","introduced_date_note":null,"introduced_by":{"paper":"/paper/unarxive-a-large-scholarly-data-set-with","title":"unarXive: A Large Scholarly Data Set with Publications' Full-Text, Annotated In-Text Citations, and Links to Metadata","first_author":"Tarek Saier","url":null},"license":null,"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Citation Recommendation","url":"/task/citation-recommendation","datasets_with_task":"/datasets/task/citation-recommendation"},{"name":"Scientific Document Summarization","url":"/task/scientific-article-summarization","datasets_with_task":"/datasets/task/scientific-article-summarization"},{"name":"Citation Intent Classification","url":"/task/citation-intent-classification","datasets_with_task":"/datasets/task/citation-intent-classification"},{"name":"Scientific Concept Extraction","url":"/task/scientific-concept-extraction","datasets_with_task":"/datasets/task/scientific-concept-extraction"},{"name":"Document Embedding","url":"/task/document-embedding","datasets_with_task":"/datasets/task/document-embedding"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["unarXive"],"data_loaders":[],"num_papers_in_archive":9,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}