{"url":"/dataset/semopenalex","name":"SemOpenAlex","full_name":null,"description_markdown":"SemOpenAlex is an extensive RDF knowledge graph that contains over 26 billion triples about scientific publications and their associated entities, such as authors, institutions, journals, and concepts. \r\n* SemOpenAlex is licensed under CC0, providing free and open access to the data. \r\n* We offer the data through multiple channels, including RDF dump files, a SPARQL endpoint, and as a data source in the Linked Open Data cloud, complete with resolvable URIs and links to other data sources (ISNI, DOI, ORCID, ROR, Scopus, DOAJ, Wikidata, \r\n* Moreover, we provide embeddings for knowledge graph entities using high-performance computing. \r\n\r\nSemOpenAlex enables a broad range of use-case scenarios, such as \r\n* exploratory semantic search via our website,\r\n* large-scale scientific impact quantification, \r\n* other forms of scholarly big data analytics within and across scientific disciplines.\r\n* enables academic recommender systems, such as recommending collaborators, publications, and venues, including explainability capabilities. \r\n* can serve for RDF query optimization benchmarks, \r\n* creating scholarly knowledge-guided language models, \r\n* as a hub for semantic scientific publishing.","description_withheld":null,"homepage":"https://semopenalex.org/","introduced_date":"2023-08-07","introduced_date_note":null,"introduced_by":{"paper":"/paper/semopenalex-the-scientific-landscape-in-26","title":"SemOpenAlex: The Scientific Landscape in 26 Billion RDF Triples","first_author":"Michael Färber","url":null},"license":{"name":"CC0","url":"https://creativecommons.org/publicdomain/zero/1.0/legalcode"},"modalities":[{"name":"Graphs","url":"/datasets/modality/graphs"}],"tasks":[{"name":"Knowledge Graphs","url":"/task/knowledge-graphs","datasets_with_task":"/datasets/task/knowledge-graphs"},{"name":"Citation Recommendation","url":"/task/citation-recommendation","datasets_with_task":"/datasets/task/citation-recommendation"},{"name":"Scientific Document Summarization","url":"/task/scientific-article-summarization","datasets_with_task":"/datasets/task/scientific-article-summarization"},{"name":"Scientific Concept Extraction","url":"/task/scientific-concept-extraction","datasets_with_task":"/datasets/task/scientific-concept-extraction"},{"name":"Science Question Answering","url":"/task/science-question-answering","datasets_with_task":"/datasets/task/science-question-answering"},{"name":"Joint Entity and Relation Extraction on Scientific Data","url":"/task/joint-entity-and-relation-extraction-on","datasets_with_task":"/datasets/task/joint-entity-and-relation-extraction-on"},{"name":"Knowledge-Aware Recommendation","url":"/task/knowledge-aware-recommendation","datasets_with_task":"/datasets/task/knowledge-aware-recommendation"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["SemOpenAlex"],"data_loaders":[],"num_papers_in_archive":6,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-25T09:33:49+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}