Datasets › Arxiv HEP-TH citation graph

Arxiv HEP-TH citation graph

archive 2025-07-28

Arxiv HEP-TH (high energy physics theory) citation graph is from the e-print arXiv and covers all the citations within a dataset of 27,770 papers with 352,807 edges. If a paper i cites paper j, the graph contains a directed edge from i to j. If a paper cites, or is cited by, a paper outside the dataset, the graph does not contain any information about this. The data covers papers in the period from January 1993 to April 2003 (124 months).

Source: https://snap.stanford.edu/data/cit-HepTh.html

Benchmarks archive 2025-07-28

All 5 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Text Summarization Arxiv HEP-TH citation graph Top Down Transformer (AdaPool) (464M) ROUGE-1 50.95 Long Document Summarization with Top-down and Bottom-up Inference lllyasviel/framepack 28 Compare
Topic Models Arxiv HEP-TH citation graph JoSH MACC 83.24 Hierarchical Topic Mining via Joint Spherical Tree and... yumeng5/JoSH 2 Compare
Document Summarization Arxiv HEP-TH citation graph DeepPyramidion ROUGE-1 47.15 Sparsifying Transformer Models with Trainable... — 1 Compare
Language Modelling Arxiv HEP-TH citation graph Gopher BPB 0.662 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
Text Classification Arxiv HEP-TH citation graph BigBird Accuracy 92.31 Big Bird: Transformers for Longer Sequences huggingface/transformers +13 1 Compare

Papers archive 2025-07-28

24 shown of 24 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 35. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Segmented Recurrent Transformer: An Efficient Sequence-to-Sequence Model 1 1 24 May 2023 not harvested
Document Summarization with Text Segmentation 0 2 20 Jan 2023 not harvested
Toward Unifying Text Segmentation and Long Document Summarization 1 2 28 Oct 2022 not harvested
Adapting Pretrained Text-to-Text Models for Long Text Sequences 1 1 21 Sep 2022 not harvested
Investigating Efficiently Extending Transformers for Long Input Summarization 2 1 8 Aug 2022 not harvested
Factorizing Content and Budget Decisions in Abstractive Summarization of Long Documents 1 1 25 May 2022 not harvested
GenCompareSum: a hybrid unsupervised summarization method using salience 1 1 1 May 2022 not harvested
HiStruct+: Improving Extractive Text Summarization with Hierarchical Structure Information 0 1 17 Mar 2022 not harvested
Long Document Summarization with Top-down and Bottom-up Inference 1 1 15 Mar 2022 not harvested
LongT5: Efficient Text-To-Text Transformer for Long Sequences 4 1 15 Dec 2021 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
Scaling Language Models: Methods, Analysis & Insights from Training Gopher 3 1 8 Dec 2021 not harvested
Sparsifying Transformer Models with Trainable Representation Pooling 0 3 16 Nov 2021 not harvested
MemSum: Extractive Summarization of Long Documents Using Multi-Step Episodic Markov Decision Processes 1 1 19 Jul 2021 not harvested
Hierarchical Learning for Generation with Long Source Sequences 0 1 15 Apr 2021 not harvested
Systematically Exploring Redundancy Reduction in Summarizing Long Documents 1 2 30 Nov 2020 not harvested
Big Bird: Transformers for Longer Sequences 14 1 28 Jul 2020 ran 10 of 15 samples (5 unverified; 11 pointer-only for licence)
Hierarchical Topic Mining via Joint Spherical Tree and Text Embedding 1 1 18 Jul 2020 not harvested
A Divide-and-Conquer Approach to the Summarization of Long Documents 1 3 13 Apr 2020 not harvested
PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization 19 1 18 Dec 2019 ran 1 of 19 samples (18 unverified)
Extractive Summarization of Long Documents by Combining Global and Local Context 1 1 17 Sep 2019 not harvested
On Extractive and Abstractive Neural Document Summarization with Transformer Language Models 1 3 7 Sep 2019 ran 0 of 6 samples (6 unverified)
TopicEq: A Joint Topic and Mathematical Equation Model for Scientific Texts 0 1 16 Feb 2019 not harvested
A Discourse-Aware Attention Model for Abstractive Summarization of Long Documents 3 1 16 Apr 2018 ran 1 of 14 samples (13 unverified)
Get To The Point: Summarization with Pointer-Generator Networks 39 1 14 Apr 2017 ran 30 of 64 samples (34 unverified; 44 pointer-only for licence)

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • arXiv-AstroPh 3-clique
  • arXiv-GrQc 4-clique
  • Arxiv HEP-TH citation graph
  • arXiv-Long Val
  • arXiv-Long Test

5 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections