Datasets › Pubmed

Pubmed

Introduced in Collective Classification in Network Data1 Jan 2008 archive 2025-07-28

The PubMed dataset consists of 19717 scientific publications from PubMed database pertaining to diabetes classified into one of three classes. The citation network consists of 44338 links. Each publication in the dataset is described by a TF/IDF weighted word vector from a dictionary which consists of 500 unique words.

Source: https://linqs.soe.ucsc.edu/data

Benchmarks archive 2025-07-28

All 19 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Node Classification Pubmed NCGCN Accuracy 91.64 ± 0.53 Clarify Confused Nodes via Separated Learning GISec-Team/NCGNN 70 Compare
Node Classification PubMed with Public Split: fixed 20 nodes per class OGC Accuracy 83.4% From Cluster Assumption to Graph Convolution:... zhengwang100/ogc_ggcm 37 Compare
Text Summarization Pubmed Top Down Transformer (AdaPool) (464M) ROUGE-1 51.05 Long Document Summarization with Top-down and Bottom-up Inference lllyasviel/framepack 29 Compare
Node Classification PubMed (0.03%) VCHN Accuracy 71.8% View-Consistent Heterogeneous Network on Graphs With Few... kunzhan/VCHN 14 Compare
Node Classification PubMed (0.05%) VCHN Accuracy 74.3% View-Consistent Heterogeneous Network on Graphs With Few... kunzhan/VCHN 14 Compare
Node Classification PubMed (0.1%) Truncated Krylov Accuracy 77.21% Break the Ceiling: Stronger Multi-scale Deep Graph... PwnerHarry/Stronger_GCN 14 Compare
Link Prediction Pubmed AUC 200% An Effective Graph Learning based Approach for Temporal... im0qianqian/WSDM2022TGP-AntGraph 13 Compare
Unsupervised Extractive Summarization Pubmed HipoRank ROUGE-1 43.58 Discourse-Aware Unsupervised Summarization of Long... mirandrom/HipoRank 9 Compare
Graph Clustering Pubmed R-GMM-VGAE ACC 74.0 Rethinking Graph Auto-Encoder Models for Attributed... nairouz/R-GAE 7 Compare
Node Classification Pubmed Full-supervised GraphSAGE+DropEdge Accuracy 91.70% DropEdge: Towards Deep Graph Convolutional Networks on... GraphSAINT/GraphSAINT +6 7 Compare
Link Prediction Pubmed (biased evaluation) GraphStar (double weight on positive examples) AP 98.64 Graph Star Net for Generalized Multi-Task Learning graph-star-team/graph_star 2 Compare
Sentence Classification PubMed 20k RCT Hierarchical Neural Networks F1 92.60 Hierarchical Neural Networks for Sequential Sentence... jind11/HSLN-Joint-Sentence-Classification 2 Compare
Community Detection Pubmed CDNMF ACC 0.6653 Contrastive Deep Nonnegative Matrix Factorization for... 6lyc/cdnmf 1 Compare
Graph Classification Pubmed Fea2Fea-s3 Test Accuracy 78.5 Fea2Fea: Exploring Structural Feature Correlations via... JIAQING-XIE/Fea2Fea 1 Compare
Language Modelling PubMed Central Gopher BPB 0.525 Scaling Language Models: Methods, Analysis & Insights... allenai/dolma +2 1 Compare
Link Prediction Pubmed (nonstandard variant) GLACE AP 97.49 Gaussian Embedding of Large-scale Attributed Graphs bhagya-hettige/GLACE 1 Compare
Node Classification on Non-Homophilic (Heterophilic) Graphs Pubmed GAT F1-Score 59.89 ± 4.12 Graph Attention Networks labmlai/annotated_deep_learning_paper_implementations +92 1 Compare
Node Classification Pubmed: fixed 20 node per class SDSS-APPNP Accuracy 82.72 Multi-task Self-distillation for Graph-based... — 1 Compare
Node Classification Pubmed random partition GraphMix (GCN) Accuracy 80.72 ± 1.08 GraphMix: Improved Training of GNNs for Semi-Supervised Learning vikasverma1077/GraphMix 1 Compare

Papers archive 2025-07-28

30 shown of 133 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 1,236. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
FIT-GNN: Faster Inference Time for GNNs Using Coarsening 1 1 19 Oct 2024 not harvested
Scaling Up Summarization: Leveraging Large Language Models for Long Text Extractive Summarization 0 1 28 Aug 2024 not harvested
Classic GNNs are Strong Baselines: Reassessing GNNs for Node Classification 1 1 13 Jun 2024 ran 2 of 9 samples (7 unverified)
GSCAN: Graph Stability Clustering for Applications With Noise Using Edge-Aware Excess-of-Mass 1 1 17 Apr 2024 not harvested
Mitigating Degree Biases in Message Passing Mechanism by Utilizing Community Structures 1 1 28 Dec 2023 not harvested
Contrastive Deep Nonnegative Matrix Factorization for Community Detection 1 1 4 Nov 2023 not harvested
From Cluster Assumption to Graph Convolution: Graph-based Semi-Supervised Learning Revisited 1 2 24 Sep 2023 not harvested
The Split Matters: Flat Minima Methods for Improving the Performance of GNNs 1 2 15 Jun 2023 not harvested
Clarify Confused Nodes via Separated Learning 1 2 4 Jun 2023 not harvested
Graph Entropy Minimization for Semi-supervised Node Classification 1 2 31 May 2023 not harvested
NESS: Node Embeddings from Static SubGraphs 1 1 15 Mar 2023 not harvested
Multi-Mask Aggregators for Graph Neural Networks 1 1 24 Nov 2022 not harvested
GoSum: Extractive Summarization of Long Documents by Reinforcement Learning and Graph Organized discourse state 1 1 18 Nov 2022 not harvested
Toward Unifying Text Segmentation and Long Document Summarization 1 2 28 Oct 2022 not harvested
Beyond Homophily with Graph Echo State Networks 0 1 27 Oct 2022 not harvested
Adapting Pretrained Text-to-Text Models for Long Text Sequences 1 1 21 Sep 2022 not harvested
GRETEL: Graph Contrastive Topic Enhanced Language Model for Long Document Extractive Summarization 0 1 21 Aug 2022 not harvested
Beyond Homophily: Structure-aware Path Aggregation Graph Neural Network 1 1 20 Jul 2022 not harvested
TREE-G: Decision Trees Contesting Graph Neural Networks 1 1 6 Jul 2022 ran 0 of 10 samples (10 unverified)
DiffWire: Inductive Graph Rewiring via the Lovász Bound 2 2 15 Jun 2022 not harvested
Factorizing Content and Budget Decisions in Abstractive Summarization of Long Documents 1 1 25 May 2022 not harvested
GenCompareSum: a hybrid unsupervised summarization method using salience 1 1 1 May 2022 not harvested
Inferring from References with Differences for Semi-Supervised Node Classification on Graphs 1 1 11 Apr 2022 not harvested
How to Find Your Friendly Neighborhood: Graph Attention Design with Self-Supervision 2 1 11 Apr 2022 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
View-Consistent Heterogeneous Network on Graphs With Few Labeled Nodes 1 3 17 Mar 2022 not harvested
HiStruct+: Improving Extractive Text Summarization with Hierarchical Structure Information 0 1 17 Mar 2022 not harvested
Long Document Summarization with Top-down and Bottom-up Inference 1 1 15 Mar 2022 not harvested
Graph Representation Learning Beyond Node and Homophily 1 1 3 Mar 2022 not harvested
An Effective Graph Learning based Approach for Temporal Link Prediction: The First Place of WSDM Cup 2022 1 1 1 Mar 2022 not harvested
LongT5: Efficient Text-To-Text Transformer for Long Sequences 4 1 15 Dec 2021 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)

The full list of 133 is in the JSON twin.

Dataset loaders archive 2025-07-28

9 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Unknown

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • PubMed-Long Val
  • PubMed-Long Test
  • Pubmed: fixed 20 node per class
  • PubMed Central
  • PubMed (60%/20%/20% random splits)
  • PubMed (48%/32%/20% fixed splits)
  • keyword_pubmed_dataset
  • enoriega/keyword_pubmed sentence
  • Pubmed (weighted evaluation)
  • Pubmed random partition
  • Pubmed (nonstandard variant)
  • Pubmed (biased evaluation)
  • PubMed 20k RCT
  • Pubmed Full-supervised
  • Pubmed
  • PubMed with Public Split: fixed 20 nodes per class
  • PubMed (0.1%)
  • PubMed (0.05%)
  • PubMed (0.03%)

19 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections