Papers › OGB-LSC: A Large-Scale Challenge for Machine Learning on Graphs

OGB-LSC: A Large-Scale Challenge for Machine Learning on Graphs

17 Mar 2021arXiv:2103.09430archive 2025-07-28

Weihua Hu, Matthias Fey, Hongyu Ren, Maho Nakata, Yuxiao Dong, Jure Leskovec

Enabling effective and efficient machine learning (ML) over large-scale graph data (e.g., graphs with billions of edges) can have a great impact on both industrial and scientific applications. However, existing efforts to advance large-scale graph ML have been largely limited by the lack of a suitable public benchmark. Here we present OGB Large-Scale Challenge (OGB-LSC), a collection of three real-world datasets for facilitating the advancements in large-scale graph ML. The OGB-LSC datasets are orders of magnitude larger than existing ones, covering three core graph learning tasks -- link prediction, graph regression, and node classification. Furthermore, we provide dedicated baseline experiments, scaling up expressive graph ML models to the massive datasets. We show that expressive models significantly outperform simple scalable baselines, indicating an opportunity for dedicated efforts to further improve graph ML at scale. Moreover, OGB-LSC datasets were deployed at ACM KDD Cup 2021 and attracted more than 500 team registrations globally, during which significant performance improvements were made by a variety of innovative techniques. We summarize the common techniques used by the winning solutions and highlight the current best practices in large-scale graph ML. Finally, we describe how we have updated the datasets after the KDD Cup to further facilitate research advances. The OGB-LSC datasets, baseline code, and all the information about the KDD Cup are available at https://ogb.stanford.edu/docs/lsc/ .

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

snap-stanford/ogb officialmentioned in paperpytorchMIT report
graphcore/distributed-kge-poplar mentioned on GitHubpytorchMIT report
graphcore/ogb-lsc-pcqm4mv2 mentioned on GitHubtf report
lars-research/3d-pgt mentioned on GitHubpytorchMIT report
shamim-hussain/egt mentioned on GitHubtfMIT report
dmlc/dgl pytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

BIG-bench Machine LearningGraph LearningGraph RegressionKnowledge GraphsLink PredictionNode Classification

Datasets

Introduced by this paper, per the archive.

OGB-LSCPCQM4Mv2-LSC

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Graph Regression PCQM4M-LSC GIN-virtual Test MAE 14.87 #2 of 11 Archive leaderboard report
Graph Regression PCQM4M-LSC GIN-virtual Validation MAE 0.1396 #2 of 11 Archive leaderboard report
Graph Regression PCQM4M-LSC GCN-Virtual Test MAE 15.79 #3 of 11 Archive leaderboard report
Graph Regression PCQM4M-LSC GCN-Virtual Validation MAE 0.1536 #3 of 11 Archive leaderboard report
Graph Regression PCQM4M-LSC GIN Test MAE 16.78 #4 of 11 Archive leaderboard report
Graph Regression PCQM4M-LSC GCN Test MAE 18.38 #5 of 11 Archive leaderboard report
Graph Regression PCQM4M-LSC GCN Validation MAE 0.1684 #5 of 11 Archive leaderboard report
Graph Regression PCQM4M-LSC MLP-fingerprint Test MAE 20.68 #6 of 11 Archive leaderboard report
Graph Regression PCQM4M-LSC MLP-fingerprint Validation MAE 0.2044 #6 of 11 Archive leaderboard report
Graph Regression PCQM4Mv2-LSC MLP-Fingerprint Test MAE 0.1760 #20 of 20 Archive leaderboard report
Graph Regression PCQM4Mv2-LSC MLP-Fingerprint Validation MAE 0.1753 #20 of 20 Archive leaderboard report
Knowledge Graphs WikiKG90M-LSC TransE-Concat Test MRR 85.48 #1 of 4 Archive leaderboard report
Knowledge Graphs WikiKG90M-LSC TransE-Concat Validation MRR 0.8494 #1 of 4 Archive leaderboard report
Knowledge Graphs WikiKG90M-LSC ComplEx-Concat Test MRR 0.8637 #2 of 4 Archive leaderboard report
Knowledge Graphs WikiKG90M-LSC ComplEx-Concat Validation MRR 0.8425 #2 of 4 Archive leaderboard report
Knowledge Graphs WikiKG90M-LSC ComplEx-RoBERTa Test MRR 0.7186 #3 of 4 Archive leaderboard report
Knowledge Graphs WikiKG90M-LSC ComplEx-RoBERTa Validation MRR 0.7052 #3 of 4 Archive leaderboard report
Knowledge Graphs WikiKG90M-LSC TransE-RoBERTa Test MRR 0.6288 #4 of 4 Archive leaderboard report
Knowledge Graphs WikiKG90M-LSC TransE-RoBERTa Validation MRR 0.6039 #4 of 4 Archive leaderboard report
Node Classification MAG240M-LSC R-GraphSAGE (NS) Test Accuracy 68.94 #1 of 4 Archive leaderboard report
Node Classification MAG240M-LSC GAT (NS) Test Accuracy 66.63 #2 of 4 Archive leaderboard report
Node Classification MAG240M-LSC GraphSAGE (NS) Test Accuracy 66.25 #3 of 4 Archive leaderboard report
Node Classification MAG240M-LSC SIGN Test Accuracy 66.09 #4 of 4 Archive leaderboard report
Node Classification MAG240M-LSC SIGN Validation Accuracy 66.64 #4 of 4 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections