Papers › HOList: An Environment for Machine Learning of Higher-Order Theorem Proving

HOList: An Environment for Machine Learning of Higher-Order Theorem Proving

5 Apr 2019arXiv:1904.03241archive 2025-07-28

Kshitij Bansal, Sarah M. Loos, Markus N. Rabe, Christian Szegedy, Stewart Wilcox

We present an environment, benchmark, and deep learning driven automated theorem prover for higher-order logic. Higher-order interactive theorem provers enable the formalization of arbitrary mathematical theories and thereby present an interesting, open-ended challenge for deep learning. We provide an open-source framework based on the HOL Light theorem prover that can be used as a reinforcement learning environment. HOL Light comes with a broad coverage of basic mathematical theorems on calculus and the formal proof of the Kepler conjecture, from which we derive a challenging benchmark for automated reasoning. We also present a deep reinforcement learning driven automated theorem prover, DeepHOL, with strong initial results on this benchmark.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

Kerram/Deephol-Bert-Zpp mentioned on GitHubtf report
Kerram/holist-train mentioned on GitHubtf report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Automated Theorem ProvingBIG-bench Machine LearningDeep LearningDeep Reinforcement LearningReinforcement LearningReinforcement Learning (RL)reinforcement-learning

Datasets

Introduced by this paper, per the archive.

HOList

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Automated Theorem Proving HOList benchmark Tactic Dependent Loop Percentage correct 38.88 #2 of 4 Archive leaderboard report
Automated Theorem Proving HOList benchmark Deeper Wider WaveNet Percentage correct 32.65 #4 of 4 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections