Papers › Machine Learning from Archives Gauquelin
Machine Learning from Archives Gauquelin
Anonymous
We apply machine learning methods to the data from Archives Gauquelin in an attempt to build a binary classifier able to distinguish between outstanding scientists and sports champions using only astronomical factors derived from their natal data. We apply a special splitting into training, validation and testing sets, and a set of three combined features, each of which combines dozens of elementary astronomical features. Our null hypothesis is that accuracy on Testing sets must be 0.5 if the Training sets contain the same number of group A and group B representatives born on each year, that is, if yearly frequencies are equal. Our external Testing sets contain only persons born later than those in Archives Gauquelin. Logistic Regression is our primary method, and Random Forest an alternative. All data and implementations of our algorithms are available from a public repository on GitLab.com.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
No leaderboard rows for this paper in the archive.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections