Papers › Tangent-Space Gradient Optimization of Tensor Network for Machine Learning

Tangent-Space Gradient Optimization of Tensor Network for Machine Learning

10 Jan 2020arXiv:2001.04029archive 2025-07-28

Zheng-Zhi Sun, Shi-Ju Ran, Gang Su

The gradient-based optimization method for deep machine learning models suffers from gradient vanishing and exploding problems, particularly when the computational graph becomes deep. In this work, we propose the tangent-space gradient optimization (TSGO) for the probabilistic models to keep the gradients from vanishing or exploding. The central idea is to guarantee the orthogonality between the variational parameters and the gradients. The optimization is then implemented by rotating parameter vector towards the direction of gradient. We explain and testify TSGO in tensor network (TN) machine learning, where the TN describes the joint probability distribution as a normalized state | ψ⟩ in Hilbert space. We show that the gradient can be restricted in the tangent space of ⟨ψ.| ψ⟩= 1 hyper-sphere. Instead of additional adaptive methods to control the learning rate in deep learning, the learning rate of TSGO is naturally determined by the angle θ as η= tanθ. Our numerical results reveal better convergence of TSGO in comparison to the off-the-shelf Adam.

PaperPDFCode

Code

crazybigcat/TSGO mentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

BIG-bench Machine Learning

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

Adam

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections