Abstract: Gradient descent is one of the most fundamental algorithms in optimization and machine learning. For smooth convex optimization, classical theory gives a convergence rate of (O(T^{-1})). A series of recent works has shown that, without introducing momentum or changing the basic algorithmic structure, carefully designed stepsize schedules alone can accelerate standard gradient descent to approximately (O(T^{-1.2716})). This naturally raises a fundamental question: Can gradient descent achieve the optimal (O(T^{-2})) convergence rate of Nesterov's accelerated methods through stepsize selection alone?
In this talk, I will present our recent results on this question. For gradient descent with predetermined nonnegative stepsize sequences, we prove a last-iterate convergence lower bound of (\Omega(T^{-1.9319})). This result shows that, regardless of how such stepsize sequences are designed, stepsize scheduling alone cannot accelerate standard gradient descent to the optimal (O(T^{-2})) rate.
The proof of this result was developed by GPT-5.6 Sol Pro under the author's guidance. The talk will also briefly discuss how large language models can participate in mathematical research and contribute to the discovery and development of rigorous proofs.
Bio: Jianhao Ma is an Assistant Professor in the Department of Industrial Engineering at Tsinghua University. His research focuses on optimization theory for machine learning. He received his Ph.D. in Industrial and Operations Engineering from the University of Michigan in 2025, where he was advised by Professor Salar Fattahi. He subsequently conducted postdoctoral research in the Department of Statistics and Data Science at the University of Pennsylvania, working with Professor Yuxin Chen. His research has appeared in journals and conferences including Journal of Machine Learning Research (JMLR), Mathematics of Operations Research, Transactions on Machine Learning Research (TMLR), COLT, ICML, ICLR, and NeurIPS.