About

Hi! I am a PhD candidate in the Siebel School of Computing and Data Science, UIUC, where I am fortunate to be advised by Professor Tong Zhang. Prior to that, I did my undergraduate in the School of Data Science at Fudan University, under the valued supervision of Professor Luo Luo.

I am broadly interested in machine learning and optimization, with a focus on the intersection of these fields.

Optimization and training foundations

Better understanding the effectiveness of practical training algorithms, from both theoretical analysis and empirical observation, and a viewpoint beyond purely loss comparison.

Algorithms design

Designing and implementing a more effective training process applied to large foundation model pretraining and finetuning, including optimizers, model architectures, training settings, and their interactions.

Efficiency of large foundation models

improving foundation model inference efficiency through architecture modifications and better inference pipelines like speculative decoding and beyond.

News

  • Our paper StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models published a new version on arXiv. We added more low-precision training experiments (FP8 and FP4), showing the effectiveness of StoSignSGD under low-precision settings like physical AI.

  • Our new paper, Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less, is now available on arXiv. We presented the optimizer-model consistency phenomenon: full finetuning with the same family of optimizer as pretraining achieves the best learning-forgetting tradeoff compared to other optimizers and even LoRA (with different optimizers), through a comprehensive Pareto frontier comparison taking learning rates into consideration.

Selected Publications

* denotes first authors. Please refer to Google Scholar for a full list.

  • Preprint2026

    Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less

    Yuxing Liu, Jianyu Wang, and Tong Zhang.

  • Preprint2026

    StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models

    Dingzhi Yu*, Rui Pan*, Yuxing Liu*, Difan Zou, and Tong Zhang.

  • Preprint2025

    Unbiased Gradient Low-Rank Projection

    Rui Pan*, Yang Luo*, Yuxing Liu*, You Yang, and Tong Zhang.

  • Preprint2025

    Theoretical Analysis on How Learning Rate Warmup Accelerates Convergence

    Yuxing Liu*, Yuze Ge*, Rui Pan, Kang An, and Tong Zhang.

  • NeurIPS2025

    ASGO: Adaptive Structured Gradient Optimization

    Kang An*, Yuxing Liu*, Rui Pan, Yi Ren, Shiqian Ma, Donald Goldfarb, and Tong Zhang.

  • ICLR2025

    Adagrad Under Anisotropic Smoothness

    Yuxing Liu*, Rui Pan*, and Tong Zhang.

  • ICLR2024

    Accelerated Convergence of Stochastic Heavy Ball Method under Anisotropic Gradient Noise

    Rui Pan*, Yuxing Liu*, Xiaoyu Wang, and Tong Zhang.

Misc.

In my spare time, I love to play badminton and go swimming. I also play Go, an ancient strategy board game. The photo here was taken when I was playing it.