Blog
Notes on reinforcement learning, humanoid agents and character animation — mostly derivations I worked through for my own understanding, written up in case they are useful to someone else.
September 20, 2026
Is OpenAI's Mysterious RL Algorithm Just an STE Approximation?
Score centering, proposed in a recent arXiv paper as a possible account of OpenAI's post-training algorithm, turns out to be an approximate straight-through estimator that compensates for the Megatron/vLLM train-inference mismatch. A derivation from the information-geometric optimization perspective, and why clipping is the wrong tool here.
I also write in Chinese on
Zhihu.