Blog

Notes on reinforcement learning, humanoid agents and character animation — mostly derivations I worked through for my own understanding, written up in case they are useful to someone else.
Is OpenAI's Mysterious RL Algorithm Just an STE Approximation?
Score centering, proposed in a recent arXiv paper as a possible account of OpenAI's post-training algorithm, turns out to be an approximate straight-through estimator that compensates for the Megatron/vLLM train-inference mismatch. A derivation from the information-geometric optimization perspective, and why clipping is the wrong tool here.
I also write in Chinese on Zhihu.