whoami
What I am working on and thinking about (ver. 2026):
- self-evolving agentic systems — agents that revise their own policies
- reasoning in latent space: thinking without decoding to tokens
- hybrid zeroth/first-order policy learning, differentiating everything
- world models learned well enough to plan against
I work on agentic post-training at AWS: supervised fine-tuning and reinforcement learning of trillion-parameter mixture-of-experts models, on production-scale clusters. I have published at ICLR, ICML, NeurIPS, EMNLP, AAAI, CVPR, etc. The newest work is always on Google Scholar.
Before that: a PhD in Electrical & Computer Engineering at Yale (2024), advised by Leandros Tassiulas, and a B.Eng in Optoelectronics at Zhejiang University (2019). Along the way, internships at HKU, CUHK, UIUC, IBM Research and Nokia Bell Labs.