My goal is to build AI agents that can make decisions in complex environments using evaluative feedback, formalized under the reinforcement learning (RL) framework. I currently work on agents for data problems, spanning the end-to-end pipeline: RL training and evaluation, environment building, and agent harness design.
Experiences
- 2025.07 – Now Senior Research Scientist at Databricks AI Research, where my research powers Genie One and Genie Code.
- 2022.08 – 2025.07 Applied Scientist, then Senior Applied Scientist, at Amazon. I worked on RL, LLMs, and agents, building the RL(HF) finetuning behind Amazon Titan models and Rufus.
- 2021.08 – 2022.08 Research Scientist at ByteDance, where I worked on RL for recommendation problems at arguably the largest scale of any recommender system.
- 2021 Ph.D. in Computer Science at Stanford, advised by Emma Brunskill. I worked on the theory and algorithms of reinforcement learning, especially in offline settings. Our work produced some of the early theoretical results on batch RL with function approximation, including the pessimistic value estimates principle, along with real-world applications of batch RL in healthcare and education.
- 2016 B.S. in Machine Intelligence at Peking University.
Selected Papers
For the full list of publications, see my Google Scholar page.
- EMNLP 2025
- ICLR 2025
- ICML 2024
- NeurIPS 2023
- NeurIPS 2020
- UAI 2019
Professional Service
Journal Reviewing: JMLR, IEEE TPAMI, Machine Learning, Artificial Intelligence, Biometrika
Conference Reviewing: NeurIPS, ICLR, ICML, AISTATS, UAI, AAAI