Agentic RL for Code
Training code agents to improve through execution feedback and trial-and-error in realistic software environments.

Alibaba Qwen · Tsinghua University · Beijing, China
Agentic RL for Code · Long-horizon Coding · User Experience
I’m currently a Researcher at Alibaba Qwen, where I build code agents for complex, long-horizon software engineering tasks through interaction, execution, and trial and error. I also study at Tsinghua University and expect to graduate in summer 2027. Before joining Qwen, I was a Research Intern at Zhipu.AI.
I work on code agents that learn from execution, operate over long horizons, and collaborate effectively with people.
Training code agents to improve through execution feedback and trial-and-error in realistic software environments.
Building agents that plan, navigate large codebases, and complete multi-step software tasks across extended trajectories.
Designing code agents that understand user intent, communicate clearly, and make complex engineering work more effective and accessible.
Researcher · Code Agents & LLM Post-training
Research Intern · Code Agents & LLM Post-training
Research Intern · Search Agents & LLM Post-training
Research Intern · Recommender Systems
Tsinghua University
Beijing Municipality
Three consecutive years · Top 1%