Skip to content
Leqi Zheng (郑乐祺)
Leqi Zheng

Leqi Zheng (郑乐祺)

Alibaba Qwen · Tsinghua University · Beijing, China

Agentic RL for Code · Long-horizon Coding · User Experience

I’m currently a Researcher at Alibaba Qwen, where I build code agents for complex, long-horizon software engineering tasks through interaction, execution, and trial and error. I also study at Tsinghua University and expect to graduate in summer 2027. Before joining Qwen, I was a Research Intern at Zhipu.AI.

Research Interests

I work on code agents that learn from execution, operate over long horizons, and collaborate effectively with people.

Agentic RL for Code

Training code agents to improve through execution feedback and trial-and-error in realistic software environments.

Reinforcement learning · Execution feedback

Long-horizon Coding

Building agents that plan, navigate large codebases, and complete multi-step software tasks across extended trajectories.

Repository-level coding · Long-horizon tasks

User Experience

Designing code agents that understand user intent, communicate clearly, and make complex engineering work more effective and accessible.

Human-agent interaction · User-centered design

Model Contributions

Selected Publications

View all on Google Scholar →
Under Review

AdaMem: Adaptive User-Centric Memory for Long-Horizon Dialogue Agents

Shannan Yan*, Jingchen Ni*, Leqi Zheng*, Jiajun Zhang, Peixi Wu, Dacheng Yin, Jing Lyu, Chun Yuan, Fengyun Rao

Under Review

Audio-Visual World Models: Grounding Multisensory Imagination for Embodied Agents

Jiahua Wang*, Leqi Zheng*, Jialong Wu, Yaoxin Mao, Shijie Cheng

Under Review

Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation

Yuying Li*, Leqi Zheng*, Yongzi Yu, et al.

EMNLP 2026

SciLENS: RL-Driven Autonomous Agents for Scientific Localized Evidence Navigation and Synthesis

Leqi Zheng*, Jinbo Su*, Yuying Li*, et al.

EMNLP 2026

Gradients Know What Outcomes Don’t: Unlocking Reinforcement Learning for LLM Reasoning with Gradient-Aligned Rewards

Leqi Zheng*, Jinbo Su*, Fang Niu*, et al.

The Web Conf 2026

What Should I Cite? A RAG Benchmark for Academic Citation Prediction

Leqi Zheng*, Jiajun Zhang*, Canzhi Chen*, Chaokun Wang, Hongwei Li, Yuying Li, et al.

NeurIPS 2025

Negative Feedback Really Matters: Signed Dual-Channel Graph Contrastive Learning Framework for Recommendation

Leqi Zheng, Chaokun Wang, Zixin Song, Cheng Wu, Shannan Yan, Jiajun Zhang, Ziyang Liu

EMNLP 2025

LAGCL4Rec: When LLMs Activate Interactions Potential in Graph Contrastive Learning for Recommendation

Leqi Zheng, Chaokun Wang, Canzhi Chen, Jiajun Zhang, Cheng Wu, Zixin Song, et al.

SIGIR 2025

Balancing Self-Presentation and Self-Hiding for Exposure-Aware Recommendation Based on Graph Contrastive Learning

Leqi Zheng, Chaokun Wang, Ziyang Liu, Canzhi Chen, Cheng Wu, Hongwei Li

CVPR 2026

Learning Cross-View Object Correspondence via Cycle-Consistent Mask Prediction

Shannan Yan, Leqi Zheng, Keyu Lv, Jingchen Ni, Hongyang Wei, Jiajun Zhang, et al.

Experience

Alibaba Qwen

Researcher · Code Agents & LLM Post-training

Zhipu.AI

Research Intern · Code Agents & LLM Post-training

Ant Group

Research Intern · Search Agents & LLM Post-training

JD.com

Research Intern · Recommender Systems

Honors & Awards

Yang Huiyan Scholarship

Tsinghua University

Excellent Higher Education Graduate

Beijing Municipality

China National Scholarship

Three consecutive years · Top 1%