👋 About Me

Hi! I am a second-year PhD student at Tsinghua University, majoring in Computer Science and Technology. I am a member of THUNLP, advised by Prof. Zhiyuan Liu. I received my bachelor’s degree with honors from Tsinghua University in June 2024. My research interests lie in natural language processing, with a focus on alignment, reinforcement learning, and self-evolving language models, including both benchmark construction and method development.

🌟 News

📝 Publications

(* denotes equal/core contribution, denotes project lead, indicates corresponding author.)

Google Scholar · 2300+ citations

📖 Educations

💼 Experience

  • 2026.03 - present, ModelBest (面壁智能), Beijing. Research Intern, Forward-Four Program (前进四计划). Working with Postdoc Chaojun Xiao.
    • Post-training of the MiniCPM4 & MiniCPM5 series: SFT, RL, and on-policy distillation (OPD).
    • AutoSFT: a coding-agent that autonomously searches SFT data recipes; the SFT data engine of the pipeline.
    • RL: a minimal, stable RL recipe landed as MiniCPM5’s math & reasoning RL; diagnosing and fixing training collapse.
    • OPD: co-developed the OPD recipe, integrated into MiniCPM5 as the cross-domain model-merging mechanism.

🎖 Honors and Awards

  • Qingyuan InnoVibe 2026 (青源最受瞩目学术新星, 25 Winners Nationwide), BAAI. 2026.06
  • ICML 2026 Gold Reviewer (Top 25%). 2026.05
  • Comprehensive Merit Scholarship of Tsinghua for 2024-2025, Dept. of CST. 2025.12
  • Outstanding Graduate Award, Beijing Municipal Education Commission. 2024.06
  • Outstanding Paper Award for Diploma Project, Tsinghua University. 2024.06
  • Comprehensive Merit Scholarship of Tsinghua for 2022-2023, Dept. of CST. 2023.10
  • Comprehensive Merit Scholarship of Tsinghua for 2021-2022, Dept. of CST (Top 1). 2022.10
  • Third Prize in THU Challenge Cup Academic Competition, Tsinghua University. 2022.04
  • Comprehensive Merit Scholarship of Tsinghua for 2020-2021, Dept. of CST. 2021.10
  • Second Prize in Freshmen Scholarship, Tsinghua University. 2020.09

💬 Invited Talks

  • Three Boundaries for Scalable Reinforcement Learning. Qingyuan InnoVibe 2026 in BAAI. 2026.06
  • AMA (Ask Me Anything) for Rethinking OPD. QingKeAI. 2026.05
  • Towards Scalable Reinforcement Learning for LLMs. BAAI. NICE. 2026.05
  • How Far Can Unsupervised RLVR Scale LLM Training? AI TIME. Synced. QingKeAI. 2026.04
  • JustRL: Scaling a 1.5B LLM with a Simple RL Recipe. QingKeAI. 2026.02
  • The Right Time Matters: Data Arrangement Affects Zero-Shot Generalization in Instruction Tuning. Alibaba Security. 2025.05
  • Tell me more! towards implicit user intention understanding of language model driven agents. Wiztalk. 2024.08

🛠️ Services

  • Conference Reviewer: NeurIPS (2024 - 2025), ICLR (2025 - 2026), ICML (2025 - 2026), ACL ARR (2024 - 2026), COLM (2025 - 2026), COLM SCALR Workshop (2025), AAAI (2026), AISTATS (2025 - 2026), ICCV (2025)