I am a Ph.D. student in Computer Science at the University of Southern California, advised by Prof. Yue Zhao in the FORTIS Lab.
Frontier Model Evaluation and Post-Training. My research focuses on discovering the capability boundaries of frontier AI systems and turning those boundaries into learning opportunities. I build challenging evaluations and interactive environments that systematically elicit model failures, and develop automated methods to identify what increasingly capable models still cannot reliably do. My broader goal is to convert these failures into scalable supervision, training data, and feedback that enable models to acquire new, generalizable capabilities.
AI, Science, and Society. I study how increasingly capable AI systems can participate in scientific discovery and other forms of complex intellectual work, and how humans can reliably verify and build upon the knowledge they produce. As AI-generated ideas and results grow in complexity, producing useful outputs is no longer enough: we need methods for evaluating their validity, exposing their assumptions and uncertainties, and making them auditable and interpretable even when their full reasoning exceeds human understanding. More broadly, I am interested in how humans can retain meaningful agency in a world where AI increasingly operates beyond our cognitive reach—deciding what to trust, what to pursue, and how to translate increasingly superhuman capabilities into knowledge and action that serve human goals.
Previously at WeChat AI, Alibaba, MBZUAI, the MINE Lab at Notre Dame, MIT CTL, Microsoft, and Sichuan University.
News
I’m joining USC in Fall 2026 to start my CS Ph.D. study! Fight On!
I’m joining WeChat AI as a Research Intern!
Every grand ambition starts with a humble beginning.
每一个远大的理想都有个微不足道的开始。
Latest Preprints
AlgoWorlds: Benchmarking Tool Use for Global Optimization in Algorithmic Worlds
Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias
GTA: Graph Theory Agent and Benchmark for Algorithmic Graph Reasoning with LLMs
Selected Work
SocialMaze: A Benchmark for Evaluating and Enhancing Social Reasoning in Large Language Models in Complex Social Environments
Cross-Lingual Pitfalls: Automatic Probing Cross-Lingual Weakness of Multilingual Large Language Models
Adaptive Distraction: Probing LLM Contextual Robustness with Automated Tree Search
DyFlow: Dynamic Workflow Framework for Agentic Reasoning
Honors and Awards
A competitive fellowship top-off awarded to selected Ph.D. students in the USC Viterbi School of Engineering.
One of 100 junior and senior undergraduates selected nationwide for academics, research, and service. Invited speaker and session host at CNCC; team placed 3rd nationwide in the CCSP Team Contest.
Gold medal in the ICPC, the world's premier collegiate algorithmic programming competition.
Silver medal in China's national informatics olympiad for secondary-school students, organized by the China Computer Federation (CCF).
Internships
Reviewing Service
AAAI 2027; COLM 2026; ICLR 2026; NeurIPS 2026; NLPCC 2025.
February and May 2025 cycles.
GroundLM and REALM at EMNLP 2026; Social Simulation with LLMs at COLM 2025.
Miscellaneous
Beyond research and competitive programming, I am a passionate enthusiast of films, games, and visual novels. I am especially fond of the works of Tanaka Romeo (田中ロミオ), particularly CROSS†CHANNEL and Rewrite. As time passes, this list has grown shorter: what remains below are the works that have left the deepest and most vivid traces in my memory.
- 蓝宝石般的被害妄想少女
- 辯護人
- 大明王朝1566
- 逆境無頼カイジ
- 銀と金
- 家族計画
- 終のステラ
- 加奈 〜いもうと〜
- サクラノ詩
- AIR
- Angel Beats!
- Summer Pockets
- Planetarian
- WHITE ALBUM2
- STEINS;GATE
- Danganronpa 2


MINE Lab at Notre Dame
Microsoft
Sichuan University