Homepage · Google Scholar · ORCID
Hi 👋, I'm Yanjun. Third-year PhD candidate at Hong Kong PolyU, joint with EIT Ningbo. Currently a research intern at Microsoft Research Asia, Singapore.
In reinforcement learning, models are trainable. The environments that train them are not.
I'm working on closing that gap.
C3 assigns exact credit in LLM agent teams: when the trace is the state, each message's worth comes from actually continuing the run, not from a prediction. Accuracy Paradox (EMNLP 2024) shows that the most accurate reward model is not the one that trains best: moderately accurate ones train better language models.
On the side: OmniSeek, a self-hosted deep-retrieval engine for AI agents, and Myco, persistent memory infrastructure. Both in the MCP ecosystem.
Also into board games 🎲, jazz 🎷, karaoke 🎤, and old-school anime that never quite went away.
Best reached by email.

