Welcome! 🖐️
My research covers agentic post-training data, infra, recipe and harness, collectively optimized as a single problem.
Selected work (as First Author) >>
🎨 Skill2Env | Reinforcing agents with collective Skills. [Code] [Dataset]
🔳 linex | (WIP) The minimal all-in-one infra for Terminal Agent RL. [Code]
⚡ FlashREINFOCE | Solving instability from async RL policy drifts and token credit mis-assignent. [Paper]
⭐ Polar | The first open Agent RL infra solving Any-harness rollout. [Code] [Paper]
🧠 NanoGPX | Clean collection of modern LLM architectures (RoPE, GQA, RMSNorm, MoE, SSM, etc.) in nanoGPT style. [Code]
🤖 Gentopia & GentPool | An Agent [Framework] & [Platform].
🚀 ReWOO | Token-efficient harness via decoupling reasoning from observation. [Code] [Paper]




