I am a researcher and engineer working on large language models, multimodal AI, and ML systems. I recently graduated from Georgia Tech with dual M.S. degrees in Computer Science (Machine Learning) and Computational Science & Engineering (Applied Mathematics). My work lives at the intersection of model behavior and deployment infrastructure: I identify failure modes in large language models and vision-language models, build methods to fix them, and design the GPU-scale systems needed to serve these models in production.
My research targets the blind spots that standard evaluations miss. I study prestige-sensitive decision revision in LLM pipelines, post-retrieval evidence neglect in multimodal RAG, reward hacking and compute-scaling collapse in RLVR, and self-correction backfire in strong reasoning models. In each case the model passes the benchmark but fails in ways that break real applications. I build audit frameworks and training objectives to close these gaps while preserving overall capability. I also work on distilling calibration signals from large VLMs into compressed models and canonicalized motion priors for few-shot robotic manipulation.
During my internship at GMI Cloud, I optimized Flux-Schnell (12B DiT) inference on H100 GPUs, achieving ~30 images/min at 1 to 2s latency through TensorRT, GPU memory persistence, and kernel-level tuning. I designed the multi-GPU serving architecture on NCCL, built the full production stack from scratch (queueing, heartbeat monitoring, structured logging, GCS integration, safety filtering), and shipped a video super-resolution pipeline (Real-ESRGAN) integrated with Wan2.2 text-to-video generation.
I grew up in China and did my undergraduate degree in Artificial Intelligence at Shandong University. Those four years gave me a strong grounding in mathematics and control theory, and more importantly, taught me how to think across disciplinary boundaries. I have carried that habit with me ever since.
