Skip to content
View yupengtang's full-sized avatar

Block or report yupengtang

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
yupengtang/README.md

Hi, I'm Yupeng Tang

I am a researcher and engineer working on large language models, multimodal AI, and ML systems. I recently graduated from Georgia Tech with dual M.S. degrees in Computer Science (Machine Learning) and Computational Science & Engineering (Applied Mathematics). My work lives at the intersection of model behavior and deployment infrastructure: I identify failure modes in large language models and vision-language models, build methods to fix them, and design the GPU-scale systems needed to serve these models in production.

My research targets the blind spots that standard evaluations miss. I study prestige-sensitive decision revision in LLM pipelines, post-retrieval evidence neglect in multimodal RAG, reward hacking and compute-scaling collapse in RLVR, and self-correction backfire in strong reasoning models. In each case the model passes the benchmark but fails in ways that break real applications. I build audit frameworks and training objectives to close these gaps while preserving overall capability. I also work on distilling calibration signals from large VLMs into compressed models and canonicalized motion priors for few-shot robotic manipulation.

During my internship at GMI Cloud, I optimized Flux-Schnell (12B DiT) inference on H100 GPUs, achieving ~30 images/min at 1 to 2s latency through TensorRT, GPU memory persistence, and kernel-level tuning. I designed the multi-GPU serving architecture on NCCL, built the full production stack from scratch (queueing, heartbeat monitoring, structured logging, GCS integration, safety filtering), and shipped a video super-resolution pipeline (Real-ESRGAN) integrated with Wan2.2 text-to-video generation.

I grew up in China and did my undergraduate degree in Artificial Intelligence at Shandong University. Those four years gave me a strong grounding in mathematics and control theory, and more importantly, taught me how to think across disciplinary boundaries. I have carried that habit with me ever since.

Pinned Loading

  1. Daily-AI-Trend-Reporter Daily-AI-Trend-Reporter Public

    Your premier source for cutting-edge AI/ML research trends

    Python

  2. yupengtang.github.io yupengtang.github.io Public

    HTML

  3. huggingface/peft huggingface/peft Public

    🤗 PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.

    Python 21.6k 2.5k

  4. huggingface/accelerate huggingface/accelerate Public

    🚀 A simple way to launch, train, and use PyTorch models on almost any device and distributed configuration, automatic mixed precision (including fp8), and easy-to-configure FSDP and DeepSpeed support

    Python 9.9k 1.5k

  5. linkedin/Liger-Kernel linkedin/Liger-Kernel Public

    Efficient Triton Kernels for LLM Training

    Python 6.6k 598

  6. NVIDIA-NeMo/RL NVIDIA-NeMo/RL Public

    Scalable toolkit for efficient model reinforcement

    Python 2k 551