Education

B.S. in Computer Science

University of California, Santa Barbara
06/2027

Relevant Courses: Foundation Models (grad), Machine Learning, Generative AI, Operating Systems, Parallel Computing

GPA: 3.98/4.0

Experience

NVIDIA

ML Research Engineer Intern
06/2026 – Present | Santa Clara, CA
  • Owned the end-to-end post-training pipeline for Nemotron on multimodal finance reasoning, spanning synthetic data generation and curation, SFT/GRPO/OPD, and evaluation.
  • Ran large-scale training experiments across hundreds of GB200 and H100 GPUs using Megatron-LM and NeMo RL, diagnosing data and model failure modes to improve omni reasoning capabilities.
  • Contributed open source training recipes to NVFlow⁠ and finance datasets to Hugging Face (releasing soon).
  • 03/2026 – 06/2026 | Santa Barbara, CA
  • Extended ChipAgents coding harness into a long-horizon, multi-agent system with persistent subagents, inter-agent communication, resumable execution, and live subagent steering.
  • Built batch execution mode for reproducible, large-scale evaluation of ChipAgents on internal RTL benchmarks
  • UCSB AI Group⁠

    Undergraduate Researcher | Advisor: Xin Eric Wang
    10/2025 – Present | Santa Barbara, CA
  • Led creation of a long-horizon coding benchmark that reframes context rot from static retrieval (needle-in-the-haystack) to agentic performance under realistic user/environment distractors. (COLM 2026 Context Beyond the Window Workshop)
  • Developed an embodied benchmark for functional Theory of Mind, testing whether agents can act on others’ beliefs during collaboration rather than merely answer explicit belief questions. (ICML 2026 AI4GOOD Workshop)
  • Pretrained a 23M language model⁠ from scratch in PyTorch, hand-writing the BPE tokenizer, RoPE/SwiGLU architecture, AdamW optimizer, and training pipeline.
  • 06/2025 – 09/2025 | San Jose, CA
  • Wrote depthwise convolution kernels in Rust to achieve 16x faster performance on our NPU. Evaluating on MobileNetV2 demonstrated a 5.8x improvement in inference latency.
  • Explored efficient autoscheduling for CNN/MLP operations by extending Halide to support in-memory compute primitives.
  • Awards
    10/2025

    Won the world's largest collegiate hackathon (3k+ participants, 700+ teams) building computer-use agents remotely controlled through FaceTime or iMessage. Released 1 month before OpenClaw.

    Publications

    ContextRotBench: Measuring Long-Horizon Degradation in Coding Agents

    COLM 2026 CBW Workshop
    07/2026

    Dylan Lu*, Surya Gunukula*, Saaket Agashe, Tengxiao Liu, Xin Eric Wang

    Gurusha Juneja*, Dylan Lu*, Saaket Agashe, Parth Diwane, Edward Gunn, Jayanth Srinivasa, Gaowen Liu, William Yang Wang, Yali Du, Xin Eric Wang