Allen Lu
ML infrastructure engineer also with a degree in Arts and Design

ABOUT ME
I work on the layer between models and hardware. At Ambient.AI I moved an eleven-model real-time computer vision pipeline off monolithic PyTorch onto NVIDIA Triton, then spent most of my time in profilers finding out why the GPUs were still idle.
The answers are rarely in the model. They are in graph breaks, kernel launch counts, a global flag contaminating a shared CUDA context, a scheduler that never overlaps preprocessing with execution. I like the part of the job where a number moves because someone finally understood the system.
Before Ambient I studied computer science at UC San Diego and National Tsing Hua University, and did research on reinforcement learning environments for LLM tool use and on cross-spectral estimators for multichannel EEG.
- Python
- C++
- CUDA
- PyTorch
- TensorRT
- NVIDIA Triton
- vLLM
- SGLang
- Nsight Systems
- OpenAI Triton
- DeepSpeed
- Docker
- Kubernetes
- Prometheus
- Grafana
- FastAPI
- AWS
- Go