Writing — 45,000 kernel launches and one global flag
< All posts
MAR 2026 · GPU OPTIMIZING · PROFILING · 9 MIN

45,000 kernel launches and one global flag

Placeholder body copy. A batch that should have issued a few hundred kernels was issuing forty-five thousand of them, and every one of those launches cost us the same fixed overhead on the host side. The GPU was not slow. It was waiting.

Nsight Systems makes this kind of thing obvious once you know what shape to look for: a dense picket fence of tiny kernels with gaps between them, and a CPU timeline that is completely saturated. The interesting question is never whether the launches are there. It is which layer put them there.

nsys profile --stats=true python infer.py
  → cudaLaunchKernel  45,182 calls
  → sm__issue_active    17.2%

The flag was global, the CUDA context was shared, and nobody had connected those two facts.

Placeholder continuation. Two more paragraphs of the actual investigation go here, then the fix and the number it moved. Kernel count fell by 500×.

fig-1-nsight-timeline.png
About aaaallleen.github.io
aaaallleen.github.io
Version 1.0 · built 2026-09-14
ENGINE
Astro 7.3.2, fully static
HOSTING
GitHub Pages
DESIGN
Retro desktop, mocked up in Claude Design
TYPE
Silkscreen & Verdana
OWNER
Allen Lu · San Diego, CA