Writing · 01 published note
Where systems meet reality.
Architecture, inference, APIs, and the uncomfortable parts of shipping them.
Why more GPUs is not enough for LLM inference
What I learned deploying and tuning large-model inference: KV cache, routing, and cache hierarchy matter as much as raw GPU count.
Read note