Chapter 0 · Inference
Scale changes the problem
0.4

Scale changes the problem

Here is the part that catches teams out. The dominant inference problem is not fixed. It changes as you grow, and each new version of it is largely unrelated to the one you just finished solving.

A team that has spent six months becoming excellent at runtime tuning does not thereby become good at capacity planning. The skills barely overlap. What was a config file becomes a conversation with a vendor about what is physically available in a given region next quarter.

1 GPU
One box

Runtime performance

Everything that matters happens inside a single instance. Batch size, cache configuration, precision, kernel selection. Nobody is paging you; you are just deciding how much of the GPU you are willing to waste.

In practice: Tuning an engine config and re-running a benchmark.

~10 GPUs
A few replicas

Autoscaling

Traffic now varies more than one instance can absorb. The questions become when to add a replica, how fast it can be ready, and what happens to requests arriving during the gap. Cold starts stop being a curiosity and start being the p99.

In practice: Arguing about scale-up thresholds and container image size.

~100s
A fleet

Capacity

You cannot get the GPUs. Not because of budget, but because the region is out. Workloads spread across regions and cloud providers because that is where the hardware physically is, and reservations become a procurement exercise rather than an API call.

In practice: Talking to three vendors about H100 availability in Q3.

1,000+
Multi-cloud

Fragmentation

Capacity exists but it is stranded. One cluster queues requests while another sits idle two regions away, and neither knows about the other. The remaining win is treating every GPU you rent, anywhere, as one pool.

In practice: Building the routing layer that makes four clusters behave like one.

What breaks next, by fleet size. Each stage has a different dominant problem, and solving one does not prepare you for the next. The work at the top of the ladder barely resembles the work at the bottom.Illustrative numbers

The last stage is the interesting one. Once workloads are spread across regions and providers, the failure mode stops being “we cannot get GPUs” and becomes “we have GPUs and cannot use them.” One cluster queues requests while another idles. The work at that point is unification: making everything you rent, everywhere, behave as a single pool of compute.

Spreading out has two side benefits worth naming. It protects you from any one region or provider having a bad day, and for a global product it puts inference physically closer to users, which shortens every round trip.