Chapter 0 · Inference
Where to put the abstraction
0.5

Where to put the abstraction

Once the runtime and infrastructure work exists, somebody has to decide how it is exposed. This applies whether you are buying inference from a provider or building a platform for your own engineers; the question is the same either way.

At one extreme, inference is a black box: hand over weights, receive an API. At the other, you get compute, network, and disk, and everything above that is yours. Both extremes are defensible and most teams belong somewhere between them.

Hand it off
01Model API
You get

A URL and a token. Somebody else picked the hardware, the engine, the batch size, and the quantization.

You give up

Nearly all of it. You cannot fix a latency problem that originates below your API call, and you cannot deploy a model nobody else is hosting.

SuitsGetting a product in front of users before optimizing anything.

02Managed deployment
You get

Weights in, autoscaled endpoint out. Engine choice and a handful of serving knobs are yours; the cluster underneath is not.

You give up

Kernel-level control and the ability to do anything unusual with scheduling or routing.

SuitsTeams serving open or fine-tuned models who would rather ship features than run Kubernetes.

03Serving framework
You get

vLLM or SGLang configured by you, running in your containers, on infrastructure you operate.

You give up

Operational time. Autoscaling, capacity, rollout, and observability are now yours to build and keep working.

SuitsTeams with a real inference workload and at least one engineer who wants to own it.

04Custom runtime
You get

Your own kernels, your own scheduler, your own cache. Everything is tunable because you wrote it.

You give up

Time, and a lot of it. This is a multi-engineer commitment that only pays back at scale or at the edge of what published engines support.

SuitsFrontier labs, and teams whose model or modality nothing off the shelf serves well.

Own every layer
How much of the stack do you want to own?. Control and productivity trade against each other along one axis. Each step down buys the ability to tune deeper layers and costs the engineering time to run them. Take as much abstraction as your requirements allow, and give up productivity only where something you need demands the control.

I would push most teams further toward abstraction than instinct suggests. Owning the whole stack is satisfying and it is occasionally correct, but it is a standing commitment of engineering time that competes directly with the product. Keep only as much control as you have a concrete use for.