Models & LabsUnited States

Why AI Inference Infrastructure Fails Differently From Traditional Services

Credited to HPCwire · hpcwire.com

Useful

Most engineers know what an overloaded web service looks like. Latency goes up. Queues start growing. CPU gets busy. Eventually requests begin timing out and autoscaling tries to catch up. AI inference systems can fail in some of the same ways, but I’ve found that the usual mental model does not always hold up very well once models, accelerators, batching, and long startup times get involved.

The confusing part is that the service may still look healthy from the outside. Pods are running. GPUs are busy. Requests are still completing. And yet users are already waiting much longer than they should. That is where inference reliability starts to feel different.

Average Latency Can Look Fine While Users Are Already Unhappy

One of the first things I would be careful with is average latency. Inference servers often batch requests together because that makes better use of expensive accelerator capacity. From a throughput point of view, that can be exactly what you want. But the request does not necessarily start executing the moment it arrives.

It may sit in a queue while the server waits for a batch to form. As traffic increases, those waits can get longer. Some requests may also require more work than others, which makes the distribution even less predictable. The average can still look surprisingly normal. The tail usually tells the real story.

If p99 latency is climbing while average latency and overall throughput still look acceptable, I would not call that a healthy serving path. I want to know how long requests are sitting in the queue, how batches are forming, and what the accelerator is doing at the same time. Looking at any one of those signals by itself is usually not enough.

A New Worker Is Not Necessarily New Capacity

Another thing that trips people up is cold-start behavior. Starting a conventional application instance may involve loading configuration, opening connections, and becoming ready to serve traffic. An inference worker can have much more to do.

The model may need to be downloaded or mounted. Large weights have to be loaded into memory. Accelerator state has to be initialized. Caches may need warming. Some serving stacks also perform compilation or optimization before the worker reaches useful performance.

So Kubernetes may tell me another pod exists. That does not mean I actually have another unit of serving capacity yet.

I think this distinction between allocated capacity and ready inference capacity is important, especially during recovery or a traffic spike.

If demand is rising quickly, adding five workers does not help much if all five are still loading models while the existing workers are already overloaded. The scheduler sees capacity coming. The users are still waiting.

GPU Utilization Needs a Different Reading

GPU utilization is another place where habits from conventional infrastructure can be misleading.

With CPU-based services, teams often get used to treating utilization as an obvious scaling signal. Accelerators are different. High GPU utilization can be a good thing. These are expensive resources, and keeping them busy is usually the point.

So when I see high accelerator utilization, my first question is not simply, “Is the GPU too busy?” I want to know what is happening around it. Is queue depth rising? Is accelerator memory becoming tight? Are batches becoming larger? Is p99 moving? Are requests spending more time waiting before execution?

A GPU running near full utilization could describe a very efficient system. It could also describe a system with almost no room left for the next traffic burst. The percentage alone does not tell me which one I am looking at.

The Q ueue Is Where a Small Problem Can Become a Large One

The failure pattern I find most interesting usually begins before anything actually crashes.

Traffic increases. The accelerators are busy, so requests start waiting. The queue grows. As queue time rises, end-to-end latency rises with it. Eventually some clients hit their timeout limits. Then some of them retry.

Those retries go straight back to the same service that was already struggling. Now the queue grows even faster. Autoscaling may eventually add more workers, but those workers still have to load the model before they can take useful traffic. Meanwhile, retries keep arriving.

Nothing in this sequence requires the model server itself to fail. Every component may be doing exactly what it was configured to do.

That is what makes this kind of incident tricky. The failure is not one broken component. It is the interaction between queueing, retries, startup time, and the amount of capacity that is actually ready.

By the time request failures become obvious, the system may already be well into the feedback loop.

Autoscaling Can Arrive After the Problem

This is why inference autoscaling deserves more thought than simply attaching a scaling policy to CPU or GPU utilization.

If queue depth is the first sign that demand is outrunning usable capacity, waiting until the accelerator is fully saturated may be too late. Startup time makes the delay worse.

Imagine the platform decides it needs more workers only after queues are already unhealthy. Those workers now need time to initialize before they help. The scaling decision is effectively responding to traffic that arrived several seconds or even minutes earlier.

For steady workloads, that may be manageable. For bursty inference traffic, it can be painful. I would rather look at queue depth, queue growth, concurrency, ready serving capacity, and worker startup time together.

The important question is not only whether the hardware is busy. It is whether the system can absorb what is coming next.

Sometimes the Right Answer Is Not More of the Same Model

Fallback behavior is another reliability decision that I think teams should make before they need it.

Suppose the preferred model is overloaded and requests are beginning to wait too long. The obvious response is to add more capacity. But that may not be the only option.

For some applications, temporarily serving a smaller model may be better than allowing requests to sit in a queue until they time out. In other cases, a reduced response may be safer than accepting work the system knows it cannot finish within its latency target.

There is no universal fallback strategy. The questions are application-specific. Which requests can use the fallback? How different can the output be? When should traffic move back to the primary model? How will operators know that fallback behavior is happening?

What I would avoid is designing that behavior for the first time during an incident.

The SRE Fundamentals Still Apply

I do not think AI infrastructure needs a completely new reliability discipline. The familiar things still matter. Latency matters. Capacity matters. Load testing matters. Error budgets matter. Observability and failure isolation matter. What changes is the shape of the failure.

An inference service can be unhealthy even while accelerators are busy and errors are low. It can appear scaled because new workers exist even though their models are not ready. It can make its own overload worse because retries arrive faster than useful capacity. And sometimes the safest response is not simply adding more replicas of the same model. Those are the behaviors I would design monitoring and recovery around.

AI infrastructure is still distributed infrastructure. Model serving just adds a few new ways for familiar reliability problems to hide. The sooner teams treat those behaviors as ordinary production concerns instead of unusual AI edge cases, the easier these systems become to operate.

About the author: Sai Joshitha Kathari is a Senior Site Reliability Engineer specializing in distributed systems, Kubernetes, AI infrastructure, observability, and production reliability.

Original · HPCwire

FrontMethod