Your GPUs Might Be Less Busy Than You Think

Many organizations overspend on AI infrastructure because standard dashboards mismeasure GPU efficiency.

Every AI infrastructure conversation seems to start in the same place right now: GPUs.

More GPUs. Bigger clusters. Longer reservation commitments. More capacity planning meetings.

That makes sense. AI demand is growing quickly, and nobody wants to be the team that runs out of compute just as the business starts scaling its AI initiatives.

But over the past year, we’ve started hearing a different question from infrastructure teams building private AI environments:

How Do We Know the GPUs We’re Already Paying for Are Actually Doing Work?

At first glance, that sounds like an odd question. Most organizations already have dashboards for that. Open any monitoring console and you’ll find utilization metrics, memory consumption, queue depths, and plenty of graphs suggesting exactly how busy a system is.

The problem is that AI workloads aren’t behaving like the workloads many of those dashboards were originally built to measure.

Several teams have discovered that a GPU showing 95% or even 100% utilization may not actually be spending most of its time performing calculations. In some cases, the hardware is spending more time waiting on memory transfers than executing math. The dashboard isn’t necessarily wrong. It’s simply measuring something different than people assume. 

That distinction might sound academic, but it has real-world implications. If a company believes its infrastructure is fully utilized, the next step is often to buy more infrastructure. If the underlying hardware still has significant headroom, that’s a very different conversation.

And that’s where things start to get interesting.

When Busy Doesn’t Mean Productive

One architect recently summarized the issue to us in a way that stuck:

“The dashboard says the GPU is busy. The silicon says it’s waiting.”

That’s obviously an oversimplification, but it gets at the heart of the problem.

Most traditional monitoring tools focus on resource allocation. If a workload has reserved GPU memory or maintains active execution streams, utilization appears high. From the perspective of the orchestration layer, everything looks healthy and fully occupied. 

The hardware may have a different story to tell.

Large language models spend a surprising amount of time moving data between memory and compute resources. While weights are being transferred and caches are being populated, Tensor Cores can sit idle waiting for the next operation. To most dashboards, utilization remains high. To the silicon itself, very little actual computation may be occurring. 

This is the phenomenon some infrastructure teams have started referring to as “phantom utilization.”

The interesting part isn’t the name. It’s what happens next.

Once organizations begin comparing virtualization-layer metrics with hardware-level performance counters, they sometimes discover a meaningful gap between what they thought was happening and what was actually happening. 

And suddenly the discussion shifts from:

“How many more GPUs do we need?”

to

“How efficiently are we using the ones we already have?”

Looking Beyond the Dashboard

Rather than relying on the same telemetry most infrastructure teams see in standard dashboards, Utilyze looks directly at hardware performance counters. The goal is to understand how much time GPUs are spending executing arithmetic operations versus waiting on memory movement. 

Whether Utilyze ends up becoming widely adopted is almost secondary.

What’s more interesting is that tools like it are starting to appear at all.

That usually happens when an industry begins questioning its assumptions.

This is the problem Systalyze has been exploring with its open-source profiling tool, Utilyze.

A decade ago, cloud teams started asking whether virtual machine density was the right way to measure efficiency. Today, AI infrastructure teams are asking a similar question about GPU utilization.

The metrics that mattered during the cloud transition may not be the metrics that matter during the AI transition.

And if that turns out to be true, many organizations may discover that their biggest AI optimization opportunity isn’t a new GPU purchase—it’s visibility.

The Agent Problem Nobody Was Talking About

The timing of this conversation is interesting because AI workloads themselves are changing.

For the past few years, most enterprise AI interactions have been human-driven. Someone asks a question, submits a prompt, or generates a document. If the response takes a second or two, nobody loses much sleep over it. 

Agentic workflows change the equation.

A financial analysis agent doesn’t stop after a single inference. It may perform dozens of reasoning cycles. A software engineering agent may continuously evaluate, revise, and re-execute its own outputs. Supply chain optimization systems can trigger chains of dependent decisions before reaching a final recommendation. 

Individually, each step might seem inexpensive.

Collectively, they can become significant.

We’ve heard several architects describe this as the difference between serving one customer at a checkout counter versus operating an assembly line. Small delays that barely register in one scenario become impossible to ignore when multiplied hundreds or thousands of times.

That’s why latency conversations are increasingly moving beyond application teams and into infrastructure discussions.

When AI systems become autonomous, efficiency stops being a user experience metric and starts becoming an economics metric.

Are We Rediscovering the Cost of Abstraction?

One thing we find particularly fascinating is how this conversation echoes debates from earlier generations of infrastructure.

For years, the industry has embraced abstraction—and for good reason.

Virtual machines simplified provisioning. Containers simplified deployment. Orchestration platforms simplified operations.

Most organizations aren’t eager to give any of that up.

But AI workloads may be exposing tradeoffs that weren’t especially visible before.

When models become large enough and workloads become demanding enough, scheduling delays, memory movement, and orchestration overhead can start showing up in ways that matter. What used to be negligible becomes measurable. 

This doesn’t mean virtualization is suddenly bad.

It does mean infrastructure teams are taking a harder look at where performance is being lost.

Systalyze‘s commercial platform is one example of that trend. Its approach focuses on executing workloads closer to the hardware itself rather than relying on traditional orchestration layers. Whether that’s the right answer for every organization is debatable, but the underlying question is becoming increasingly common:

How much of our AI infrastructure spend is going toward actual computation, and how much is being consumed by everything around it?” 

What Early Adopters Are Seeing

The organizations exploring these questions span a surprisingly wide range of industries.

In pharmaceutical research, teams processing large radiology datasets and molecular models are under constant pressure to reduce execution times. According to Systalyze, Bayer was able to dramatically compress processing timelines after shifting to a more hardware-aware architecture, reducing workflows that previously took weeks down to days. 

Financial services presents a different challenge altogether.

In payment processing and fraud detection systems, latency isn’t merely an inconvenience—it’s part of the business model. Systalyze reports significant performance gains compared to several common orchestration approaches in these environments, largely by reducing layers of infrastructure overhead. 

The specific numbers will naturally vary by workload.

What’s more interesting is that organizations in completely different industries appear to be arriving at similar conclusions: understanding hardware behavior matters more than it used to.

The Bigger Takeaway

The AI infrastructure market spends a lot of time talking about scale.

That’s understandable. Scale is easy to visualize. Bigger clusters. More GPUs. Larger models.

But after digging into this topic, we are not convinced scale is the most interesting story anymore.

Efficiency might be.

The teams getting the most value from AI infrastructure over the next few years may not necessarily be the ones with the largest GPU budgets. They may be the teams that develop the clearest understanding of what’s actually happening beneath the dashboard.

Because at some point, every organization will face the same question:

Do we need more infrastructure?

Or do we simply need better visibility into the infrastructure we already have?

For an industry that’s become accustomed to solving performance problems by adding more compute, that’s a surprisingly uncomfortable—and increasingly important—question.

Related

Key players

Enter a search