Measure Before You Scale

Measure Before You Scale

Your proof of concept passed. The natural next move is to scale up — more GPUs, bigger model, longer runs. Don't. Not yet. The fix for a compute-bound workload and the fix for an I/O-bound one are almost opposite, and scaling the wrong axis just buys you a bigger version of the same bottleneck. This piece is about the metrics that tell you which one you have.

Vishwas Narayana

Core GPU Cost Series

Your proof of concept passed. Loss dropped where you wanted it to, on real data, on real hardware. The natural next move is to scale up — more GPUs, bigger model, longer runs.

Don't. Not yet. Scaling before you've measured why your POC ran the way it did is how a team ends up with 8 GPUs solving a problem that had nothing to do with GPU count in the first place. Before you scale, you need to know what's actually limiting your current run — because the fix for a compute-bound workload and the fix for an I/O-bound one are almost opposite, and scaling the wrong axis just buys you a bigger version of the same bottleneck.

This piece is about the metrics that tell you which one you have.

Utilization is a symptom, not a diagnosis

nvidia-smi gives you GPU utilization as a single percentage, and it's the most commonly misread number in this entire series. High utilization doesn't necessarily mean "compute-bound and working hard." Low utilization doesn't necessarily mean "wasted capacity." Both can be true for reasons that have nothing to do with what you'd assume.

A GPU at 95% utilization spending most of its cycles waiting on memory reads and writes still reports high utilization — the SM is busy, just not busy doing the arithmetic you care about. A GPU at 20% utilization might be correctly, efficiently waiting on a data loader that hasn't kept up. Utilization tells you the GPU isn't idle. It doesn't tell you what it's doing instead of the work you wanted.

You need a small set of metrics together to actually diagnose a bottleneck.

The core metric set

Compute utilization (SM activity)

What nvidia-smi's utilization number approximates. Rising, sustained SM activity during the compute-heavy parts of a step is what you want to see. If this is consistently low during what should be the heaviest part of your workload, something upstream is starving the GPU.

Memory bandwidth utilization

How much of the GPU's memory bandwidth is actually in use, distinct from how much memory is allocated. A workload can use very little of its available VRAM while still being memory-bandwidth-bound — small tensors moved constantly can saturate bandwidth long before they fill capacity.

Tensor Core / mixed-precision utilization

If your workload is meant to run in fp16/bf16 with Tensor Cores, check that it actually is. It's common to write code that should hit Tensor Cores and silently falls back to fp32 paths due to an unsupported op, a dtype mismatch, or a shape that doesn't meet alignment requirements. This shows up as "GPU busy, throughput far below spec" — a classic false-compute-bound signal.

PCIe / NVLink throughput

Data movement between host and device, or between GPUs. A workload with a data loader that can't keep pace, or a multi-GPU setup with unnecessary cross-device communication, will bottleneck here — and it'll often look like low GPU utilization, which gets misdiagnosed as "the GPU isn't the bottleneck" when the real fix is still squarely in the data pipeline.

Data loader throughput vs. model consumption rate

The most common bottleneck for small-to-mid POCs, and the easiest to miss because it's not a GPU metric at all. Compare samples/sec your data pipeline can produce against samples/sec your model can consume. If the loader is slower, the GPU spends real, measurable time waiting — and no amount of additional GPU hardware fixes that.

Step time breakdown

Wall-clock time per training step, split into: data loading, forward pass, backward pass, optimizer step, logging/checkpointing overhead. This is the single most diagnostic metric on the list, because it turns "GPU utilization is low" from a mystery into a specific line item.

Reading the combinations

Metrics mean little in isolation. Here's how they combine into an actual diagnosis:

  • High compute · High memory bandwidth · Loader keeping up → Genuinely compute-bound — this is the good case for scaling GPU count/size.
  • High compute · Low memory bandwidth · Loader keeping up → Compute-bound, likely underusing Tensor Cores or running suboptimal kernels — check dtype/shapes before scaling.
  • Low compute · Low memory bandwidth · Loader not keeping up → I/O-bound — data pipeline is the bottleneck, not the GPU.
  • Low compute · High memory bandwidth · Loader keeping up → Memory-bandwidth-bound — small/frequent tensor ops; consider batching or fusing ops before adding hardware.
  • Variable/bursty compute · Variable/bursty memory bandwidth · Loader keeping up → Overhead-bound — checkpointing, logging, or synchronization stalls eating step time.

Only the first row is a workload where "add more GPUs" straightforwardly helps. Every other row has a fix that doesn't involve more hardware — and scaling GPU count against those bottlenecks just reproduces the same wait states across more expensive machines.

A worked example

Say your POC step time is 800ms, and profiling breaks it down as:

Step time breakdown (800ms total)
──────────────────────────────────
Data loading        420ms   (52%)
Forward pass        180ms   (22%)
Backward pass       150ms   (19%)
Optimizer step       30ms    (4%)
Logging/checkpoint   20ms    (3%)

GPU compute utilization: 24%
Data loader throughput:  340 samples/sec
Model consumption rate:  890 samples/sec

The GPU is idle more than three-quarters of the time, and the reason is right there: the data loader produces samples at roughly a third of the rate the model could consume them. Scaling this workload to 4 GPUs would give you four GPUs each waiting on the same slow loader — a 4x hardware bill for close to 0x improvement in wall-clock training time. The actual fix is in the data pipeline: more loader workers, prefetching, caching preprocessed data, or a faster storage backend. Only after that gap closes does "add GPUs" become the correct next lever.

What to actually check before scaling

  1. Run the step-time breakdown first, always. It's the fastest path from "something feels slow" to "here's the specific stage that's slow."
  2. Confirm compute utilization is genuinely high during the compute stages, not just averaged over the whole step including idle waits.
  3. Verify mixed precision is actually engaging Tensor Cores, if that's the plan — don't assume the dtype you configured is the dtype actually executing.
  4. Compare data loader throughput to model consumption rate explicitly. Don't infer it from GPU utilization alone; measure both sides directly.
  5. Only scale the resource that profiling identifies as the constraint. More GPUs for an I/O-bound workload, more VRAM for a compute-bound one, and faster storage/more loader workers for a data-bound one are three different purchases — buying the wrong one doesn't just waste money, it leaves the real bottleneck exactly where it was.

Why this belongs before the scale-up decision

Every piece before this one in the series has been about not paying for GPU capacity you don't need — starting on CPU, sizing the smallest useful experiment. This is the same principle applied one step later: before you multiply GPU spend by scaling up, multiply your certainty that GPUs are actually what's limiting you. Utilization alone won't tell you that. The full metric set will.

More Stories

Arrow leftArrow left
Try our Planetary Grid of Compute Now!