Running GPU Workloads on Kubernetes: Where Networks Like GPU.net Fit In

Running GPU Workloads on Kubernetes: Where Networks Like GPU.net Fit In

Kubernetes was never really designed with GPUs in mind. It gives you placement, not insight — it'll happily schedule eight GPU replicas for you, but it has no opinion on whether you needed eight. That gap, between what Kubernetes manages and what it actually understands about GPU usage, is exactly where a layer like GPU Net fits. This is Part 4 of the Core GPU Cost Series.

Vishwas Narayana

Core GPU Cost Series — Part 4

Kubernetes was never really designed with GPUs in mind. It was extended to support them — device plugins, node labels, resource requests bolted onto a scheduler built for CPU and memory. That extension works well enough that most GPU fleets run on it today. But it also means Kubernetes gives you placement, not insight. It'll happily schedule eight GPU replicas for you. It has no opinion on whether you needed eight.

That gap — between what Kubernetes manages and what it actually understands about GPU usage — is exactly where a layer like GPU Net fits.

What Kubernetes already gives you

To be clear about the gap, it's worth being clear about what's already there.

  • Scheduling: the device plugin model lets you request GPUs the same way you request CPU/memory, and the scheduler places pods on nodes with available GPUs.
  • Isolation: each GPU (or MIG slice) is exclusively assigned to a pod, so workloads don't silently collide.
  • Autoscaling primitives: HPA and cluster autoscalers can react to custom metrics, including GPU utilization, to add or remove replicas and nodes.
  • Declarative state: your Deployment says "8 replicas," and Kubernetes enforces that until you say otherwise.

This is a solid foundation. It's also entirely mechanical. Kubernetes will maintain "8 replicas" with total conviction whether that number reflects real demand or a launch-week decision nobody revisited. It has no native concept of a GPU that's technically busy but not actually earning its keep — the ghost GPU problem from earlier in this series.

Where the gap actually shows up

Three places, consistently:

1. Utilization metrics exist, but nobody's watching them as a fleet

kubectl top and DCGM exporters will tell you what a single node or pod is doing right now. They won't tell you that twelve inference workloads across the cluster are each sitting at 15% utilization and collectively represent a recoverable chunk of your fleet. That's a fleet-level question, and Kubernetes objects are inherently workload-level.

2. HPA scales on thresholds, not on sustained evidence

Horizontal Pod Autoscaler can react to a custom GPU utilization metric, but it's a threshold-crossing tool by design — cross above X, add a replica; cross below Y, remove one. It doesn't distinguish a 30-second dip from a genuine two-hour trend, and it has no built-in concept of confidence tiers or "recommend first, act later." That nuance has to live somewhere else.

3. Nobody owns cross-workload consolidation

Kubernetes will happily let workload A run at 12% on its own GPU and workload B run at 15% on a separate GPU forever. Nothing in the base scheduler asks "could these share a GPU?" Bin-packing exists for CPU/memory requests; it doesn't extend to GPU-level consolidation without additional tooling (GPU sharing, MIG-aware scheduling, or a layer that makes that recommendation explicitly).

Where GPU Net sits in the stack

GPU Net isn't a replacement for the Kubernetes scheduler or HPA — it's a layer that sits above workload-level autoscaling and answers the question Kubernetes was never built to ask: across the whole fleet, where is capacity being wasted, and is it safe to reclaim?

GPU Nodes
   │
   ├── dcgm-exporter
   │        ↓
   │    Prometheus
   │        ↓
   │  Ghost Detector / Ghost Score      ← GPU Net layer
   │        ↓
   │  Recommendation / Policy Layer     ← GPU Net layer
   │        ↓
   │    Kubernetes API
   │        ↓
   └── Deployment / HPA → GPU replicas

Kubernetes still does the placement and the mechanical scaling. GPU Net does the judgment layer underneath it: turning raw per-node metrics into a fleet-wide Ghost Capacity number, deciding whether a detected pattern meets the bar for a recommendation, and — once a workload has earned high confidence — issuing the scale/rebalance/consolidate action through the Kubernetes API rather than around it.

That last part matters. GPU Net shouldn't fight Kubernetes for control of a Deployment. It should act as a client of the Kubernetes API, the same way a human operator running kubectl scale would — just with better evidence and a documented reason attached to every change.

A concrete example

Take the earlier scenario from this series:

Deployment: llm-api
Replicas: 8
GPUs:      8

GPU utilization:
GPU 0: 14%   GPU 4: 15%
GPU 1: 11%   GPU 5: 13%
GPU 2: 17%   GPU 6: 10%
GPU 3: 12%   GPU 7: 16%

Kubernetes sees eight healthy pods, each with a GPU, doing exactly what they were told to do. Nothing here trips an alert — the pods aren't crashing, latency isn't degraded, nothing is unhealthy by Kubernetes' definition of healthy.

GPU Net sees eight GPUs sustaining sub-20% utilization well past the duration threshold, computes a ghost score for the workload, and — at Medium confidence — surfaces a recommendation: 8 GPUs → 3 GPUs, ~5 GPUs recoverable. A human approves it. GPU Net calls the Kubernetes API to update the Deployment's replica count. Kubernetes takes it from there — terminating pods, rescheduling, all the mechanical work it's already good at.

Neither layer replaces the other. Kubernetes enforces state; GPU Net decides what the state should be.

Why this belongs as a layer, not a patch to Kubernetes itself

It's tempting to imagine folding ghost detection directly into the scheduler or into HPA's metric logic. Resist that. A few reasons:

  • Blast radius. A bug in scheduling logic can break pod placement cluster-wide. A bug in a recommendation layer produces a bad recommendation that a human can reject.
  • Auditability. A separate policy layer gives you a clean log of why an action happened — which Ghost Score, which confidence tier, which approval — instead of burying that reasoning inside scheduler internals.
  • Portability. Fleets are rarely single-cluster or single-orchestrator forever. A layer that talks to the Kubernetes API (or any orchestrator API) generalizes better than logic wired into one scheduler's internals.
  • Staged trust. As covered in the last piece in this series, automation has to be earned in tiers. That staging is much easier to build and reason about as its own layer than as a patch scattered through core scheduling code.

The takeaway

Kubernetes answers "where should this workload run and how many replicas does it have right now." GPU Net answers "does it need that many, and is it safe to change." Those are different questions, sitting at different layers, and conflating them is how you end up with a scheduler that's technically doing its job perfectly while quietly costing you a fortune.

More Stories

Arrow leftArrow left
Try our Planetary Grid of Compute Now!