The Human-in-the-Loop Approach
Most GPU cost overruns aren't the result of a bad decision. They're the result of no decision at all — spend that happens by default because nobody was explicitly asked to approve it at the moment it mattered. The fix isn't more process for its own sake. It's putting a human judgment call at the specific points where GPU spend is about to change in kind, and automating execution, not judgment.
Most GPU cost overruns aren't the result of a bad decision. They're the result of no decision at all — spend that happens by default because nobody was explicitly asked to approve it at the moment it mattered. The fix isn't more process for its own sake. It's putting a human judgment call at the specific points where GPU spend is about to change in kind, not just in degree.
Automation Is Not the Enemy Here
To be clear upfront: this isn't an argument against automation. Automated provisioning, autoscaling, and scheduled teardown are genuinely useful, and later posts in this series cover them directly. The point of a human-in-the-loop approach isn't to slow everything down with manual approval — it's to identify the small number of moments where a person should look at what's about to happen before it happens automatically forever.
Where the Loop Actually Belongs
Not every GPU decision needs a human check. Scaling an already-approved, already-understood inference deployment up and down with traffic doesn't need someone to sign off every time — that's exactly the kind of decision automation should own. The moments that warrant a human look are the ones where the shape of spend is about to change:
- Before the first GPU gets provisioned for a new workload. Has anyone confirmed a GPU is actually necessary, or is it the default starting point because that's how the last project started?
- Before scaling from an experiment to a real allocation. A small test run proving a concept works is a different decision from committing a multi-GPU cluster to it. The second decision deserves a second look, informed by what the first one actually showed.
- Before a training or fine-tuning job gets re-run. The first re-run after a failed attempt is normal. The third re-run of a job that still isn't converging is a different situation, and it's worth someone explicitly deciding to continue rather than continuing by default.
- Before a development or staging resource becomes long-running. The difference between "I need this for the next two hours" and "this is now part of our permanent infrastructure" is a decision, even when nobody makes it on purpose.
- Before inference infrastructure gets sized. This is the single highest-leverage checkpoint in the whole list, because inference cost is continuous. Getting the initial sizing decision reviewed by a person who understands both the traffic pattern and the cost model prevents a mistake that otherwise compounds for as long as the model stays in production.
What a Good Checkpoint Actually Looks Like
A checkpoint doesn't need to be a formal approval process with a ticket queue. In practice, the most effective version is lightweight: a short, specific question that has to be answered — in writing, even briefly — before spend proceeds. "What does this experiment need to show for us to scale it?" "What's the traffic estimate this inference sizing is based on?" "What happened in the last run that this re-run is meant to fix?"
The value isn't in the bureaucracy. It's in forcing the assumption behind the spend to be stated explicitly, where it can be checked, instead of remaining implicit, where it can't.
This Gets Easier, Not Harder, With Flexible Infrastructure
A common objection to adding any checkpoint is that it slows teams down — and if provisioning a GPU takes three days of quota approval to begin with, adding a review step on top feels like a real cost. But on infrastructure where provisioning is fast and commitment-free — a decentralized network like GPU.net, for instance, where compute can be spun up in minutes without a long procurement cycle — a human checkpoint costs almost nothing in time. The review happens, the answer is usually "yes, proceed," and the team moves on just as fast as if the checkpoint didn't exist. The difference only shows up on the rare occasion when the answer is "wait, let's rethink this" — which is precisely the moment a checkpoint is supposed to catch.
Automate execution, not judgment
There's a version of this series that ends with a fully autonomous system: telemetry in, GPUs scaled down automatically, nobody in the loop. That version is tempting to build and wrong to ship first.
Not because automation is bad. Because judgment and execution are two different things, and confusing them is how automated systems lose trust — usually in one incident, right before someone disables the whole thing and goes back to manual spreadsheets.
The right split is simpler than it sounds: automate execution, not judgment.
Two different jobs, wearing one hat
"Should we scale this workload down?" and "scale this workload down" feel like one action. They're not.
The first is a judgment call. It weighs signals against context nobody wrote down: is this dip normal for a Tuesday, or is this the calm before a launch? Did someone on the team mention a traffic spike is coming? Is this workload tied to something in beta that's about to get real users? Some of that context lives in dashboards. A lot of it lives in Slack threads, calendars, and institutional memory that no metrics pipeline captures.
The second is mechanical. Once the decision is made, executing it — changing a replica count, reallocating a GPU, updating a scheduler config — is exactly the kind of repetitive, error-prone-when-done-by-hand task that automation is good at.
Systems that blur these two jobs tend to fail in a specific way: they make a judgment call correctly 95% of the time, execute it flawlessly, and then on the 5% — the launch week, the unannounced traffic shift, the workload nobody flagged as special — they execute a bad judgment call just as flawlessly. Nothing catches it, because nothing was designed to. Confidence and correctness aren't the same thing, and a system that's fast and wrong is worse than a system that's slow and asks first.
What this looks like in a GPU cost system
Applied to the ghost-capacity detector from earlier in this series, the split looks like this:
Automated (execution)
- Collecting GPU telemetry continuously
- Computing the ghost score against defined thresholds
- Generating a specific, actionable recommendation ("8 GPUs → 4 GPUs")
- Executing an approved scaling action via the orchestrator API
- Monitoring latency/throughput after the action and rolling back if SLOs are breached
Human (judgment)
- Approving the first N scale-down recommendations for a given workload
- Deciding whether a detected pattern reflects a real trend or a known, temporary anomaly
- Setting the guardrails themselves — min/max replicas, cooldown windows, which workloads are eligible for automation at all
- Overriding the system when context outside its telemetry says the recommendation is wrong
The system never has to guess whether to trust its own judgment, because it isn't making the judgment call in the areas where trust hasn't been earned yet. It's proposing, executing what's approved, and reporting back.
Recommendation-only isn't a lesser version — it's the actual product for a while
It's tempting to treat "recommend, don't act" as a stepping stone to the real system. In practice, it is the real system for longer than most teams expect, and that's fine.
Ghost detected
Workload: llama-api
GPUs: 8
Average utilization: 14%
Recommended action: 8 GPUs → 4 GPUs
Estimated GPU capacity recovered: 4 GPUs
[ Approve ] [ Dismiss ]
A recommendation like this does almost everything a fully automated action does — surfaces the problem, quantifies the opportunity, proposes a specific fix — minus the one thing that requires trust the system hasn't earned yet. And every approval or dismissal is a labeled data point: it tells you where the model's judgment matches human judgment, and just as importantly, where it doesn't. You can't get that signal from a system that skipped straight to automation.
Earning automation, tier by tier
The path from "always ask" to "sometimes act on its own" shouldn't be a single switch. It should track confidence:
- Low confidence — Basis: underused, but demand unpredictable. Behavior: log only — no recommendation surfaced.
- Medium confidence — Basis: sustained underuse (~30 min), pattern not yet proven. Behavior: recommend, require approval.
- High confidence — Basis: sustained underuse (~2 hrs), stable demand, healthy latency, spare capacity confirmed. Behavior: eligible for automatic action, within hard guardrails.
Even at High confidence, "automatic" doesn't mean "unsupervised." It means the human isn't approving this specific action — they already approved the conditions under which actions like it are allowed to happen. The guardrails (max scale-down percent, latency SLOs, cooldown periods, minimum replica floors) are still a judgment call a person made in advance. The system is executing within boundaries a human set, not inventing new boundaries on the fly.
That's the actual definition of human-in-the-loop: not a human clicking approve on every action forever, but a human's judgment embedded in every action the system is permitted to take — either directly, one approval at a time, or indirectly, through the limits they set.
Why this holds up as GPUs get more expensive, not less
There's a version of this argument that only works while GPUs are cheap enough that mistakes don't matter. That's the opposite of the current reality. As GPU costs climb, the cost of a bad automated decision climbs with it — a wrongly scaled-down inference service isn't a rounding error, it's an incident with a user-facing latency spike and a very expensive root-cause meeting.
That's exactly why judgment doesn't get automated away as systems mature. What gets automated is the distance between a judgment and its execution — how quickly a human's decision turns into action, and how reliably it's monitored afterward. The judgment itself stays anchored to a person for as long as the stakes justify it. For most GPU fleets, that's a long time.
Automate the toil. Keep the judgment. That's not a temporary compromise on the way to full autonomy — for a system spending real money on real infrastructure, it's the design.
The next post in this series starts putting this into practice: what it actually looks like to begin a project with no GPU at all.




![GPUNET Verifiable Exchange: The Next Frontier for $GPU, Nodes and Ecosystem [TEASER]](https://i.ibb.co/Z1JWjN7r/Article-Cover.png)










