The $500M Bill Was Always Going to Happen
A $500M Claude bill sounds like a freak accident. It wasn't. The controls to stop it existed the whole time - they just didn't enforce. Alerts notify. Caps enforce. Here's what governance that actually holds looks like.
Let me be honest with you. A $500M Claude bill an enterprise spent on Claude sounds like a horror story. A freak accident. Something that happened to *them*, not you.
It's not. It was the plan working exactly as designed.
Not their plan. The tool's plan. Because here's the thing nobody says out loud: the controls to stop that bill existed the whole time. They were sitting in a dashboard. Someone just had to know to turn them on, harden them, and wire up attribution before the spend ran off a cliff.
Nobody did. Because nobody does.
That's the whole story of AI coding cost governance right now. The controls are *available*. They just don't *enforce*. And "available" has never once stopped an invoice.
Alerts don't stop anything. That's the entire problem.
Quick gut check. When your budget tool "hits a limit," what actually happens?
Most of the time: a Slack ping. A red badge. An email that says you crossed 80%.
Cool. And then the tokens keep flowing.
That's not a control. That's a smoke detector that watches the house burn and files a report. An alert tells you spend crossed a line. A cap stops spend at the line. Those are not two versions of the same feature. One notifies. One enforces. If your cost strategy is a wall of alerts, congrats, you bought a very expensive way to find out you already overspent.
And I get why everyone ships alerts. They demo *beautifully*. Set a threshold, trigger a ping, everyone in the room nods. Looks like control.
But an alert has zero authority to stop anything. Put 150 engineers on frontier models through an agent all day and the gap between "you're at 80% of budget" and "you blew past it hours ago" closes fast. The notification fires right on schedule. The spend doesn't care.

That's the shape of the $500M bill. Every time.
So let's actually look at how the big tools handle this
Because the details are where the whole thing falls apart.
Copilot. There's a governance dashboard, and it looks the part. But "stop usage" is off by default. Budgets act as alerts unless someone explicitly hardens them. Metered usage is on by default. Credits get forfeited monthly, used or not. In a demo you see a ceiling. In production you get a speedometer that watches the overage happen.
Claude Code. Seat allowance isn't metered in dollars until you opt in to usage credits. And there's no native, real-time, per-user attribution. Want to know which engineer or which key is torching the budget right now? You're building your own OpenTelemetry stack to find out. The controls to prevent a runaway bill were *available*. They weren't *enforced*. That's the $500M incident in one sentence.
Cursor. You get usage analytics. Requests, diffs, lines of code. What you don't get is dollar attribution. Billing runs in arrears, so there's no hard ceiling by design, just an invoice that shows up after the damage. And real cost governance is locked behind the Enterprise tier, so the teams most likely to overspend early are the ones least likely to have guardrails.
See the pattern?
Every one of these vendors can say, with a straight face, that governance is "available." And every single time, enforcement is something you opt into, harden yourself, or upgrade to reach.
Available is not enforcing. The bill doesn't care what was available.
OK, so what does "governance that enforces" actually mean?
Here's the bar I'd hold any tool to before I put my company's card behind it. Two things. That's basically it.
One: hard budget caps at every level. User, team, API key, org. Not thresholds that ping. Caps that enforce when configured.
And I want to be specific about what "enforce" means, because this is where the marketing usually gets vague and I hate vague. In Kimchi Coding, every inference request runs through the proxy, and the proxy checks it against the budget *before it runs*. If you're at or over your limit, you get an HTTP 429 back and the request never executes. The tokens never leave the building. Spend stops.

That's a cap. If it doesn't stop the request, it's an alert wearing a cap's clothes.
The caps cascade too. Org sets the ceiling. Team, user, and API key nest under it. Tightest scope wins, and per-model limits still bind if they're tighter. So a service account on a CI runner can't quietly outspend the whole org because someone forgot a limit.
Two: per-user cost attribution, in real time. You should be able to see, *right now*, what every user, team, key, and org is spending. Not reconstruct it next week from logs.
And it should be real because of how it's built, not because of a dashboard someone bolted on. The cost gets computed at the proxy from the model plus the actual token counts as requests flow through. Want per-team numbers? Tag requests with a team header and the teams show up on their own. No admin setup, no config screens, no pre-registering every squad. First tagged request arrives, the team exists. Works across the agents your engineers already use.
Put those two together, per-user cost attribution and hard budget caps, and you've got a governed AI coding platform instead of a dashboard.
And let me kill one myth right now, because I'd rather you hear it from me. "Governed" means the governance is *built in and ready to enforce*. It does not mean it's magically running before you set it up. Nobody should tell you the caps are on before you've configured them. They're not. They're available from day one. You turn them on as part of setup. That honesty matters, because the tools overselling "default" governance are the exact ones that produced the surprise invoices.
Our own bill is the proof
Easy to preach enforcement in the abstract. So here's our receipt.
CAST AI runs 150 engineers through this. Roughly 36 billion tokens a month of real agent work, not a benchmark. Over a 30-day window, measured against a 100% Anthropic Sonnet and Opus baseline, we ran 12x cheaper than frontier-only. Same engineers. Same work. Same expectation that the agent had better be good.
The 12x didn't come from a magic model or a throttle that quietly hands people worse output. It came from smart routing plus governance that actually holds. When the caps are enforced and the attribution is real-time, you stop paying for the runaway tail, the exact tail that turns a normal month into a board-level incident.
Kimchi Coding is made by the creators of CAST AI. This is the bill we cut for ourselves before we ever shipped it to you. And we don't train on your data. Ever. Sovereignty isn't a slide, it's the architecture.
The part everyone gets backwards
One pushback I hear constantly: "governance sounds like it's built to slow my developers down."
Nope. And if you frame it that way you'll build the wrong thing and lose your engineers.
Your developers don't care about your budget dashboard. At all. They care about one thing: is the agent good enough to trust with real work? That's the user's gate, and it's non-negotiable. The agent has to be genuinely good or none of the rest matters.
Governance is a *different* axis. It's the buyer's problem, not the developer's. And the split is clean: your engineers get the best agent, you keep control of the bill.
Those two things are not in tension. The tools that made them feel like a tradeoff did it by making enforcement expensive or optional, so the only lever left for cost was to degrade the experience. Caps that enforce at org, team, key, and user level let you protect the invoice without the developer ever feeling it.
Where this leaves you
The $500M bill wasn't a freak event. It was the predictable end state of governance that notifies and hopes. Every tool shipping alerts-as-governance is one bad month from the same headline.
So if you're picking an AI coding tool this quarter, ask the vendor exactly one question and make them answer it straight:
*When a team hits its budget, does spend stop, or does someone just get a notification that it didn't?*
Everything else is detail.

That's the whole game. Caps that enforce, not alerts that just notify. Real-time attribution you don't have to build yourself. The best agent your engineers will actually use, and a bill you actually control.
See how Kimchi Coding does it at kimchi.dev.