I used to work on a support team of a well known backend type service that had hard budget caps.
It was, unfortunately, a nightmare. There were tons of tickets and even threats of lawsuits from customers whose service got cut off hard at the worst possible time due to organic growth/going viral/big event/nobody knew about the limit/etc. Not only did they lose all the leads and revenue they would have gotten from that bump, but they also pissed off their own existing users who suddenly couldn't use the service either.
Generally speaking, it's much better to use alerts instead of hard limit. Even in the worst case (hackers pwn your credentials and mine Bitcoin or whatever) the rest of your business is unaffected and you can negotiate with the billing department at comparative leisure.
This is all assuming you have humans operating the service. If you're letting AI agents yolo infra in prod, you have a whole series of new problems.
> Generally speaking, it's much better to use alerts instead of hard limit.
Probably for companies and opportunistic individuals with a high risk tolerance. As for me, just give me hard budget caps. It shouldn't be a choice between supporting one of the approaches, when you could let each client pick what they want.
The thing I see is folk putting a limit when given the choice. Later the situation changes, they want mega scale. Then don't fiddle the settings. I think it's a documentation & process problem.
1000 years ago, before CI/CD took over folk would have launch/go-live events, where teams would do a good checklist. There were pre-launch meetings to review, which included load checking.
That doesn't mean you couldn't have a service which by default has no hard cap, and people have to opt into it. You could even put a user interface thing where people have to type a whole sentence perfectly matching and hit OK, like "I understand that enabling a hard billing cap will shut off services if it exceeds my monthly quota". Wrap it in as much service agreement contract, TOS language as is necessary.
Heck, have it do the equivalent of send people a DocuSign equivalent PDF to sign acknowledging the risk before enabling it. Would it still stop pissed off people? Probably not. Would it help with the risk of lawsuits, very possibly.
As somebody developing hobby apps, there are things i just can’t use. Take an example cloudflare workers. I recently worked on a poc for an app. I ended up hosting it at home due to the uncertainty.
The caps were very generous but I wasn’t going to spend a minute worrying about what if a malicious actor took over
Yes, this is an important point although it has changed with A.I. Software is traditionally very high margin and so a $10k bill can be written off by the provider without any meaningful loss.
As a customer, the big number is scary and causes panic but for the provider… customers constantly fail to pay bills, providers are constantly writing off bills because it just isn’t worth the cost to chase, if a customer says “hey that usage was a mistake” it’s usually worth it to write it off to save the relationship. If you write off a big bill that wouldn’t have been paid anyway, the customer will perceive you as wonderful and benevolent and be loyal for life when they are ready to spend their money.
With tokens though the actual cost being incurred is much, much higher. If your service is just a wrapper around tokens, and a customer incurs $10k of usage that you paid OpenAI $5k for, it becomes much more difficult to write off.
Google Cloud is one of the few services that actually pursues unpaid bills even on their high margin services.
Google Cloud is also the scariest because of how much damage it can do and how bad their payment system can be.
I recently loaded up on prepaid api credits for gemini and it somehow triggered some billing shenanigans in my linked accounts where it said I had a negative balance (from the credits), and they were going to discontinue my services. I had to reset some settings to sort it out, mainly using their chat ai and mine (because theirs gave me right status info, but wrong conclusions).
It's pretty messy across like aistudio.google.com, and their typical console, and google workspace business account. I'd be so fucked if they froze my account, I'd rather just pay openrouter to access credits in the future.
> Google Cloud is also the scariest because of how much damage it can do
Google is the only cloud platform I'll not just never use, but personally discourage anyone from using it.
Simply because there have been way too many horror stories on here about people who had gotten their personal gmail accounts frozen for whatever BS reason - and absolutely zero recourse. With any other large service you can always get ahold of a human, with anything tied to Google it's impossible and even raising a major stink on HN or "legacy media" often does not help.
That will change as data centers become more "compute hubs" than service providers. When a good chunk of your bill is the cost of the electricity the compute center will have little margin to negotiate. Power companies do not care about customer mistakes.
My mobile phone contract reduces my speed a lot when the hard limit it reached. I still get some connectivity. What Never happens is that I have to pay extra. If a company can’t give me hard limits on the cost I have to pay, I will not do business with them. Notifications when close to the limit are fine, but when I don’t react to them I want things to crash rather than having to pay indefinite amounts.
Isn't that a soft limit? I guess it's a hard limit from the "zero additional cost" perspective but sift from the "stuff still works but in a degraded state" perspective.
It's indirectly a hard limit. You can calculate at the reduced speed that you're given exactly how much data you can consume after you reach that limit. And there's your true hard limit.
I think you may be confusing the issues here, especially if the "well known backend type service" you are referring to is AWS.
I got bit by some undocumented (or at least very poorly documented) hard limits in AWS when our app went viral. It was extremely difficult to just find out what these hard limits actually were. And most importantly, we never explicitly set them, they were just hidden defaults in AWS.
That's very different from having an easy to use, visual dashboard of where and what all your hard limits actually are, and make it it extremely easy to turn them off or on at a moments notice.
I'm of course lacking specifics, but it sounds like this take-away might not be the best solution?
A better approach to that scenario might be to split into base load and peak load infra, similar to how we do it with the power grid.
Base load could be an actual metal server that you pay x amount of money for and is fully yours, with peak load being handled by autoscaling dynamic stuff.
Both things being on a fixed budget, if that budget would run out, the service would not degrade as in "grind to a halt" but as in "slows down", which is probably not a downtime as part of a communicated SLA.
My point being that cloud and on-demand is useful, but a hybrid approach might in many cases make more sense. Though of course YMMV. I do not know what the requirements of your product are.
Depending on the size of the peak, overloading your base load could very well grind it to a halt. At times of overload, you need loadshedding to recover, which is exactly the opposite of what you are proposing here.
What use case do you have in mind? In my experience every use case I can think of won’t work with your model. The diurnal variation between nighttime demand and daytime demand (in a given timezone) is too great.
I can understand how that causes issues, but surely it also shows that alerts don't work either? These companies were presumably missing/ignoring alerts with scary messaging that they're about to get shut off, and would also miss/ignore alerts with scary messaging that their bill is getting too high.
I think it's not generally 'much better' to use alerts. People - end users - are by now quite used to seeing things go down for a while. No biggie. But a infra oopsie can kill a company in ways a short outage won't.
And its of course not just people yolo'ing with AI. People were quite capable of causing such outages themselves just fine. Distributed, serverless systems are hard.
Alerts are only good if somebody responds to them each and every single time. If you don't have someone with authority who can respond to every alert and will do that as part of their job, then they're worthless. Part of responding of course is understanding what it means.
Really it depends on the average loss for the service provider by customers that can't pay. If overages don't get paid by customers than all service providers will move to hard caps to prevent loss. If all providers do it, customers won't have much of a choice, except for higher caps with credit checks first.
It may be a dumb observation, but until I actually read the phrase "AI agents yolo infra in prod", I hadn't internalized that each agent is effectively always YOLOing! Because, of course they are. It is their purpose.
Why is this mutually exclusive. Is this even an argument? Some people complain becuase the hard cap kicked in. Others want hard caps. There is absolutely nothing stopping both from being implemented and offered as a choice. But we have spent decades being told "we know what's better for you". As an engineer, I think we need to come back down from the stratosphere.
> you can negotiate with the billing department at comparative leisure.
I'm curious. How likely is the billing department to waive off a huge bill as bad debt because an inexperienced builder misconfigured their infra or was hacked?
Not the OP but run a high margin service that has customers run up accidental bills often. Customers running up bills intentionally and then not paying is even more common. There is almost no situation where trying to force a customer to pay makes sense, we write off any amount without question. The goodwill is worth it every time. Most SaaS companies don’t even have the processes in place for debt collection anyway.
I don't see how it can be a problem if a service has mandatory settings for soft and hard limits. Those who don't want to disrupt services just unset the hard limit. Those who complain after the fact can be pointed at their explicit choices and have them explained to them - the service provider can hardly be held liable for a customer choosing the wrong value when it's easy to set the right value.
> Generally speaking, it's much better to use alerts instead of hard limit.
You can't extrapolate your own experience and generalize like that. There is a huge number of people who explicitly want hard caps, period. AWS giving in and finally offering this option after 20 years of people begging them to do that is a sign of that.
I mean, I understand why _you_ want it that way, but that doesn’t mean that there can’t be hard budget caps for other people. You could even have both, with alerts at a level lower than the cap. I just want some kind of control over it.
This goes beyond dollar charges. For production code to be reliable, everything needs to have a hard limit.
Queue lengths, request sizes, response wait duration, message payload size, authentication attempts, allocation rates -- there's always some upper number beyond which the system is so messed up you'd rather it crashes.
> An argument against this is that businesses don’t want their hosted applications to start throwing errors because some budget was exceeded. I expect that most businesses and individuals would prefer errors to a surprise $10,000+ bill.
Indeed. If you want a surprise $10,000 bill that's still not an argument against a hard cap -- just set it at $9,999,999 instead, or wherever you don't want the surprise bill. There's always a number that indicates something has gone insane. There's always a sensible upper limit to any operation.
I find Google AI Studio / Dev Platform / whatchamacallit one of the worst in this.
When video models just came out, I wanted to make a short clip. Gemini didn't have enough control back then, so I started in Studio. I had $10 on my account. Try, try, try, not good, retry, altogether maybe 20-30 retries with 4 choices for a 10 second video.
Wake up next morning to an email from Google - my Studio account is frozen because of negative balance.
I check it, it's at -$160. Not financial ruin, but a painful sum for a 10-second video I didn't even use in the end.
I don't think there was (or is) a setting in there that says "stop everything when I'm in the negative".
Almost a year later, it's still as bad. Trying to make some music via Lyria, I need to wait 10-12 hours for the API charge to show in my costs. Now I had the bitter lesson, I wait after every major novel experiment to see how much it costed me.
Had a similar experience with OpenRouter, put 10$, goes -50$, I just stopped using it. To their credit support agreed to reset my balance (AI refusal at first, then re-contacted me months later to agree). I'm being more careful now but I shouldn't have to be IMHO.
These shouldn't even exist without a negotiated contract.
I can subscribe to your service for a specific fee on a monthly basis ($20/month say), and take the risk of losing that month's fee if your service or I make mistakes, or I can choose to drop another $20 mid-month, or anything for convenience.
Saying that the computer will "control" the billing and can run haywire tells me that I don't want to be anywhere near your pile of bad incentives.
Monthly electricity bills are based on usage, and it works well but there’s a limit to how surprising a bill can be. The difference is the relative orders of magnitude you can be charged for these services you can go from 20$/month to 200k/month without warning.
You need to be aware of all ways outside people can rack up your bill. If you've got experience it's okay, but getting hit with a huge bill because you left an S3 bucket out in the open is something that has happened too many times to people starting out.
In the UK there are prepaid electricity providers: there's a card reader at the power panel. When you run out of credit, power cuts off, it resumes when you pay again.
>It’s such a basic thing, not having it has to be deliberate to make you accidentally spend more than you’d like.
This is unlikely. The big cloud providers routinely cancel bills based on accidental usage. The real explanation is just that it's technically difficult to calculate all usage in real time and shut down systems immediately as soon as a certain cap is reached.
> This is unlikely. The big cloud providers routinely cancel bills based on accidental usage
That is not an evidence alone. They also don't cancel bills in many cases. Most people also just swallow the damage and pay. Like gym subscriptions where people just pay because cancelling was too hard and they postpone the canceling because it takes so much effort. And they also think it is their fault of not cancelling, because it took that much of effort.
> The real explanation is just that it's technically difficult to calculate all usage in real time and shut down systems immediately as soon as a certain cap is reached.
That is trivial to add these days. Cloud providers usually say that "we care about your business and don't want to make your systems go down unexpectedly if you did not pay your bills".
The difference between Amazon and gym subscriptions is that Amazon can easily make more money by upselling to happy customers than it can by screwing individuals or small businesses out of relatively small amounts of money for their accidental usage.
In cpu compute margins are high and there is a lot of excess compute, you can give some away and make money.
In the age of AI this is gone. These companies have taken on huge debts. The operating margins are going to crash as interests increase. AI enabled fraud is going to make stolen service even more common.
That may be true, but I don't think it really affects the point. "Make people use our services accidentally and then force them to cough up" is just never going to be a good business plan for AWS. AWS makes money selling to big customers who willingly come back for more. AI might make it harder for AWS to turn a profit, but it's not going to make a lack of hard usage caps suddenly become central to AWS's business strategy.
I was a member of AWS customer advisory board for years. Two dozen of largest enterprises and unicorns and so on. Nobody wanted the billing department "inline" of systems controls.
AWS provides all the controls you need to flip the switch yourself, and provides the CDK / TF / etc. patterns for it. Plenty GitHub projects to host your own light switch service.
Amazon didn't have to build a hosted db, queue, mail or anything else that is value add as long as it provides basic hosted compute, storage and network.
But it still does and can only claim having those after building them.
Their customers have been asking for that hard cap feature for a long long time, not building that feature while saying customers have all the tools to build it themselves is just a deliberate misdirection at this point.
> Two dozen of largest enterprises and unicorns and so on. Nobody wanted the billing department "inline" of systems controls.
And of course, two dozen enterprises trump hundreds if not thousands of other small businesses right? Did AWS think of, oh I don't know, a 'feature flag' or somesuch thing that enterprises and unicorns are so fond of to provide this feature to those who don't have VC money to flush?
Again, they give you the tools for this. You can understand your usage by API, you can set threshold actions on it, you can switch things on and off, and applyable templates exist, all of which are under your control and "out of band" from the systems delivering traffic or compute to your paying customers.
Using these tools can you switch everything off immediately (or at least within a minute) of total billing going over $X? Because if not, you're being disingenuous
Not to mention that one of the things people want a hard limit for is to limit the financial blast radius of misconfiguration. Saying "oh sure, you can cobble that together yourself" is super missing the point.
The value proposition for software - and the reason why it's highly paid and has eaten the real world - is its "write once run forever (ish)" property. Unlike every other consumable, it doesn't expire.
In this proposed world where this static software is replaced with dynamic software ("AI will build it JIT"), doesn't that destroy the value proposition?
And in the world where normal software becomes AI-integrated ("software has intelligence "), isn't this tantamount to very very high hosting costs - where each execution now costs centi-dollars instead of nano-dollars?
Seems like the only use case where the cost model is equal or better is "AI writes the software".
Network saturation is difficult. Even if you turn off the endpoint you can still saturate the network in between. And it's still bandwidth.
I actually think network ACL triggers based on billing might be the only way to really enforce this.
I witnessed a DDoS attack once that changed how I think about billing. It was locally provisioned hardware and the attackers had saturated the switches. Naively I said "just block the CIDRs" but the problem was the incoming ram is so saturated that it can't even get to the point of "deny" in the firmware.
So from a technical perspective if there's an internal DDoS at AWS what do you do? Do you turn off the endpoint? Do you drop the sources from hitting it at the router? And even that costs money. Anyway that incident gave me a different level of appreciation for this challenge.
Edit: this is mainly targeted at the people complaining why this took so long. At some point in scaling even telling you "no sorry" in a nice way is expensive. I'm sure recruiters can sympathize with this nowdays.
Hard caps are rare because companies find it more profitable to forgive sympathetic individuals' bills while raking in profits from corporations whose services have gone awry
It’s incredible that in 2026, AWS and GCP are only just now introducing this. It’s possibly one of the most obviously needed features for a cloud provider.
Also, does anyone know why it’s taken this long? I suspect it’s a technical reason. While one could be cynical, I doubt it’s an intentional business/product decision. Hard spending caps are both excellent product differentiators and could possibly save these providers money as they don’t have to forgive their users when they accidentally over-use a service.
It's one part technical, one part a product decision. The technical part is that billing is not actually instant. As a most basic example, a VM reports its billing units every X period of time it is active. If there is some network blip but it's still running, then that billing data could be delayed.
The product level decision is that "shut down everything" is something the customers you want to target don't actually want. Are we including deleting RDS data? S3? Glacier storage? If so then the headline will just change from "Hobbyist got charged XXXXXX on AWS" to "Business literally had all their data deleted because a hacker took over their VM and mined bitcoin". The only people who really want this are hobbyists and it's not a market segment that's worth chasing. Easier to do the status quo of forgive afterwards then even open the can of worms of deleting all of a business's data and all their backups just because they had a 100k overrun.
Yeah, you can't just implement it as a pure "stop all services immediately once I hit a set amount"
It needs to be more like "don't allow spinning up additional services after you hit this amount", although that still allows you to go over the limit by a lot, since most services are billed hourly.
It really is difficult to implement a spending cap that doesn't risk shutting down important things.
It really is difficult to implement a spending cap that doesn't risk shutting down important things.
That's a checkbox decision for the customer. There needs to be the option of "This is important, never turn it off and I'll pay for any overages." versus "I want an entirely predictable bill up to $xxx, so stop my stuff as soon as possible over that."
It's not up to a cloud service to decide my website is more important than my money for me. That's my decision to make.
and they would still complain if they got it wrong - it's always the platform/company's fault.
Look at banks and fraudulent transfers that customers themselves get phished into doing. The bank in the end usually take the hit (after the customer complains long enough). That's why there's all sorts of hoops and such to prevent customers from failing - and that causes friction for people regularly.
Therefore, the cloud company's decision to default safer is more correct from this perspective.
You might be perfectly okay with having certain systems shut down, but you probably still want to pay for the archival storage of your important files.
That archival storage might be in several places, including one S3 bucket, whereas there might be another one that, contains copies of scraped Craigslist for X where you'd actually be happy that it just shut down.
This makes it far more complicated to do correctly, and as others pointed out, mostly relevant for hobbyist — this is not something you are going to make a lot of money from.
Better to spend engineering hours making an MCP for the dashboard or improving your Databricks setup.
Regarding hobbyists: some clouds, and also the bigger clouds, do target hobbyists quite a lot, presumably with the intention that some of those hobbyists will eventually grow and stick with them. So I don't think "not wanting hobbyists" is entirely true.
But yes, the amount of money they spend is less, so it makes less sense to implement features that only hobbyists want.
What surprised me about GCP was that yes, you could load all of your training data and model weights (terabytes) into a free starter account, then create a new account after 30 days and transfer the billing obligation from account A to account B.
It was such a blessing for hobbyists, back in ye olde 2019.
I think there is a middle ground between deleting data and allowing 5000 VMs to be created to mine bitcoin. Obviously there are a lot of different scenarios to consider but the explosive costs seem to be constrained mostly to a couple of features which would be fairly safe to cap.
There is a middle ground here: Make this setting configurable, off by default, and with all the associated warnings of what will happen if you switch it on. Better yet, make it settable on each billable service. Keep Route 53 going, but halt and delete any VMs that exceed some limit.
> The only people who really want this are hobbyists and it's not a market segment that's worth chasing. Easier to do the status quo of forgive afterwards then even open the can of worms of deleting all of a business's data and all their backups
Hit the nail on the head.
Nearly all businesses would prefer a cost overrun than services going offline.
and if they're going to be a good decision maker they need to put their personal feelings from hobbyist times aside and realize stuff is different in the two scenarios.
Meanwhile while in the real business world, Cloud decisions are shaped by marketing campaigns and it's the hobbyists and old grey-beards who provide the empirical push-back and reality checks.
If your peak monthly cost for AWS services as a business is $100k, do you not think setting a $200k spending limit is reasonable?
Obviously, the system should provide ample time by warning in advance of reaching it (and could even offer suggestion to keep it at N times your peak from M months ago).
If as a business you set your spending limits so tight that you frequently run into them and it's not some unusual activity, the problem is not that spending limits are available :)
It is mostly about protecting from the unknown, likely unbounded attack on your infrastructure, where your spend might grow 100x: even if you can take $100k, you might not be able to take $10M in a month.
And if you use enough services, a global per-account spending limit can't distinguish between peak usage versus one service being abused.
Imagine you spend $100k on average each month, but spend around Christmas rises to $1m because of the specifics of your industry. With a global spending limit, you can't distinguish between $200k of general spend increases due to Black Friday versus $200k in fraudulent 2FA SMS to South Sudan.
I agree but I think also, as a programmer, surely this is not an insurmountable problem. The idea that pops into my head immediately when hearing this problem is
1. of course spending limits are monthly or even weekly basis. You set a yearly limit, probably most people will do something like November to January 4 times the limits set rest of year.
2. as you near limit calls go to people on your team to tell them you are getting near your limit. Estimate is 2 hours, what do you want. Double Limit for this time period? Triple Limit for This Time Period? Remove Limit Entirely, you will get back to us with potential new limit? It's the beginning of Christmas, the next time your limit resets is the 3rd of January, remove limit entirely and you will get back to them. Why do you decide to remove limit entirely, because you have info that Amazon doesn't, specifically your assassin Nisse doll has gone viral for this Christmas season! The shit is making bank!!
3. When setting limit you say "expected low usage", "expected high usage", "limit". Limit should be significantly above expected high usage. Service informs you - you have been over your expected high usage by 20% last three time periods. Would you like to increase limit and high usage by 20%? Please Look at your settings otherwise.
4. Phone calls when there is an unexpected peaking in usage, like one hour we are 300 thousand which is very high for you, next hour it is 1.5 million.
Obviously none of this stuff helps hobbyists but even the worst run businesses I've worked at would handle this. Otherwise they deserve to be hobbyists, there's no reason to be an organization if you're not organized.
Obviously these things do not stop fraudulent attacks abusing your service, but it does make it harder for them, at the same time making it more difficult for your stuff to just get shut off without you knowing.
Of course, as a programmer I am aware that all these services are created by programmers as effective or more effective than I, and who have undoubtedly thought about it more than I, so I must also assume there are reasons why my off the cuff suggestions are ludicrously unhelpful, but I lack the knowledge as to why this should be so.
You are correct; however, you're also putting a lot of faith in people's ability to make decisions.
Another thing (that does not really apply at AWS anymore), is that todays's enthusiasts are going to be the future CTOs, and the easiest time to recruit them to your service is when they are still an enthusiast who gets to make decisions on their own because there's exactly one decision maker you have to appeal to and that person really likes to try new stuff.
That's why you can get a free fly.io and why we all use Tailscale. And it works too — if I was in charge and needed it, I would immediately go with Tailscale for a business; I know it and I use it.
for personal/hobby accounts sure. for a business, it’s much better to negotiate around billing or adjust systems/processes post-facto than it is to have service cut off unexpectedly.
debts are easier to manage when you have an active (ideally growing) customer base. you don’t have customers anymore if your cloud account takes down your service for the rest of the month due to spending limits.
If you are actually a business, a common warning that you are near the limit should mostly resolve it. It might only be tricky because the estimated time remaining is really short if it's a huge recent spike: eg. nobody is looking forward to a notice of "you'll use up your spending limit in 4h" on the weekend.
This type of warning should give you enough time to investigate if the warning is real and adjust the spending limits.
But then again, even if you hit them and your services get paused, you'd be increasing the spending limits and restoring services after you are back at work and notice they are down, so it mostly comes down to your incident response times.
> Are we including deleting RDS data? S3? Glacier storage?
If you're billing per GB of storage, then you can put hard caps on storage capacity, and then hard-reject any operation that would take the total stored size over that capacity.
A lot of billing systems are organized around event delivery. The system does what it does and reports usage. This reporting is asynchronous and can be done E.G. via cron jobs running on a 24 hour cadence in certain cases. There's an internal guarantee that billing records for a given period are delivered by a certain time. Nobody checks whether the user has enough money to do what they're trying to do, just whether they're authorized to access the system in the first place. Shutting down accounts due to non-payment is more of an abuse / fraud concern, and happens long after the bill is delivered.
Wow, I wish my my cell provider worked like that! Instead, they somehow manage to track and bill me literally in real time. How do they do it? Do the telcos have some magic pixie dust or what?
Storage is almost never the main thing racking up bills so it should be handled in a different way. If you hit a limit you can't store any more data but everything you have stays there and readable.
This is more about compute, VMs, LLM inference and services like hosted database. These are all safe to stop if the system triggers a normal shutdown when costs hit a limit.
I dunno, the largest accidental AWS overages I’ve triaged in my career (at very different companies) were all storage. Dangling EBS, forgotten S3 accumulating data after a deletion cron broke, incompletely retention-policy’d CloudWatch buckets, Aurora snapshots…the list is pretty long.
“Large” is relative to the business of course, but the biggest storage-related overages I’ve personally triaged are in the high $100ks to low $millions per month. Colleagues have heard of orders of magnitude more costly.
That’s not storage costs, that’s activity adding to storage used. The question is whether it would be cheap enough for them to keep all your existing data frozen until you paid the bill.
For S3 a reasonable option is a data limit as the primary limit. If you set it a TB over your real needs you only waste one dollar per day after it locks. And there's no need for any other paused service to charge more than that for idle data.
You're right that nobody wants deletion. Spending limits do not imply deletion.
I think a more sensible default is S3/etc locking access to data for 30 days, maybe even as little as 7 days, while you're able to still list and DELETE said data as you wish, without being able to get or put.
A very charitable take, in light of tech industry habits of exorbitant rent-seeking in scenarios of Platform Dominance (e.g. Google and Apple on the app store). We should remember AWS and Google companies are among the best in the world at A/B testing and extracting revenue from cloud services.
When you're one of only two real options out there, you can afford to demand users put up with things that on their surface seem ridiculous. Such as a billing system that (oops!) makes it difficult for customers to see where their costs are coming from, trim their largest sources of spend, notice meaningful changes in line item prices, or limit their spend. Wild how they can figure out a million different advanced services but gosh-darn-it can't figure out the hardtech of displaying line items.
Large enterprises can afford employees who are tasked full time with unwinding this capacity to mitigate the impact of these billing headaches. But I think this measure is introduced now because LLMs introduced a risk that these billing specialists could not control without caps.
I had a $.20/month recurring charge from AWS that I could only remove¹ by completely deleting my AWS account. That was enough to get me to give up on AWS for personal projects.
⸻
1. The key word was “I.” Maybe someone more skilled at navigating AWS’s menu structure than I could would have done it quickly and easily, but even though I knew what it was for, turning it off and not getting billed for it turned out to be a huge challenge. Thankfully it was only $.20, but if I were using the service for something that generated actual bills, that $.20 (and possibly more) would end up quietly siphoning money out of my pocket into Amazon’s).
Sounds familiar. I’m being billed £0.01/month for something in GCP, I don’t know what even after digging, but I’m too fearful to complain about it or disable the account lest it somehow gets my main Gmail account blacklisted somehow.
It's possible to get billed for things not exposed in the UI. ~10 years ago, I had one of these. A Glacier upload chunk was stuck in a staging area for over a month. I couldn't complete or cancel the upload, or remove the chunk. Support was able to, and they refunded the money without much hassle. But I did have to reach out to support (as a <$100 per month customer) to ask. Glacier was new, so I chalk it up to it being a new product, not anything greedy or malicious.
Remember when they decided the best UI experience was to give everything a vague abstract collection of shapes? Early 2010s or so. Couldn't tell a damn thing apart.
It's definitely technically difficult. You can't easily estimate how much an operation is going to cost before you kick off that operation, which means as soon as you get close to the limit you are at risk of tripping it.
Consider something like a "select * from bigtable" SQL query that might process a trillion rows. Hard to know that's going to cost $100 until after you have run it.
> You can't easily estimate how much an operation is going to cost before you kick off that operation
I recently had a debate with a colleague on this topic but concerning estimating the costs of AI agent work. For example, if you prompt an AI to refactor your codebase, the final cost can't be estimated perfectly, but I'm sure it can at least be estimated with some amount of precision! Like simply knowing that it will cost < $100 is actually great information even if the final work only ends up costing $5.
I think there actually might be a business opportunity (or at least the opportunity to build something cool here) if anyone wants to work in the AI cost estimation space. It's not exactly an idea I want to pursue, but just thought I'd put it out there. AI cost estimation (even with wide confidence bands) would be very useful to a lot of people.
Stop everything is pretty damaging any real business though. Things were better in the era of VPSs. You paid for a fixed amount of compute, if you ran a stupidly expensive operation than it just maxed out your system for a certain amount of time and things slowed down. But you didn’t kill the service entirely and you didn’t have unlimited potential price
In this example, that would require the “big query” to have billing baked into its actual query runtime, which isn’t impossible, just not how one would design a query planner per se. Usually such services emit metrics of usage units, then the billing calculation happens in a completely different system taking into account discounts, promotions, contracts, regional and currency differences, etc.
Suddenly a database, a storage service or a computer service needs to be aware of the billing situations and make behavioral decisions based on the billing status. Again, not impossible, but something that suddenly promotes billing from an async/non-crucial background service that can be paused, replayed, adjusted by account teams etc, into a crucial hot-path service.
The way you'd usually handle that AFAIK is to have the service ask the billing system for a "reservation" in its native units, likely with an attached TTL. Then, the service would translate those units to U.S. Dollars (or possibly Indian Rupees), taking your plan, discounts, vouchers, contracts, grandfathered pricing and all that into account. It would then "lock" the calculated amount of money, denying the reservation if total_spent + total_locked > spending_limit. After finishing the operation, the service would ask for actual billing and free the unused units.
Yes, thats one way to implement it. It still moves that global billing system into a ring 0 importance with 99.9999% availability and performance requirement, where before it wasn’t even in the picture. Not impossible, just suddenly every single AWS (or GCP or Azure etc) has a global single point of failure that gates all their performance and runtime actions.
Not to mention having no way to network isolate such service as every single service in every single network boundary needs to be able to contact such service, and every service needs (more or less) same auth permissions on all user accounts it’s running.
Again, solvable problems with enough code, but not simple problems by any means if you operate at a large scale.
Yes and then the choice is run it and forgive it, or, stop the process midway.
If you stop then you have to decide whether to charge for uncompleted work.
Interesting tradeoffs.
For very small ops e.g. individual Lambda invocation you have similar concerns especially if lots are fired at once from a queue or schedule or fanout.
If these cloud providers had needed a standard customer acquisition strategy to grow to their current size, hard caps and other “training wheels” features would already be in place to get people interested in and comfortable using the platform, with the hope of eventually getting a foothold into Enterprise like most SaaS startups have to do (“enjoy our product on a side project and then recommend us to your CTO!”). But AWS and GCP got to start as in-house providers for their own constellations of massive sites and back out from that to serving other hyper scale businesses first. The lack of friendly on-ramps and starter account features is a reflection of that origin more than anything.
Because most enterprise users would much rather have overages in billing than outages. The opportunity costs on any serious service I deploy dwarfs usage pricing, at least at the level a generic cloud can determine.
So it’s a feature the best customers don’t want, that adds risk to those customers deployments, to appease the worst customers.
At least historically. Perhaps Simon is right that the calculation has changed.
It’s not a binary decision though. Any sensible enterprise has many AWS accounts. Often hundreds or thousands. It’s the only clear separation of privilege.
There’s no reason for the majority of them to have an infinite budget cap. Prod? Sure. UAT? Why not. The sandbox environment Johnny just spun up to test some new agentic workflow? Hell no.
> While one could be cynical, I doubt it’s an intentional business/product decision
Why is egress more expensive than ingress in clouds? To lock in users.
Why cloudformation in some cases leaves s3 buckets laying around after destroy? To keep charging those cents.
Why no caps? To get the user into the mindset of we ll pay whatever they say, and charge the ones that dont notice or dont ask for the refund. Same reason why my newspaper subscription autorenews.
Nothing cynical about both of these. I think assuming it is technical is naive.
This is one of those features that customers think they want without having thought it through:
“Never let me spend more than $X” also means, “Shut down my business-critical app/service/solution at 2 am on a Sunday morning because Joel in IT forgot to plan for the new report runs.”
The product design work to let customers have the first thing without risk of major pain from the second thing is non-trivial.
Counter argument, this is the sort of thing that, especially for a smaller business or individual, can be the difference between a bad night and bankruptcy.
Sure it sucks that critical services blinked out at 2am. But what sucks even more is finding out the image on my ASG had a vulnerability that allowed someone to install a bunch of bitcoin miners which kept me fully scaled from midnight to 2am.
Or more likely, that a mistake in terraform 1000xed my spending.
Most people have predictable spending and could easily say "don't spend more than 10x what I normally spend". Or 1.5x, or 2x, 3x, etc. All depending on how they want to balance a runaway cloud expense.
For a large enterprise spending millions on AWS, 1.5X is already a budgetary disaster. Unfortunately, shutting off critical IT infra because it hit 1.4X spend this month is a business disaster.
There’s no magic wand that produces good outcomes when planning or execution goes awry at scale.
> especially for a smaller business or individual, can be the difference between a bad night and bankruptcy.
Just because this isn't a good solution for everyone, doesn't mean it's not a good solution for a large number of people and businesses.
A lot of businesses can tolerate outages. In fact, even very big businesses come out mostly unscathed when they have multi-hour outages. (how many is it for github this year?)
An outage causes a reputational black eye. It does not necessarily translate to lost income.
The alternate conversation is "the new report run had a bug and cost us $1,000,000 over the weekend" and I think that one's usually worse.
But one could just have two categories of service - the default capped plan, and a special Enterprise one where you sign a contract making it clear you understand the consequences of not having a budget limit.
Also, if your average usage is $900, set your hard limit at $2,000, not $1,000. Then when the report runs $500 over expected, you get a soft limit email and still have your report. Even a "business critical" run is probably not actually worth more than double your average spend.
This is the right perspective, but the folks who would be setting this cap for the customers that matter likely have no idea how to price that, or the price would be so absurd as to make the cap meaningless.
How much would a hospital pay to avoid unexpected downtime of their software systems?
I would think once you reach the scale that this becomes an issue you can afford someone or a team to be monitoring the system 24/7 able to respond to a price spike.
Price caps are for small scale stuff where you wake up on Monday and see 1000x the normal bill.
I imagine that the product folks at places like AWS are averse to introducing discontinuities in the experience based on scale. Little customers get the same experience as big customers who get the same experience as mega customers.
Obviously they’ve changed their mind about cost management in light of the scale and dynamism of agents, which isn’t too surprising.
A hospital needs to be able to handle a full cloud outage. So I'd be worried if they're near the top of the list of how much they'd be willing to pay here.
Sure, but prior to AWS introducing a notion of “project”, then they’d be at risk for those $1M bills from the data science team looking for agentic magic to reduce readmissions.
My point is this was never as simple as, “Give me a dial to set my maximum account spend.”
Big companies have thousands of budgets. An email is _worthless_. In fact, it would probably cause me to lose faith in a cloud that provided that as the control.
It is a technical reason. Basically cloud billing is much more granular and across many more services / line items than most things that basically the pipelines that figure out how much you have spent take a long time to know how much you have consumed. I believe all cloud providers with granular usage based billing have this problem.
> It’s incredible that in 2026, AWS and GCP are only just now introducing this. It’s possibly one of the most obviously needed features for a cloud provider.
GCP did have a budget cap previously. I think the new one is just more fine-grained to apply to specific services.
Simple: when data is added to S3, immediately bill the full cost of storing that data until the end of the month (or fail the upload request if there's not enough quota left). If the data is deleted before the end of the month, refund the remaining storage time.
To prevent cases where users can upload an excessive amount of data on March 31 that they then can't afford the April bill for, AWS should also maintain a "next month's balance" limit that gets handled in the same way.
It would take a bit more work to correctly handle things like ephemeral data and tiered storage classes, but it's not insurmountable.
You could apply the same kind of billing to most other long-lived resources, including VMs that run business-critical services.
There is a huge difference between "stop accruing new spend" and "nuke everything".
The horror stories I have seen are of the type: some big artifact was getting pulled in a loop, causing TBs of network traffic or access keys were leaked and malware spun up 1000 xxxlarge instances.
The ability to stop the bleeding is the bare minimum people want. Not, "Well, you made a boo-boo so now you lost everything."
As a hobbyist, I pulled completely out of AWS when I realized the billing issues could never be resolved. Even/especially with tight monitoring and hard caps.
Even as a hobbyist, even as a most careful and judicious architect and admin, I could not prevent my VPS from incurring costs beyond my control. That means that the entire Internet, anyone with some kind of material access to the VPS, and especially any user or authorized entity, they could incur costs to me without bound and without notice until slapped with a bill.
Even something as simple as egress charges aren't under your control. So if people download enough data, you pay for all of it? It seems like an absurd proposition.
It's like opening a business somewhere in a war zone, and vandals and squatters are constantly attacking your storefront, and maybe you have a band of toughs as security and some good cops to defend awhile, but you're utterly in a war zone with adversaries acting far beyond your control.
As a hobbyist, I could never again justify running a pure VPS with the Linux and stack on top, as I ran before like the MediaWiki server. I was excited to learn all the vocabulary and skillsets of cloud services, but on the "free tier" uncapped, there was no telling when I'd be presented with costs beyond my ability to handle. And I do not see how a Fortune 500 would have any different calculus in this regard.
Most businesses in recent memory were "their own landlord" of on-prem equipment and machine rooms. Yeah, they began to outsource even their IT admin, but the machinery was in-house until the cloud services took over. Did we go through a phase of collective machine rooms or data centers with a collection of tenants? It seems we skipped from "homeownership" to "feudalism" with the Cloud Providers being the Princes [beyond mere Lords] who provide minimal resources to the serfs now. I can see many corporate execs who begin to hate "AI Data Centers" just for what they have become: a very attractive and irresistable way to reduce your capex and footprint and physical plant, by "migrating to the cloud" but is that really a better status quo after all your revenue is being pumped into AWS?
And in view of what I just wrote, a "hard budget cap" checkbox is even worse for you than runaway costs, because it will allow any determined adversary, butterfingered DevOps, or innocent fuckup, to shut you down and deny your service by hitting that cap. When a business experiences a service or infrastructure outage, they count that in dollars of revenue. Your "budget cap" will cost you money because it "paused your project" and your customers all got burned in turn.
Moving toward real solutions, just spitballing, but I can envision throttling and capping of everything, every billable service, monitored by the cloud provider and ensured that your services don't spill over into unmanageable territory. If your spend could be throttled and capped by-the-minute, by-the-hour, daily, then there is a start for it. But really, any service that incurs costs to you should have reasonable rate-limiting, throttling, caps and alerts that can help you manage it. I think all that stuff is currently missing but I am not a cloud admin, obviously. Cloud services obviously have perverse incentives to open the floodgates and bill as much as possible for any possibly legitimate usage that doesn't exceed their aggregate capacities. It's like LLMs today that just burn tokens like there's no tomorrow, because someone [you, not Mexico] will eventually pay for it all.
Edit: Ugh it's fake. Literally only works for four random services, unsupported for all the rest. Completely useless for all of my projects. Also dumb that the only supported term is "monthly" considering that months are different lengths, and that they don't bother to account for credits or discounts. Maybe next decade they'll get around to implementing something useful.
The reason I will never use AWS again was all related to billing and budget caps. We had a small kubernetes cluster that failed to remove all of itself on a shutdown, it hung around like a ghost consuming resources for 3 months and nothing could see or find it, but we got billed!
Was a 6 month nightmare proving we did nothing wrong.
Let’s just call it what it is. We’ve been conned into accepting microtransactions for everything in computing. At the same time centralising and exposing everything publicly thus creating a huge risk surface area with respect to billing and security.
The whole idea was insane. We shunned paid compute services in favour of personal computing and then sold our souls back again.
Having a monthly summary or estimate of how your spending is going would be really useful, too.
Even if we have a negotiated yearly contract for $X spend per year, maybe we’ll hit that spend in 6 months instead of 12. Having some kind of automated telemetry saying how we’re trending would be so useful.
I’ve gotten vague warnings from customer success people saying vaguely that, but without any warning of what’s truly happening. The more we abstract away from money (tokens, credits, etc), the more we need a way to translate right back to money, to see how close we’re getting to any limits over time.
I don’t want to find out in month 5 that the contract which was expected to cover a year is now going to run out in 14 days.
Just imagine we could buy one piece of hardware where we could run all our software on, wouldn’t that be awesome? We could call it a „Personal Computer“ and it could start a revolutionary new way of doing business independent of gatekeepers! Freedom for free people!
This is something payment providers and banks should be offering. Otherwise you’re hoping each of N sellers will spend engineering money to individually implement a feature that realistically reduces their potential profit, on the vague promise that it’s some kind of beneficial feature that will bring them profit, which is never going to work like you hope.
My Dutch bank offers such a thing on SEPA direct debits. If a company sends a debit request, it must stay within X euros per month or otherwise it's not automatically accepted and I'll have to explicitly allow/reject it.
Rejection will block the payment and the company will likely go after you to settle the dispute.
Yeah, I don't think its particularly useful with hostile cloud providers that want to keep their legally unchallenged discretion over which of the clearly-not-inteded invoices to cancel and which ones to enforce.
But is is useful in related cases. I have now had two separate incidents where this would have saved me the hassle. Both were service providers which unilaterally tried to deviate from contract terms. One was possibly a unintended billing software upgrade where they suddenly added an extra item to a 10 year contract that did not belong there, the other was an apparently profitable "lets just try and see what percentage of customers notice" dick move. I should not have needed to care about this, both companies should have just been slapped with a hefty fee for requiring their and my bank to take a look, and then figure it out on their own instead of me having to call once to fix the problem for the future, and then call again to remind them about reimbursing the delta from past billing intervals.
It should be possible to open my bank app, see all my recurring payments (subscriptions), and be able to cancel them with one button. And this should count as an official termination.
I was amazed when, of all the banks, PayPal was the one that has this "show all subscriptions" screen (somewhere tucked away a few settings links deep).
Learnt recently DigitalOcean does not have hard budgets limits, only emails.
I am surprised and torn between this being intentional vs. them not caring.
One solution is banking apps that let you create extra cards and assigning hard limits to spending. Wise & Revolut apps can do that.
Have to mention OpenRouter also, I logged into OpenRouter through pi, on the auth screen it optionally let me set a limit which is very smart and user friendly.
It’s generally a solution in the US since invoicing for non-enterprise software is almost non-existent. The only time I’ve seen it was from a HN post complaining about being charged for after a free trial that didn’t require a CC.
You can still rack up a bill even if you can't pay for it. That solution just means the card declines when the monthly bill needs to be paid, not that you don't owe money.
I know AWS and similar sites have no such concepts, but other than those, this is an existing option of most kinds of service providers.
Realistically, most providers can't even allow you to consume $10k usage if they don't have the certainty that you can pay up. Prepaid is what gives them that certainty.
I thought Microsoft used to give student's some free Azure credits - say $200. Rumor is, Microsoft was very good at pulling the plug the instant you went over that billing limit.
This. For me it’s insane that the default is “you got the service, we send the bill after, and you’re expected to pay” and not “if I got the service without explicitly agreeing to pay set amount for it (without signs of fraud), it’s your loss for providing it”. Obviously things change for corporate agreements, but then both sides are expected to have lawyers.
Not making this political since this is true across countries and the various political spectrums, but you can see the same thing happening in national budgets in terms of programs set up by legislation that have built in financing formulas. Years later people wake up and realize the program is now spending an order of magnitude more money than anyone would have supported in the beginning. It's the same phenomenon over a much bigger scale and timeframe. If only they'd inserted a hard budget cap.
My org has a leaderboard for AI spending each month, and I have found it interesting how fast the distribution decays, just within the top 10 users. I often think “what did these people do with all those tokens?” It’s interesting to think the answer to that question is “maybe not a lot?”
Right? If spending the most is lauded, why wouldn't I use the most expensive model, automate things that don't need automating, build things I don't need to build etc just to jack the spend up?
And everywhere in Europe you have clear price sheets (unlike the deliberately opaque mess of the US price sheets with more small-print than a packet of pills), which means even if you are at an EU provider with no hard caps you can still accurately reason and predict your costs.
Just a few examples....
Cloud providers:
- Upcloud
- Exoscale
Inference providers:
- Verda
- PrivateMode
I really don't buy the stories the US providers tell you that "its too difficult" or "what if you suddenly go viral".
The "viral" bit is easily solved through basic monitoring of metrics that everybody should be doing. I believe the cool-kids give it the fancy name of Site Reliability Engineering (SRE). All you need to do is top-up your balance / adjust your cap if your metrics are trending upwards for an explainable reason. Its not rocket science.
As for the "too difficult" that's just a lie. It just suits the US cloud providers better to have you spend spend spend on their messy soup of random interdependent microservices.
Nice to see the tech giants finally launch hard spending caps in 2026, matching some but not all of the resource/spend limiting features ($, compute, disk, tape, print, time) that were baked into mainframes and time sharing systems in the early 1970s! Everything old is "modern" again!
I call this "blast radius". Earlier this year during the Claw hype I was reading about all kinds of elaborate schemes to prevent the agent from getting API keys.
I realized, what am I actually afraid of. Well, overspend. So I just set then all to disable auto-reloading. Now if it blows up, I'm down $5.
Same story with containers. Just give it root on a VPS, and if it blows up, I'm down $3.
It should definitely be default for all payed APIs. I think every AI provider have this by default.
I've had an AI coding agent enter a doom loop for hours multiple times now. having a budget limit for my openrouter key is helpful, but it's already spent then.
I've been building a thing you can hook into ai workflows that uses multiple detectors if an agent is starting to loop and will then send a kill command to that chain. I don't have any testers for it though.
To create these caps wouldn’t AI companies need to tell us how much we are spending in realtime? Not sure they want to do that. The complete lack of transparency is a feature and not a bug. Just like in the healthcare space.
The challenge with cloud has always been the selling point of “infinite scaling” vs the flip side “infinite billing”.
It’s not in their interest to make cost controls work well.
Ideally you’d be able to set something granular like “allow this service to scale up only 10x, measured at an hourly level, and alert me when it happens. Drop all requests that exceed 10x”.
And then you are mostly in a throttling situation until the burst clears or a human can review & accept increased usage is ok/increase thresholds. I’d rather have services go slow during excess load (like a real server) than go dark for remainder of month.
This seems a lot better than brute force “turn everything off at $X level of monthly billing” or “no limits you can charge me infinity dollars”.
More than just hard budget caps, we need agents to be aware of those hard budget caps, and plan accordingly. Too much agent behavior is driven by the immediate proximal goal instead of long term, strategic considerations. What we have now is agents deciding to on embark on expensive, circuitous routs to a goal, perhaps with potential but not certain in any case, benefits, and intervention depends on either an attentive user or forces stop after the money pot dries up. Prior research (#?) has Dem nsttated agent behavior to adjust behavior in response to budget considerations that appeared to enable more efficient token usage, and if that means it needs to be imposed on a custom harness (and not the self interested ai provider), so be it.
Surprised and happy to see this. Made an account just to comment.
Just last week I was looking for similar functionality in Cloudflare as I am exploring publicly exposing apps I have built (as opposed to just me and some friends on a home lab).
My main concern on public cloud platforms is costs. I never want to spend more than, say, 10 EUR on a simple service.
In my mental risk matrix, the chance of a cost-related issue has been increasing more-and-more. I am creating and hosting more than ever, and the chances of malicious activity/abuse are in my opinion higher than ever before.
It feels like cloud providers either are unaware of this, or riding a wave of increased turnover - which I think a more likely.
I learned this the hard way with OpenAI last week. A key got hacked and a Chinese-language bot used $300 in tokens in an hour.
I had a spending limit on for $30, so why did it keep charging? Because the spending limit is meaningless without a hidden checkbox called “enforce spending limit,” which is (or at least was for me) off by default.
I am a little bit surprised by the lack of depth in discussing such a major product change.
I get it, the problem is definitely worth solving for. Waking up with a $100k bill isn’t great.
At the same time, from a product perspective the proposed solution might be a bad idea. Simply having hard caps as default will definitely turn out to be as bad for some people as a $100k bill, see for example (1).
You can’t come up with good product changes if you don’t discuss the potential negative effects of a change.
You could set the hard cap to a billion dollars. Everyone has a different hard cap — every company and person has an amount they simply couldn’t pay. This is just asking you to decide what that is and write it down.
Then add the option, but make it default to soft caps. Or bake the default into the tiers. If I sub with a $20 spending limit, I will probably be ruined by a $2000 bill. On the other hand, if a company subs with a $1000 limit, it can probably cover the occasional $10000.
At the very least, there should be an optional hard limit that is obviously indicated in the UI. When you're signing up where you set your "usage cap" warning, next to it should be an optional hard cap with big bold red letters "THIS WILL CUT YOU OFF THE MOMENT YOU GO ONE CENT OVER". So I can set e.g. a warning at $20 and a hard cap of $100.
> In an ideal world, our agents could help with this.
Not exactly the sort of case Simon has in mind, but I tell Claude to keep to hard daily limits on its OpenRouter spending for two long-running projects [1, 2].
A Routine for each project fires ever few hours, and Claude decides itself what to do in each session. It does tasks that require calls to other models through OpenRouter only when it is still within its daily budget for that project; after it reaches that cap, it does other tasks that don’t require extra spending.
I work at a place that has an eight figure monthly AWS bill. They won’t use this.
I had a personal development account for ~15 years. I tinker with infrastructure stuff and had built some centralized event reporting. One day about two years later I turned on sqs data events into cloudtrail. What I didn’t realize was that this closed a feedback loop and over the next couple of hours my run rate went to about $4k per day in cloudtrail+sqs usage.
I didn’t realize it until I hit the next months billing alarm immediately the next month. I’d racked up $25k in usage fees.
I’ll be using this feature. Nothing I run is worth that risk.
We always did, the clouds convinced us that overages were the norm. You can blame credit ratings as another vector for big business to screw everyone over. Everything should have been pay in advance with an alternate billing method for overages if you want it.
I had an api key set to read only that somehow ran up a $400 bill, I contacted openai about it and never heard back. Not quite the same thing, but still, I find this very annoying.
When I've used GCP in the past this has really really annoyed me.
Someone told me some years ago that it was "impossible" for Google to apply an upper usage caps due to technical design reasons which I found absurd. You can build a globally distributed continent-scale SQL database with consistency guarantees, but you can't stop me paying more than X USD when my usage ticks over that? Huh?
It smelt much more to me like Google's business model made it impossible to stop customers spending money, not their engineering.
I understand this is snark, but if you think about it, this is already implemented in electrical infrastructure. If I use too much power, the circuit breaker trips to protect me and protect the electrical grid. OP is about a billing breaker, but the parallels should be obvious.
If the CEO of the electric company didn't mandate circuit breakers, he should go to jail.
The peak throughput does result in an overall monthly limit though. For a house with a 200A main breaker, that effectively limits your electric bill to $7,000/month, which is very reasonable compared to the tens of thousands of dollars in a single day that a lot of cloud billing disasters end up costing.
Why do people think new laws are needed to solve every last problem in the world?
Google implemented caps because their competitors offered them. Before that, customers could choose one of several competitors, rent the GPUs at a fixed rate, or buy the GPUs and install them on-premises.
At no point was any law needed to solve any of this.
> Nobody wants to wake up to an email sent at midnight warning about a budget limit and find that, while they slept, their rogue service had consumed several hundred (or several thousand) more dollars of usage.
Monitoring and a circuit breaker. If you are given the tools, make your own heuristic and flip the breaker yourself. Don't let someone else turn it off, and be hostage to their process failures for getting it back on. By self-selection, if you are a hard limit customer, you are not front of their service line.
> An argument against this is that businesses don’t want their hosted applications to start throwing errors because some budget was exceeded. I expect that most businesses and individuals would prefer errors to a surprise $10,000+ bill.
Based on a career serving the Fortune 50, CTO for trillion dollar bank, etc.: "No." Assuming actual business is being done by the machinery in question...
The lights must stay on.
Cost management cannot shut off the enterprise. Cascading costs, even before reputation, are incalculable.
Ubicloud does not have hard budget caps, which I only realized this morning after moving all my CI over to them over the past few months. Fortunately I didn't learn the hard way.
It'd be nice if more than AI spend worked this way, autoscaling is almost a mixed blessing because unpredictable pricing can be worse than the cost savings...
Counterpoint: if you can automate API calls on the client side, why can't you automate billing caps? If you want a machine that can run 24-7 and make money for you while you sleep (which let's face it is the motivation for a lot of AI takeup), isn't the onus on you to install cicuit-breakers?
Because many cloud services have incredibly complex or opaque pricing structures that make it difficult to impossible to determine how much something is going to cost you ahead of time, especially if it's usage-based a la network egress (and the usage statistics don't update frequently enough to make such circuit breakers possible to implement client-side).
They might not be able to predict your bill but how much time do they need to add up what you already spent to minimize your overage? And TBH how much time should be acceptable to exceed your cap before it's their fault for the lag in their software.
I would just not sign up for a service without price transparency, or pre-calculate my liability based on available information before pushing the (metaphorical) Deliver Now button.
Making incredibly complex and opaque pricing structures is not necessary for the providers to charge for and make a profit on their service. And being technically difficult is a lazy excuse. Cloud platforms have to solve many, much more difficult challenges to offer their services at all, they just don’t want to invest the time in more customer friendly billing because they expect it will result in reduced revenues.
I know your post isn’t explicitly defending the platforms, but the arguments they use feel transparently flimsy.
I'm not sure where you got the impression that I'm making excuses for cloud providers. I'm just stating the way things are, not the way I think they should be.
Yes and no. I suspect many of the hard limits were set arbitrarily, and we'll see a relaxation of limits as people get frustrated with the limited use they get out of them. And some services will genuinely need to be re written to support higher rps or risk losing customers
The solution is to not give agents access to MCP servers.
The entire MCP ecosystem is ludicrous. You’re paying for inference for an agent to make the same decisions over and over again, and yet the actions they’re taking can be so easily written by those same agents into a bash script you can run again and again, deterministically and for free.
MCP is the problem. Having agents “use a product on your behalf” is the problem.
The pattern you’re looking for is that agents should write scripts that use products for us, and that doesn’t need a new protocol.
This is nothing new. This has been defacto standard for companies using cloud, which has burst or semi-predictable spikes in usage.
But the idiocy is incredible, even allowing for this to be happen in a business is so infantile that the only hard cap that should be important is not to allow stupid people in the machine room.
I'd support this provided we have the converse as well: if the customer doesn't pay their bill on time, the service gets shut down immediately. (Disclosure: I sell SaaS services to people who don't pay their bills on time).
Because it’s unlikely they’ll actually be able to collect that million dollars from a lot of those customers. Rephrased: why would your vendor want to make it harder to accidentally give you a million dollars of services in exchange for debt of dubious quality?
One of the biggest benefits of not engaging with LLMs or any of this nonsense is you dont have to care about all these "self made" problems of the LLM-gliteratti.
AWS, is a loot box... Tokens are just in game currency, and that sales person is just metrics that have identified your spending as making you a whale.
Your average CTO from the last decade turned a fixed cost into variable spending that looks like a mobile game.
This is another reason why cryptocurrency wins. Hard limits are baked in. There is no automatic charge like with a credit card or debit card or bank account. The payer has to initiate the payment.
the premise seems a bit faulty to me. why should we be giving next token predictors access to spend our money? like what great benefit do we get from this that we should allow them unfettered access, but with safeguards in the form of hard budget caps?
I don't think Simon means you should hand off the spending to agents/LLM (which would also make me uneasy) but that if you're probing one for hosting/SaaS providers they should default to recommending ones with budget caps
This is about budget caps (“I don’t want to spend more than $100, cut me off once I spend that much”) not price caps (“no one is allowed to charge more than this price per token”).
Guy seems to have lost his mind to AI psychosis and is just posting unmitigated obvious low-grade crap of late, which gets picked up here by his fan base.
whttering / wittering: To chatter, babble, or ramble on at length about trivial matters. About right.
It was, unfortunately, a nightmare. There were tons of tickets and even threats of lawsuits from customers whose service got cut off hard at the worst possible time due to organic growth/going viral/big event/nobody knew about the limit/etc. Not only did they lose all the leads and revenue they would have gotten from that bump, but they also pissed off their own existing users who suddenly couldn't use the service either.
Generally speaking, it's much better to use alerts instead of hard limit. Even in the worst case (hackers pwn your credentials and mine Bitcoin or whatever) the rest of your business is unaffected and you can negotiate with the billing department at comparative leisure.
This is all assuming you have humans operating the service. If you're letting AI agents yolo infra in prod, you have a whole series of new problems.
Probably for companies and opportunistic individuals with a high risk tolerance. As for me, just give me hard budget caps. It shouldn't be a choice between supporting one of the approaches, when you could let each client pick what they want.
Want an alert? Configure one.
Want a hard limit? Configure one.
(just make the configuration easy and visible)
1000 years ago, before CI/CD took over folk would have launch/go-live events, where teams would do a good checklist. There were pre-launch meetings to review, which included load checking.
Heck, have it do the equivalent of send people a DocuSign equivalent PDF to sign acknowledging the risk before enabling it. Would it still stop pissed off people? Probably not. Would it help with the risk of lawsuits, very possibly.
The caps were very generous but I wasn’t going to spend a minute worrying about what if a malicious actor took over
As a customer, the big number is scary and causes panic but for the provider… customers constantly fail to pay bills, providers are constantly writing off bills because it just isn’t worth the cost to chase, if a customer says “hey that usage was a mistake” it’s usually worth it to write it off to save the relationship. If you write off a big bill that wouldn’t have been paid anyway, the customer will perceive you as wonderful and benevolent and be loyal for life when they are ready to spend their money.
With tokens though the actual cost being incurred is much, much higher. If your service is just a wrapper around tokens, and a customer incurs $10k of usage that you paid OpenAI $5k for, it becomes much more difficult to write off.
Google Cloud is one of the few services that actually pursues unpaid bills even on their high margin services.
I recently loaded up on prepaid api credits for gemini and it somehow triggered some billing shenanigans in my linked accounts where it said I had a negative balance (from the credits), and they were going to discontinue my services. I had to reset some settings to sort it out, mainly using their chat ai and mine (because theirs gave me right status info, but wrong conclusions).
It's pretty messy across like aistudio.google.com, and their typical console, and google workspace business account. I'd be so fucked if they froze my account, I'd rather just pay openrouter to access credits in the future.
Google is the only cloud platform I'll not just never use, but personally discourage anyone from using it.
Simply because there have been way too many horror stories on here about people who had gotten their personal gmail accounts frozen for whatever BS reason - and absolutely zero recourse. With any other large service you can always get ahold of a human, with anything tied to Google it's impossible and even raising a major stink on HN or "legacy media" often does not help.
I got bit by some undocumented (or at least very poorly documented) hard limits in AWS when our app went viral. It was extremely difficult to just find out what these hard limits actually were. And most importantly, we never explicitly set them, they were just hidden defaults in AWS.
That's very different from having an easy to use, visual dashboard of where and what all your hard limits actually are, and make it it extremely easy to turn them off or on at a moments notice.
A better approach to that scenario might be to split into base load and peak load infra, similar to how we do it with the power grid.
Base load could be an actual metal server that you pay x amount of money for and is fully yours, with peak load being handled by autoscaling dynamic stuff.
Both things being on a fixed budget, if that budget would run out, the service would not degrade as in "grind to a halt" but as in "slows down", which is probably not a downtime as part of a communicated SLA.
My point being that cloud and on-demand is useful, but a hybrid approach might in many cases make more sense. Though of course YMMV. I do not know what the requirements of your product are.
And its of course not just people yolo'ing with AI. People were quite capable of causing such outages themselves just fine. Distributed, serverless systems are hard.
Anyone who says its 'much better' to use alerts instead of hard caps needs to Google 'alert fatigue'.
Alerts are soft. You ignore them or miss them, nothing happens except you spending $$$$$$$$$$ more.
Great if you're the cloud provider raking in the cash, but a poor way to run your infrastructure.
Hard caps force you to implement correctly (to control costs in the first place) and have correct monitoring in place (to keep a healthy cap buffer).
So it means you can't just vibecode some slop and blindly devops it via Github CI/CD. You actually need to think and reason about your infrastructure.
It should be the choice of the customer.
Imposing one option: only hard budget cap or only alerts is the wrong design decision.
I'm curious. How likely is the billing department to waive off a huge bill as bad debt because an inexperienced builder misconfigured their infra or was hacked?
This is not rocket salad to implement correctly.
You can't extrapolate your own experience and generalize like that. There is a huge number of people who explicitly want hard caps, period. AWS giving in and finally offering this option after 20 years of people begging them to do that is a sign of that.
Queue lengths, request sizes, response wait duration, message payload size, authentication attempts, allocation rates -- there's always some upper number beyond which the system is so messed up you'd rather it crashes.
> An argument against this is that businesses don’t want their hosted applications to start throwing errors because some budget was exceeded. I expect that most businesses and individuals would prefer errors to a surprise $10,000+ bill.
Indeed. If you want a surprise $10,000 bill that's still not an argument against a hard cap -- just set it at $9,999,999 instead, or wherever you don't want the surprise bill. There's always a number that indicates something has gone insane. There's always a sensible upper limit to any operation.
When video models just came out, I wanted to make a short clip. Gemini didn't have enough control back then, so I started in Studio. I had $10 on my account. Try, try, try, not good, retry, altogether maybe 20-30 retries with 4 choices for a 10 second video.
Wake up next morning to an email from Google - my Studio account is frozen because of negative balance.
I check it, it's at -$160. Not financial ruin, but a painful sum for a 10-second video I didn't even use in the end.
I don't think there was (or is) a setting in there that says "stop everything when I'm in the negative".
Almost a year later, it's still as bad. Trying to make some music via Lyria, I need to wait 10-12 hours for the API charge to show in my costs. Now I had the bitter lesson, I wait after every major novel experiment to see how much it costed me.
Coding is solved, my ass.
I can subscribe to your service for a specific fee on a monthly basis ($20/month say), and take the risk of losing that month's fee if your service or I make mistakes, or I can choose to drop another $20 mid-month, or anything for convenience.
Saying that the computer will "control" the billing and can run haywire tells me that I don't want to be anywhere near your pile of bad incentives.
So my powerbills are predictable.
Whereas traffic spikes to websites are not.
This age of abusive AI crawlers and the non-revenue generating traffic has been a very real problem for me!
But it would be trivial for the platform to rate limit traffic.
I think they explicitly said that.
It’s such a basic thing, not having it has to be deliberate to make you accidentally spend more than you’d like.
This is unlikely. The big cloud providers routinely cancel bills based on accidental usage. The real explanation is just that it's technically difficult to calculate all usage in real time and shut down systems immediately as soon as a certain cap is reached.
That is not an evidence alone. They also don't cancel bills in many cases. Most people also just swallow the damage and pay. Like gym subscriptions where people just pay because cancelling was too hard and they postpone the canceling because it takes so much effort. And they also think it is their fault of not cancelling, because it took that much of effort.
> The real explanation is just that it's technically difficult to calculate all usage in real time and shut down systems immediately as soon as a certain cap is reached.
That is trivial to add these days. Cloud providers usually say that "we care about your business and don't want to make your systems go down unexpectedly if you did not pay your bills".
The difference between Amazon and gym subscriptions is that Amazon can easily make more money by upselling to happy customers than it can by screwing individuals or small businesses out of relatively small amounts of money for their accidental usage.
In cpu compute margins are high and there is a lot of excess compute, you can give some away and make money.
In the age of AI this is gone. These companies have taken on huge debts. The operating margins are going to crash as interests increase. AI enabled fraud is going to make stolen service even more common.
I was a member of AWS customer advisory board for years. Two dozen of largest enterprises and unicorns and so on. Nobody wanted the billing department "inline" of systems controls.
AWS provides all the controls you need to flip the switch yourself, and provides the CDK / TF / etc. patterns for it. Plenty GitHub projects to host your own light switch service.
Don't relinquish the switch.
But it still does and can only claim having those after building them.
Their customers have been asking for that hard cap feature for a long long time, not building that feature while saying customers have all the tools to build it themselves is just a deliberate misdirection at this point.
And of course, two dozen enterprises trump hundreds if not thousands of other small businesses right? Did AWS think of, oh I don't know, a 'feature flag' or somesuch thing that enterprises and unicorns are so fond of to provide this feature to those who don't have VC money to flush?
In this proposed world where this static software is replaced with dynamic software ("AI will build it JIT"), doesn't that destroy the value proposition?
And in the world where normal software becomes AI-integrated ("software has intelligence "), isn't this tantamount to very very high hosting costs - where each execution now costs centi-dollars instead of nano-dollars?
Seems like the only use case where the cost model is equal or better is "AI writes the software".
I actually think network ACL triggers based on billing might be the only way to really enforce this.
I witnessed a DDoS attack once that changed how I think about billing. It was locally provisioned hardware and the attackers had saturated the switches. Naively I said "just block the CIDRs" but the problem was the incoming ram is so saturated that it can't even get to the point of "deny" in the firmware.
So from a technical perspective if there's an internal DDoS at AWS what do you do? Do you turn off the endpoint? Do you drop the sources from hitting it at the router? And even that costs money. Anyway that incident gave me a different level of appreciation for this challenge.
Edit: this is mainly targeted at the people complaining why this took so long. At some point in scaling even telling you "no sorry" in a nice way is expensive. I'm sure recruiters can sympathize with this nowdays.
The fancy alternative is to anycast that network and do distributed filtering so that the flood is manageable.
The first can be done by your ISP but it does lead to the temporary loss of traffic via that IP.
The second one is what DDoS mitigation services do and AWS has a basic one built in ( AWS Shield ) and several additional services you can get.
Plus, nobody wants to be the fired PM who said "I spent our eng. hours to achieve -20% revenue".
Also, does anyone know why it’s taken this long? I suspect it’s a technical reason. While one could be cynical, I doubt it’s an intentional business/product decision. Hard spending caps are both excellent product differentiators and could possibly save these providers money as they don’t have to forgive their users when they accidentally over-use a service.
The product level decision is that "shut down everything" is something the customers you want to target don't actually want. Are we including deleting RDS data? S3? Glacier storage? If so then the headline will just change from "Hobbyist got charged XXXXXX on AWS" to "Business literally had all their data deleted because a hacker took over their VM and mined bitcoin". The only people who really want this are hobbyists and it's not a market segment that's worth chasing. Easier to do the status quo of forgive afterwards then even open the can of worms of deleting all of a business's data and all their backups just because they had a 100k overrun.
It needs to be more like "don't allow spinning up additional services after you hit this amount", although that still allows you to go over the limit by a lot, since most services are billed hourly.
It really is difficult to implement a spending cap that doesn't risk shutting down important things.
That's a checkbox decision for the customer. There needs to be the option of "This is important, never turn it off and I'll pay for any overages." versus "I want an entirely predictable bill up to $xxx, so stop my stuff as soon as possible over that."
It's not up to a cloud service to decide my website is more important than my money for me. That's my decision to make.
and they would still complain if they got it wrong - it's always the platform/company's fault.
Look at banks and fraudulent transfers that customers themselves get phished into doing. The bank in the end usually take the hit (after the customer complains long enough). That's why there's all sorts of hoops and such to prevent customers from failing - and that causes friction for people regularly.
Therefore, the cloud company's decision to default safer is more correct from this perspective.
I mean, it's really not unless they don't build it in the first place—that really is the platform's fault.
You might be perfectly okay with having certain systems shut down, but you probably still want to pay for the archival storage of your important files.
That archival storage might be in several places, including one S3 bucket, whereas there might be another one that, contains copies of scraped Craigslist for X where you'd actually be happy that it just shut down.
This makes it far more complicated to do correctly, and as others pointed out, mostly relevant for hobbyist — this is not something you are going to make a lot of money from.
Better to spend engineering hours making an MCP for the dashboard or improving your Databricks setup.
But yes, the amount of money they spend is less, so it makes less sense to implement features that only hobbyists want.
> If you take no action within 90 days of your project being paused, AWS permanently deletes your project data.
From https://docs.aws.amazon.com/accounts/latest/reference/create...
It was such a blessing for hobbyists, back in ye olde 2019.
They'll ban you after a year because it will be against their TOS.
But sure, go for it.
Hit the nail on the head.
Nearly all businesses would prefer a cost overrun than services going offline.
Some were replaced with a competing service which has a limit, others replaced by a self-hosted alternative.
I think many small businesses would prefer to be offline or have a degraded service than pay $X000.
Obviously, the system should provide ample time by warning in advance of reaching it (and could even offer suggestion to keep it at N times your peak from M months ago).
If as a business you set your spending limits so tight that you frequently run into them and it's not some unusual activity, the problem is not that spending limits are available :)
It is mostly about protecting from the unknown, likely unbounded attack on your infrastructure, where your spend might grow 100x: even if you can take $100k, you might not be able to take $10M in a month.
Imagine you spend $100k on average each month, but spend around Christmas rises to $1m because of the specifics of your industry. With a global spending limit, you can't distinguish between $200k of general spend increases due to Black Friday versus $200k in fraudulent 2FA SMS to South Sudan.
1. of course spending limits are monthly or even weekly basis. You set a yearly limit, probably most people will do something like November to January 4 times the limits set rest of year.
2. as you near limit calls go to people on your team to tell them you are getting near your limit. Estimate is 2 hours, what do you want. Double Limit for this time period? Triple Limit for This Time Period? Remove Limit Entirely, you will get back to us with potential new limit? It's the beginning of Christmas, the next time your limit resets is the 3rd of January, remove limit entirely and you will get back to them. Why do you decide to remove limit entirely, because you have info that Amazon doesn't, specifically your assassin Nisse doll has gone viral for this Christmas season! The shit is making bank!!
3. When setting limit you say "expected low usage", "expected high usage", "limit". Limit should be significantly above expected high usage. Service informs you - you have been over your expected high usage by 20% last three time periods. Would you like to increase limit and high usage by 20%? Please Look at your settings otherwise.
4. Phone calls when there is an unexpected peaking in usage, like one hour we are 300 thousand which is very high for you, next hour it is 1.5 million.
Obviously none of this stuff helps hobbyists but even the worst run businesses I've worked at would handle this. Otherwise they deserve to be hobbyists, there's no reason to be an organization if you're not organized.
Obviously these things do not stop fraudulent attacks abusing your service, but it does make it harder for them, at the same time making it more difficult for your stuff to just get shut off without you knowing.
Of course, as a programmer I am aware that all these services are created by programmers as effective or more effective than I, and who have undoubtedly thought about it more than I, so I must also assume there are reasons why my off the cuff suggestions are ludicrously unhelpful, but I lack the knowledge as to why this should be so.
Another thing (that does not really apply at AWS anymore), is that todays's enthusiasts are going to be the future CTOs, and the easiest time to recruit them to your service is when they are still an enthusiast who gets to make decisions on their own because there's exactly one decision maker you have to appeal to and that person really likes to try new stuff.
That's why you can get a free fly.io and why we all use Tailscale. And it works too — if I was in charge and needed it, I would immediately go with Tailscale for a business; I know it and I use it.
for personal/hobby accounts sure. for a business, it’s much better to negotiate around billing or adjust systems/processes post-facto than it is to have service cut off unexpectedly.
debts are easier to manage when you have an active (ideally growing) customer base. you don’t have customers anymore if your cloud account takes down your service for the rest of the month due to spending limits.
This type of warning should give you enough time to investigate if the warning is real and adjust the spending limits.
But then again, even if you hit them and your services get paused, you'd be increasing the spending limits and restoring services after you are back at work and notice they are down, so it mostly comes down to your incident response times.
If you're billing per GB of storage, then you can put hard caps on storage capacity, and then hard-reject any operation that would take the total stored size over that capacity.
This is more about compute, VMs, LLM inference and services like hosted database. These are all safe to stop if the system triggers a normal shutdown when costs hit a limit.
I’m not saying it’s normal. Just that AWS is complicated, and people use it in a plethora of ways, thus billing caps aren’t easy to get right either.
“Large” is relative to the business of course, but the biggest storage-related overages I’ve personally triaged are in the high $100ks to low $millions per month. Colleagues have heard of orders of magnitude more costly.
You're right that nobody wants deletion. Spending limits do not imply deletion.
When you're one of only two real options out there, you can afford to demand users put up with things that on their surface seem ridiculous. Such as a billing system that (oops!) makes it difficult for customers to see where their costs are coming from, trim their largest sources of spend, notice meaningful changes in line item prices, or limit their spend. Wild how they can figure out a million different advanced services but gosh-darn-it can't figure out the hardtech of displaying line items.
Large enterprises can afford employees who are tasked full time with unwinding this capacity to mitigate the impact of these billing headaches. But I think this measure is introduced now because LLMs introduced a risk that these billing specialists could not control without caps.
⸻
1. The key word was “I.” Maybe someone more skilled at navigating AWS’s menu structure than I could would have done it quickly and easily, but even though I knew what it was for, turning it off and not getting billed for it turned out to be a huge challenge. Thankfully it was only $.20, but if I were using the service for something that generated actual bills, that $.20 (and possibly more) would end up quietly siphoning money out of my pocket into Amazon’s).
Starting from the login point, who asks to login to root or IAM user account in 2026?
Or having to change regions from a dropdown to see resources you own in those regions?
It's really in top 5 messy UI i have ever seen.
Remember when they decided the best UI experience was to give everything a vague abstract collection of shapes? Early 2010s or so. Couldn't tell a damn thing apart.
Consider something like a "select * from bigtable" SQL query that might process a trillion rows. Hard to know that's going to cost $100 until after you have run it.
I recently had a debate with a colleague on this topic but concerning estimating the costs of AI agent work. For example, if you prompt an AI to refactor your codebase, the final cost can't be estimated perfectly, but I'm sure it can at least be estimated with some amount of precision! Like simply knowing that it will cost < $100 is actually great information even if the final work only ends up costing $5.
I think there actually might be a business opportunity (or at least the opportunity to build something cool here) if anyone wants to work in the AI cost estimation space. It's not exactly an idea I want to pursue, but just thought I'd put it out there. AI cost estimation (even with wide confidence bands) would be very useful to a lot of people.
Green means go Orange means finish what you're doing but don't start anything new Red means stop everything
And probably a special rule to permit stable, critical spend through regardless, the same way we allow police and ambulance to run lights.
Suddenly a database, a storage service or a computer service needs to be aware of the billing situations and make behavioral decisions based on the billing status. Again, not impossible, but something that suddenly promotes billing from an async/non-crucial background service that can be paused, replayed, adjusted by account teams etc, into a crucial hot-path service.
Not to mention having no way to network isolate such service as every single service in every single network boundary needs to be able to contact such service, and every service needs (more or less) same auth permissions on all user accounts it’s running.
Again, solvable problems with enough code, but not simple problems by any means if you operate at a large scale.
If you stop then you have to decide whether to charge for uncompleted work.
Interesting tradeoffs.
For very small ops e.g. individual Lambda invocation you have similar concerns especially if lots are fired at once from a queue or schedule or fanout.
So it’s a feature the best customers don’t want, that adds risk to those customers deployments, to appease the worst customers.
At least historically. Perhaps Simon is right that the calculation has changed.
There’s no reason for the majority of them to have an infinite budget cap. Prod? Sure. UAT? Why not. The sandbox environment Johnny just spun up to test some new agentic workflow? Hell no.
Why is egress more expensive than ingress in clouds? To lock in users.
Why cloudformation in some cases leaves s3 buckets laying around after destroy? To keep charging those cents.
Why no caps? To get the user into the mindset of we ll pay whatever they say, and charge the ones that dont notice or dont ask for the refund. Same reason why my newspaper subscription autorenews.
Nothing cynical about both of these. I think assuming it is technical is naive.
“Never let me spend more than $X” also means, “Shut down my business-critical app/service/solution at 2 am on a Sunday morning because Joel in IT forgot to plan for the new report runs.”
The product design work to let customers have the first thing without risk of major pain from the second thing is non-trivial.
Sure it sucks that critical services blinked out at 2am. But what sucks even more is finding out the image on my ASG had a vulnerability that allowed someone to install a bunch of bitcoin miners which kept me fully scaled from midnight to 2am.
Or more likely, that a mistake in terraform 1000xed my spending.
Most people have predictable spending and could easily say "don't spend more than 10x what I normally spend". Or 1.5x, or 2x, 3x, etc. All depending on how they want to balance a runaway cloud expense.
There’s no magic wand that produces good outcomes when planning or execution goes awry at scale.
Just because this isn't a good solution for everyone, doesn't mean it's not a good solution for a large number of people and businesses.
A lot of businesses can tolerate outages. In fact, even very big businesses come out mostly unscathed when they have multi-hour outages. (how many is it for github this year?)
An outage causes a reputational black eye. It does not necessarily translate to lost income.
If you're a big business that wants to spend unlimited money, you apply for unlimited credit with a credit check.
If you don't, you get hard caps.
But one could just have two categories of service - the default capped plan, and a special Enterprise one where you sign a contract making it clear you understand the consequences of not having a budget limit.
Also, if your average usage is $900, set your hard limit at $2,000, not $1,000. Then when the report runs $500 over expected, you get a soft limit email and still have your report. Even a "business critical" run is probably not actually worth more than double your average spend.
The price cap should be the number you’d be willing to spend to avoid an outage vs when you’d rather kill everything and work out what happened.
How much would a hospital pay to avoid unexpected downtime of their software systems?
Price caps are for small scale stuff where you wake up on Monday and see 1000x the normal bill.
Obviously they’ve changed their mind about cost management in light of the scale and dynamism of agents, which isn’t too surprising.
My point is this was never as simple as, “Give me a dial to set my maximum account spend.”
And presumably getting bill for bajillion dollars would be bad also for them.
Sending an email when your budget gets low shouldn't be a big lift.
GCP did have a budget cap previously. I think the new one is just more fine-grained to apply to specific services.
"Sorry, you had a hard cap on AWS spend so we deleted all your S3 data on August 27th". Yeah not going to fly.
To prevent cases where users can upload an excessive amount of data on March 31 that they then can't afford the April bill for, AWS should also maintain a "next month's balance" limit that gets handled in the same way.
It would take a bit more work to correctly handle things like ephemeral data and tiered storage classes, but it's not insurmountable.
You could apply the same kind of billing to most other long-lived resources, including VMs that run business-critical services.
The horror stories I have seen are of the type: some big artifact was getting pulled in a loop, causing TBs of network traffic or access keys were leaked and malware spun up 1000 xxxlarge instances.
The ability to stop the bleeding is the bare minimum people want. Not, "Well, you made a boo-boo so now you lost everything."
grace period as AWS did or set aside x% of data spend limit on holding existing data for N days
Even as a hobbyist, even as a most careful and judicious architect and admin, I could not prevent my VPS from incurring costs beyond my control. That means that the entire Internet, anyone with some kind of material access to the VPS, and especially any user or authorized entity, they could incur costs to me without bound and without notice until slapped with a bill.
Even something as simple as egress charges aren't under your control. So if people download enough data, you pay for all of it? It seems like an absurd proposition.
It's like opening a business somewhere in a war zone, and vandals and squatters are constantly attacking your storefront, and maybe you have a band of toughs as security and some good cops to defend awhile, but you're utterly in a war zone with adversaries acting far beyond your control.
As a hobbyist, I could never again justify running a pure VPS with the Linux and stack on top, as I ran before like the MediaWiki server. I was excited to learn all the vocabulary and skillsets of cloud services, but on the "free tier" uncapped, there was no telling when I'd be presented with costs beyond my ability to handle. And I do not see how a Fortune 500 would have any different calculus in this regard.
Most businesses in recent memory were "their own landlord" of on-prem equipment and machine rooms. Yeah, they began to outsource even their IT admin, but the machinery was in-house until the cloud services took over. Did we go through a phase of collective machine rooms or data centers with a collection of tenants? It seems we skipped from "homeownership" to "feudalism" with the Cloud Providers being the Princes [beyond mere Lords] who provide minimal resources to the serfs now. I can see many corporate execs who begin to hate "AI Data Centers" just for what they have become: a very attractive and irresistable way to reduce your capex and footprint and physical plant, by "migrating to the cloud" but is that really a better status quo after all your revenue is being pumped into AWS?
And in view of what I just wrote, a "hard budget cap" checkbox is even worse for you than runaway costs, because it will allow any determined adversary, butterfingered DevOps, or innocent fuckup, to shut you down and deny your service by hitting that cap. When a business experiences a service or infrastructure outage, they count that in dollars of revenue. Your "budget cap" will cost you money because it "paused your project" and your customers all got burned in turn.
Moving toward real solutions, just spitballing, but I can envision throttling and capping of everything, every billable service, monitored by the cloud provider and ensured that your services don't spill over into unmanageable territory. If your spend could be throttled and capped by-the-minute, by-the-hour, daily, then there is a start for it. But really, any service that incurs costs to you should have reasonable rate-limiting, throttling, caps and alerts that can help you manage it. I think all that stuff is currently missing but I am not a cloud admin, obviously. Cloud services obviously have perverse incentives to open the floodgates and bill as much as possible for any possibly legitimate usage that doesn't exceed their aggregate capacities. It's like LLMs today that just burn tokens like there's no tomorrow, because someone [you, not Mexico] will eventually pay for it all.
https://cloud.google.com/blog/topics/cost-management/new-ear...
Edit: Ugh it's fake. Literally only works for four random services, unsupported for all the rest. Completely useless for all of my projects. Also dumb that the only supported term is "monthly" considering that months are different lengths, and that they don't bother to account for credits or discounts. Maybe next decade they'll get around to implementing something useful.
Edit: sadface
It works for some non-B2Bs.
Was a 6 month nightmare proving we did nothing wrong.
The whole idea was insane. We shunned paid compute services in favour of personal computing and then sold our souls back again.
Even if we have a negotiated yearly contract for $X spend per year, maybe we’ll hit that spend in 6 months instead of 12. Having some kind of automated telemetry saying how we’re trending would be so useful.
I’ve gotten vague warnings from customer success people saying vaguely that, but without any warning of what’s truly happening. The more we abstract away from money (tokens, credits, etc), the more we need a way to translate right back to money, to see how close we’re getting to any limits over time.
I don’t want to find out in month 5 that the contract which was expected to cover a year is now going to run out in 14 days.
And you could complement people's personal computers with "dedicated servers" for companies (and even individuals).
Rejection will block the payment and the company will likely go after you to settle the dispute.
But is is useful in related cases. I have now had two separate incidents where this would have saved me the hassle. Both were service providers which unilaterally tried to deviate from contract terms. One was possibly a unintended billing software upgrade where they suddenly added an extra item to a 10 year contract that did not belong there, the other was an apparently profitable "lets just try and see what percentage of customers notice" dick move. I should not have needed to care about this, both companies should have just been slapped with a hefty fee for requiring their and my bank to take a look, and then figure it out on their own instead of me having to call once to fix the problem for the future, and then call again to remind them about reimbursing the delta from past billing intervals.
It should be possible to open my bank app, see all my recurring payments (subscriptions), and be able to cancel them with one button. And this should count as an official termination.
Also, limits, etc.
One solution is banking apps that let you create extra cards and assigning hard limits to spending. Wise & Revolut apps can do that.
Have to mention OpenRouter also, I logged into OpenRouter through pi, on the auth screen it optionally let me set a limit which is very smart and user friendly.
That's not a solution because not paying your debts doesn't make your debts go away.
They may well decide not to chase customers for small unpaid amounts. But that's not something you should rely on as the issue is with large amounts.
None of this has anything to do with invoicing, but I would be surprised if they didn't issue invoices to all US customers.
I know AWS and similar sites have no such concepts, but other than those, this is an existing option of most kinds of service providers.
Realistically, most providers can't even allow you to consume $10k usage if they don't have the certainty that you can pay up. Prepaid is what gives them that certainty.
Phone service, bank card, home internet etc.
If you don't pay your bill than they just cancel your membership and it works ok.
People in western countries are just getting shafted by companies for (mostly) no reason because an alternative balance is just inconceivable.
https://devhumor.com/media/dilbert-s-team-writes-a-minivan
If you don’t have a mechanism for enforcing hard caps, you don’t get to send customers a bill for unlimited amounts.
And everywhere in Europe you have clear price sheets (unlike the deliberately opaque mess of the US price sheets with more small-print than a packet of pills), which means even if you are at an EU provider with no hard caps you can still accurately reason and predict your costs.
Just a few examples....
Cloud providers:
Inference providers: I really don't buy the stories the US providers tell you that "its too difficult" or "what if you suddenly go viral".The "viral" bit is easily solved through basic monitoring of metrics that everybody should be doing. I believe the cool-kids give it the fancy name of Site Reliability Engineering (SRE). All you need to do is top-up your balance / adjust your cap if your metrics are trending upwards for an explainable reason. Its not rocket science.
As for the "too difficult" that's just a lie. It just suits the US cloud providers better to have you spend spend spend on their messy soup of random interdependent microservices.
I realized, what am I actually afraid of. Well, overspend. So I just set then all to disable auto-reloading. Now if it blows up, I'm down $5.
Same story with containers. Just give it root on a VPS, and if it blows up, I'm down $3.
I've had an AI coding agent enter a doom loop for hours multiple times now. having a budget limit for my openrouter key is helpful, but it's already spent then.
I've been building a thing you can hook into ai workflows that uses multiple detectors if an agent is starting to loop and will then send a kill command to that chain. I don't have any testers for it though.
Hobbyist needs to not get wiped out. Business need stuff to stay online.
In order to learn how to use the platform, employees in business and startups need a relatively safe space to learn too so even they have mixed needs.
Big tech is definitely the worst offender though on handing out footguns and relying on "beg support for mercy" model.
It’s not in their interest to make cost controls work well.
Ideally you’d be able to set something granular like “allow this service to scale up only 10x, measured at an hourly level, and alert me when it happens. Drop all requests that exceed 10x”.
And then you are mostly in a throttling situation until the burst clears or a human can review & accept increased usage is ok/increase thresholds. I’d rather have services go slow during excess load (like a real server) than go dark for remainder of month.
This seems a lot better than brute force “turn everything off at $X level of monthly billing” or “no limits you can charge me infinity dollars”.
Turns out it's impossible to do. You have to delete your entire account.
Just last week I was looking for similar functionality in Cloudflare as I am exploring publicly exposing apps I have built (as opposed to just me and some friends on a home lab).
My main concern on public cloud platforms is costs. I never want to spend more than, say, 10 EUR on a simple service. In my mental risk matrix, the chance of a cost-related issue has been increasing more-and-more. I am creating and hosting more than ever, and the chances of malicious activity/abuse are in my opinion higher than ever before.
It feels like cloud providers either are unaware of this, or riding a wave of increased turnover - which I think a more likely.
I had a spending limit on for $30, so why did it keep charging? Because the spending limit is meaningless without a hidden checkbox called “enforce spending limit,” which is (or at least was for me) off by default.
To OpenAI’s credit, they refunded the money.
I get it, the problem is definitely worth solving for. Waking up with a $100k bill isn’t great.
At the same time, from a product perspective the proposed solution might be a bad idea. Simply having hard caps as default will definitely turn out to be as bad for some people as a $100k bill, see for example (1).
You can’t come up with good product changes if you don’t discuss the potential negative effects of a change.
(1) https://news.ycombinator.com/item?id=49950196
At the very least, there should be an optional hard limit that is obviously indicated in the UI. When you're signing up where you set your "usage cap" warning, next to it should be an optional hard cap with big bold red letters "THIS WILL CUT YOU OFF THE MOMENT YOU GO ONE CENT OVER". So I can set e.g. a warning at $20 and a hard cap of $100.
Not exactly the sort of case Simon has in mind, but I tell Claude to keep to hard daily limits on its OpenRouter spending for two long-running projects [1, 2].
A Routine for each project fires ever few hours, and Claude decides itself what to do in each session. It does tasks that require calls to other models through OpenRouter only when it is still within its daily budget for that project; after it reaches that cap, it does other tasks that don’t require extra spending.
[1] https://github.com/tkgally/je-dict-1
[2] https://github.com/tkgally/eex-dict
I had a personal development account for ~15 years. I tinker with infrastructure stuff and had built some centralized event reporting. One day about two years later I turned on sqs data events into cloudtrail. What I didn’t realize was that this closed a feedback loop and over the next couple of hours my run rate went to about $4k per day in cloudtrail+sqs usage.
I didn’t realize it until I hit the next months billing alarm immediately the next month. I’d racked up $25k in usage fees.
I’ll be using this feature. Nothing I run is worth that risk.
A regular AWS account doesn't have the notion of "project" and Cost Management doesn't have those spend limits.
Someone told me some years ago that it was "impossible" for Google to apply an upper usage caps due to technical design reasons which I found absurd. You can build a globally distributed continent-scale SQL database with consistency guarantees, but you can't stop me paying more than X USD when my usage ticks over that? Huh?
It smelt much more to me like Google's business model made it impossible to stop customers spending money, not their engineering.
I voted with my wallet.
If the CEO of the electric company didn't mandate circuit breakers, he should go to jail.
You can rack up an outrageous monthly electricity bill without tripping a breaker.
Google implemented caps because their competitors offered them. Before that, customers could choose one of several competitors, rent the GPUs at a fixed rate, or buy the GPUs and install them on-premises.
At no point was any law needed to solve any of this.
> Nobody wants to wake up to an email sent at midnight warning about a budget limit and find that, while they slept, their rogue service had consumed several hundred (or several thousand) more dollars of usage.
Monitoring and a circuit breaker. If you are given the tools, make your own heuristic and flip the breaker yourself. Don't let someone else turn it off, and be hostage to their process failures for getting it back on. By self-selection, if you are a hard limit customer, you are not front of their service line.
> An argument against this is that businesses don’t want their hosted applications to start throwing errors because some budget was exceeded. I expect that most businesses and individuals would prefer errors to a surprise $10,000+ bill.
Based on a career serving the Fortune 50, CTO for trillion dollar bank, etc.: "No." Assuming actual business is being done by the machinery in question...
The lights must stay on.
Cost management cannot shut off the enterprise. Cascading costs, even before reputation, are incalculable.
Seriously. You all asked for this.
I can’t think of anyone saying they would hate for AWS to support hard spending caps.
It seems like you’re saying “hey, you asked for a product, so you deserve for it to have a user-hostile feature”
Like hey, you asked for trains? Well, then you have no right to complain about any aspect of a train.
he's just projecting his coke induced thoughts
I know your post isn’t explicitly defending the platforms, but the arguments they use feel transparently flimsy.
Or if there was info that was after potentially horrible expensive operation.
That was blocking automated checking.
The entire MCP ecosystem is ludicrous. You’re paying for inference for an agent to make the same decisions over and over again, and yet the actions they’re taking can be so easily written by those same agents into a bash script you can run again and again, deterministically and for free.
MCP is the problem. Having agents “use a product on your behalf” is the problem.
The pattern you’re looking for is that agents should write scripts that use products for us, and that doesn’t need a new protocol.
But the idiocy is incredible, even allowing for this to be happen in a business is so infantile that the only hard cap that should be important is not to allow stupid people in the machine room.
Maybe because good for me but not good for the majority, or just the big corps.
Cheers!
AWS, is a loot box... Tokens are just in game currency, and that sales person is just metrics that have identified your spending as making you a whale.
Your average CTO from the last decade turned a fixed cost into variable spending that looks like a mobile game.
Could not agree more, this needs to be table stakes for any usage based service.
What on earth is Simon whittering on about?
This may happen also without vibecoding.
Article seemed clear on that?
Guy seems to have lost his mind to AI psychosis and is just posting unmitigated obvious low-grade crap of late, which gets picked up here by his fan base.
whttering / wittering: To chatter, babble, or ramble on at length about trivial matters. About right.