
2026-09-16Chandu
What AI Actually Costs To Run, And Why Unlimited Is Not A Plan
You are comparing two AI assistants for your website. One charges a flat monthly fee and promises unlimited messages. The other gives you an allowance and tells you what happens when it runs out.
The first one looks like the better deal. It is the one to be more careful about.
The Price Is Flat. The Cost Is Not.
Every answer an AI assistant gives costs its vendor money, paid to whoever runs the model, and the amount depends on how much text goes in and comes out. Software used to be different. Once a feature was built, the thousandth customer using it cost roughly nothing. AI work is not like that. Each question is a small purchase made on your behalf, and the vendor pays for it whether or not you do.
So a flat price is not a statement about cost. It is a bet. The vendor is betting that most customers will use far less than they pay for, and that the ones who use more will be quietly absorbed by the rest.
Not All AI Work Costs The Same
It gets less flat the closer you look. The unit everything is measured in is a token, roughly a short word, and a token does not have one price.
When we built the meter for Kakapo, we priced every model against the small one most answers come from. A token from the larger model costs about seventeen times as much. A token spent training the assistant on your own content, which is done with a much cheaper kind of model, costs about a fourteenth as much.
That is a spread of more than two hundred to one between the cheapest work and the dearest, all of it called "usage".
Counting Messages Punishes The Wrong Customer
The tempting fix is to count something simpler. Messages, requests, or raw tokens, with one price for all of them.
Counting raw tokens is worse than it sounds, because of what a new customer does first. Before an assistant can answer anything, it has to read your website, your documents, and your FAQ. That is mostly training work, the cheap kind. Under a flat count it is also by far the largest amount of work you will ever do in one go. So the first useful thing a new customer does is also the fastest way to burn through their allowance, for work that cost the vendor almost nothing.
Then it swings the other way. A customer who switches their assistant to the larger model spends the same allowance on work that costs seventeen times more. The vendor is now losing money on exactly the customers who value the product most.
Unlimited Means One Of Two Things
Put those two facts together and "unlimited" has to be hiding something.
Either there is a cap you have not been shown. A fair use clause, a rate limit, a model quietly downgraded when you get expensive. You will find it the week you need the product most.
Or there is no cap, and the price has been set for the worst case. Everybody pays for the heaviest user, which is a strange thing to sell to a small business that will never be one.
Neither is dishonest, exactly. But neither is a plan you can budget against, because you do not know which one you bought.
What We Did Instead
Every plan for Backoffice includes an AI allowance, and the allowance is measured in one unit: the equivalent of a token from the small default model. Work on the large model draws about seventeen units per token. Training draws a fraction of a unit. An answer draws what it actually cost, and you can see what is left.
One detail matters more than it looks. Each piece of work is weighted at the moment it happens and written down that way. When model prices change, and they do, the weights change for work from that day on. What you already spent stays spent at the rate it was spent at. A month that was inside its allowance cannot fall outside it later because somebody updated a price list.
The pricing page turns the allowance into a rough number of answers, and it says "about". A bot answering from the larger model gets fewer. Stating a flat number would be a promise the meter does not keep.
When It Runs Out, It Stops
An allowance is only honest if the edge is honest too.
When an account has used its allowance for the period, the assistant stops answering. Chat, training, and the AI reply drafts in the inbox each check the meter before they spend, and each is refused when it is empty. Nothing is charged past the limit. There is no overage bill arriving at the end of the month.
Your website visitors are told the assistant is unavailable, and nothing else. Whose budget ran out is not their business. You are told in the product exactly what happened and what to do about it. Nothing is lost: the assistant comes back when the period resets or the plan changes.
The failure runs in your favour too. If our own meter cannot be read because of a fault on our side, the work is served and we carry the cost. A limit exists to protect you from a surprise, not to protect us from a bug of our own.
An Allowance Is Not A Paywall
There is one more way to get this wrong, and it is common: charging for AI as a separate module. Every plan unlocks every part of Backoffice, the assistant included. What a plan decides is how much AI work is included, not whether you may use it. The plans differ in the allowance, and in nothing else.
That is a quota inside something you already have, not a second door with a second price on it.
"Free Forever" Is About The Price
Our free plan is free for as long as you keep it. That is a promise about the price, and it is one the system enforces: withdrawing the offer to new sign-ups never moves anyone already on it.
It is not a promise that the allowance will never change. Allowances can change, on any plan, and the pricing page says so in plain words, on the same page as the offer. The alternative is to let people read "free forever" as "this will always work the way it does today", and then discover otherwise. A reasonable person would read it that way. So we say it before they do.
What To Ask Before You Buy One
Three questions, and they work on any vendor.
- What does a message cost you, and does it change? If they cannot say, the price you are paying was not built from cost, and the gap will close somehow.
- What happens when I hit the limit? You want to hear "it stops" or "you pay this much more", stated up front. "There is no limit" is the beginning of the answer, not the end.
- If your costs change, what happens to what I already used? A meter that can be rewritten backwards is not a meter.
A price that says unlimited is easy to sell. A price that says what it covers, what it costs, and what happens at the edge is the only kind a business can plan around.
If you are working out what AI would cost to run inside your own business, our AI workflow integration work starts with that question rather than with a demo. Tell us what you want it to do and we will tell you what it would take.