"How much will this cost per month" is the first question anyone asks, and it almost never gets an honest answer. Price lists quote a rate per million tokens, and nobody knows in advance how many tokens you will use. Not even the person selling it to you.
We pay for models every day and track the cost per answer, because our own margin depends on it. What follows is a calculation on a specific scenario, not a range.
The headline first: for a small company, model costs come to less than people expect. You still need to do the arithmetic, but for a different reason than you might think.
The scenario
A coffee shop in Belgrade. Two locations, eight staff. An assistant answers staff questions: how many grams in a double, what to do with expired milk, how to process a refund, where the gloves are kept.
This is the most common first use case, and it is convenient for costing: questions are short, answers come from uploaded documents, volume is predictable.
| What we count | Value |
|---|---|
| Questions per month | 300 |
| Question length | 20-40 words |
| Answer length | 80-150 words |
| Documents in the base | 40 pages of recipes and procedures |
| What goes to the model | the question + 2-3 retrieved fragments |
Here is the detail people forget: the question is not the only thing sent to the model. Fragments of your documents travel with it, otherwise there is nothing to answer from. Those fragments make up most of the volume, and they are what determines the bill.
In our scenario the question is around 40 tokens and the retrieved fragments about 900. In other words, you are mostly paying not for your employee's question but for the slice of your own knowledge base that rides along with it.
Three ways to spend money on the same 300 questions
One: the most capable model everywhere
This is what almost everyone does at the start. Pick the model you have heard of, point every task at it.
One answer costs roughly 0.55 cents, so 300 questions come to $1.66 a month. The answers are excellent. The only question is whether you need excellence when someone is asking how many grams of coffee go into a double espresso.
Two: the cheapest model everywhere
The opposite extreme. An answer costs 0.025 cents and the month costs seven cents. On simple questions the result is indistinguishable from the expensive model.
The difference shows up on the hard ones. When a barista asks about returning expired stock, and the procedure covers it in two places with caveats, a cheap model confidently answers with half the truth. The manager sorts out the aftermath, and one hour of their time costs more than a year of the savings.
Three: the model that fits the task
Simple questions go to the cheap model, hard ones to the capable model. The split is made by the service, not by a person: based on the type of question, the length of the retrieved context, and whether several sources need to be reconciled.
In our scenario about forty questions out of three hundred turn out to be hard. That comes to 29 cents a month, with answer quality indistinguishable from option one.
| Approach | 300 questions | 3,000 questions | 30,000 questions | Quality on hard ones |
|---|---|---|---|---|
| Always capable | $1.66 | $16.56 | $165.60 | excellent |
| Always cheap | $0.07 | $0.75 | $7.47 | lets you down |
| Fit to the task | $0.29 | $2.86 | $28.55 | excellent |
The interesting column is not the first one. At 300 questions the gap between approaches is a dollar and a half, and nobody would notice. At 30,000 it is $137 a month for the same quality.
The saving holds at 83 percent regardless of volume. It is simply invisible when you are small and pays for a part-time employee when you are not.
Where the overspending actually comes from
Since the model bill is small to begin with, it is worth understanding what makes it grow several times over. There are two causes, and neither is the price of the model.
First: a capable model answering a simple question. Forty questions out of three hundred are hard. If all three hundred go to the capable model, you are paying its premium seven times out of eight for nothing.
Second, and less obvious: the size of your document base. The more that gets retrieved per question, the longer the request. A tidy base of forty pages produces 900 tokens of context and $1.66 per 300 questions. A dumping ground where everything was uploaded "just in case" produces 2,700 tokens and $3.28 for the very same 300 questions.
Twice the cost for the same number of questions and, worse, with weaker answers: among four hundred pages of noise the right paragraph is found less often than among forty pages of substance.
What this means in practice
Four conclusions worth the arithmetic.
The barrier to entry is lower than it looks. For a small company, models cost single-digit dollars a month. If you were postponing this over cost, you were postponing it for no reason.
A cheap model does not save money, it defers the cost. The seventy cents you save come back as an hour of a manager untangling the consequences of a wrong answer.
An expensive model does not guarantee quality, it guarantees a bill. On simple questions there is no difference in the answer, and you pay the premium every time regardless.
Tidying your documents saves more than choosing a model. Cleaning the base is a day of work once; paying for excess context happens on every single request.
How this works with us
The service picks the model for the task, not you. Nothing to configure, the split works by default.
Plans include a clear number of answers rather than abstract tokens. You can see what is left, and there is never a bill you have no way to interpret.
You can start free, no card: a hundred answers is enough to run the scenario against your own documents and see the real cost rather than an estimate from an article. Including this one.