What It Costs to Run an AI Feature in Production
By Usama Arif, CTO at Prograsec ·
The cost to run an AI feature is the provider's token rates applied to one finished job, multiplied by the attempts that job takes, then by the jobs in a month. Add the minutes a person spends checking answers. On the GPT-6.1 Sol rates in our 1 October note, one call with 20,000 uncached input tokens and 5,000 output tokens is $0.09. The same token counts on GPT-6 Astra are $0.45. Those two figures are list-price arithmetic. They become a budget only after you add cache, retries, and review.
Usama Arif, CTO, wrote this from the check we use on a model release: one real task, in the harness we would ship, with cost counted per finished job. The dollars below use published list prices. The 1,000 jobs, the 1.2 attempts, and the $40 an hour are example assumptions for a document-question feature. A client invoice, and a Prograsec fee to build the feature, would each be written from an agreed scope.
Price one finished job
Rates are dollars per million tokens. Model bill = jobs per month × attempts × calls per job × (uncached input tokens × input rate + cached input tokens × cache-read rate + output tokens × output rate) ÷ 1,000,000.
Add cache-write tokens × cache-write rate ÷ 1,000,000 once for each prefix you store. If the prefix changes on every job, multiply that write by the number of jobs. Count a job when it passes your success check, and count every model call inside it, including tool steps and retries.
The $0.09 Sol call is 20,000 ÷ 1,000,000 × $2, plus 5,000 ÷ 1,000,000 × $10. Astra is 20,000 ÷ 1,000,000 × $10, plus 5,000 ÷ 1,000,000 × $50. If that Sol call can wait, Batch and Flex are half of standard, so the same tokens are $0.045. Fast mode is double, $0.18, and EU data residency turns Fast mode off. Regional processing, Microsoft Foundry zones, and Bedrock inference profiles add their own percentages. Use the rate card of the endpoint that will send the invoice. Recheck the OpenAI pricing page before you quote anyone.
Rates on this sheet
These are the standard list prices already cited in our model notes, read on the dates those notes name. A hosted gateway, a region premium, or a long prompt can change the invoice. Recheck the provider page before you send a number to a client or a finance lead.
Scroll the table sideways to read all columns.
| Model | Input / 1M | Cached read / 1M | Output / 1M | Whole-request reprice |
|---|---|---|---|---|
| GPT-6.1 Sol | $2 | $0.10 | $10 | Above 272,000 input tokens: $4 / $0.20 / $15. Cache writes are $2.50, or $5 above that line. |
| GPT-6 Astra | $10 | $1 | $50 | Same 272,000-token rule. Cache writes on the standard card are $12.50. |
| Claude Opus 5.5 | $4 | $0.20 | $20 | The cache-read rate is the figure in our 22 September Sol and Luna note. |
| Grok 4.7 | $2 | $0.50 | $6 | On the API, a prompt at or above 200,000 tokens bills the whole request at $4 / $1 / $12. |
Claude Sonnet 5.5 lists at the same $2 and $10 as Sol. DataCamp's comparison, cited in the Sol note, puts Sonnet cache reads at $0.20 and reads Sonnet as keeping standard rates across its million-token window. Confirm that on Anthropic's pricing page before it goes into a quote. Gemini 4 Argon, announced 30 September 2026, has an introductory card of $2 and $10, then $4 and $20, with cached input at 95% off the input rate. Google is rolling it out to trusted cyber defenders first. Leave it off a budget until you can call the API you will actually be invoiced on.
A month of one document question
This month is an illustration, built so the arithmetic can be checked. It is one document question: 20,000 tokens of instructions and source text, and a 5,000-token answer. One model call per job. 1,000 finished jobs. On the cached row, all 20,000 input tokens are cache reads, the prefix is written once in the month, and every request stays under 272,000 input tokens. Replace any of those assumptions and the total moves.
Scroll the table sideways to read all columns.
| Line | Amount | Arithmetic |
|---|---|---|
| One uncached call | $0.09 | 20,000 × $2 + 5,000 × $10, per million tokens |
| Same call, prefix cached | $0.052 | 20,000 × $0.10 + 5,000 × $10, per million tokens |
| 1,000 cached jobs, one attempt | $52.00 | 1,000 × $0.052 |
| One cache write of that prefix | $0.05 | 20,000 × $2.50, per million tokens |
| Month, warm cache, one attempt | $52.05 | $52.00 plus the single write |
| Same month, every call uncached | $90.00 | 1,000 × $0.09. The cache saves $37.95 on Sol. |
| Same month, 1.2 attempts, cached | $62.45 | 1,200 × $0.052, plus the $0.05 write. One in five jobs is run twice. |
| Same 1,000 calls on Astra, uncached | $450.00 | 1,000 × $0.45. This is the model choice, with no cache. |
On this shape, choosing Sol over Astra at full input price changes the month by $360. Warming the Sol cache changes it by $37.95. A retry rate of 1.2 adds $10.40 to the cached month. Output is a large part of the Sol call: 5,000 tokens cost $0.05 of the $0.09. Run the formula on your own token counts before you treat this table as the decision.
The 272,000-token reprice
Once Sol input passes 272,000 tokens, every token in that request uses the long-context rates, including tokens under the line. On a call with 300,000 input tokens and 2,000 output tokens, input is 300,000 ÷ 1,000,000 × $4 = $1.20, and output is 2,000 ÷ 1,000,000 × $15 = $0.03. The call is $1.23. Short-context rates would have priced that input at $0.60, so the rule adds $0.60 on this call and moves output from $10 to $15 per million. Thirty of those calls cost $36.90 before retries. Grok 4.7 applies the same kind of rule at 200,000 prompt tokens on the API. Compact the prompt, or quote the long-context card, before you promise anyone the short rate.
Review time and the build
The model line is one bill. A person checking answers is another. In the example, a reviewer spends 2 minutes on 10% of the 1,000 jobs. That is 200 minutes. At an internal rate of $40 an hour, which you replace with your own cost, the review line is 200 ÷ 60 × $40 = $133.33. On the warm-cache, one-attempt month, review costs more than the model.
The engineering quote is the third bill. It covers the evaluation set, the workflow, failure behavior, and handover. The hiring guide lists the record to ask a team for before you treat a demo as production. We built a 12-step workplace investigation that keeps a person on the consequential steps. Those minutes belong on the reviewer line. Building the workflow is the engineering quote. LangChain versus LangGraph is the decision about how much of that workflow the model is allowed to choose.
If a wrong answer can change an account, send a customer email, or file a record, price the reviewer before you hunt for a cheaper model.
Copy this worksheet
Copy the lines into the message when you ask for a scope. The example column is the month above, at 1.2 attempts. Swap in your task and your rates.
Scroll the table sideways to read all columns.
| Line | What to write | Example above |
|---|---|---|
| Task | One sentence a stranger can grade. | Answer a document question from the supplied passage. |
| Success check | What must be true of a finished job. | The cited passage supports the answer. |
| Calls per finished job | Model requests, including tool steps. | 1 |
| Uncached input tokens | Tokens that miss cache, per call. | 0 when the prefix hits; 20,000 when it misses. |
| Cached input tokens | Tokens read from cache, per call. | 20,000 |
| Output tokens | Per call. | 5,000 |
| Attempts | Full repeats per finished job. | 1.2, because one in five jobs is run twice. |
| Jobs per month | Finished jobs. | 1,000 |
| Rates | The card of the endpoint you will be invoiced on. | GPT-6.1 Sol standard, under 272,000 input tokens. |
| Model bill | The formula in the first section. | $62.45 |
| Review | Minutes per reviewed job, the share of jobs a person sees, and your hourly cost. | 2 minutes, 10% of jobs, $40/hour → $133.33 |
Send the worksheet with a scope request
Paste the worksheet into the message on the scope form. Add:
- The task in one sentence, and the check that decides a job succeeded.
- One redacted input, and the answer you would accept for it.
- The worksheet, with your reviewer rate written on it.
- Which actions a person must approve, and what the system must refuse.
- Whether the prompt can cross the provider's long-context line.
Request a project scope and send that packet. The fit review is free. We agree the next step and a written price before paid work starts. If the open question is whether the task is feasible on your data, the first paid piece is a separate evaluation, with its own deliverable and a stopping point. Ongoing monitoring and model changes are quoted on their own, the same way the AI services page separates them from the build.
Keep reading
Want this priced on your task?
Send the task, the success check, and a redacted sample. The fit review is free. We agree the next step and a written price before paid work starts.
Request a project scope