← All articles

AI strategy and economics

The cost of a right answer

What cloud cost governance taught me about pricing an AI feature, and why cost per request is the wrong number.

Before I worked on AI products, I worked on cloud cost. At Microsoft I led a product that helped enterprises estimate, budget and govern what a move to the cloud would cost them. The lesson I took from it applies almost unchanged to AI: teams do not overspend because they are careless. They overspend because nobody can see the cost at the moment a decision is made.

AI features have the same problem, with one extra twist. The unit you are charged for is not the unit your user cares about.

Cost per request is the wrong number

Model pricing is quoted per token. Dashboards show cost per request. Neither tells you what the product costs to run, because a request is not a result.

A single user task can involve several requests: a first attempt, a retry when the output fails a check, a second model reviewing the first, and sometimes a person correcting the answer. The number that matters is the cost of one task done correctly.

What you count What it misses
Cost per token How long your prompts and outputs actually are
Cost per request Retries, review calls and failed attempts
Cost per successful task Nothing important

A simple way to estimate it: take total spend for a week, add the cost of human review time, and divide by the number of tasks that ended in an accepted result. The figure is often well above what the per-request number suggested.

Where the money goes

When I break down the cost of an AI feature, the same four things show up.

  1. Context. Long prompts and large retrieved documents are sent on every call. Trimming what the model reads is often the cheapest saving available.
  2. Model choice. The largest model is rarely needed for every step.
  3. Retries. Every failed attempt is paid for. Poor quality is a cost problem as well as a user problem.
  4. Review. A person checking outputs is part of the cost of the feature, even though it never appears on the model invoice.

Choosing a model is a product decision

It is tempting to pick the most capable model and move on. That works for a prototype. For a product, the better question is: what is the smallest model that meets the quality bar for this step?

Most workflows have steps of different difficulty. Classifying a request is easier than drafting a response to it. Routing easy steps to a smaller, faster model and reserving the large one for hard steps can cut cost substantially while improving speed.

This only works if you have a quality bar and a test set, which is one more reason evaluation is a product decision. Without them, you cannot tell whether the cheaper model is good enough, so you keep paying for the expensive one out of caution.

Make cost visible where decisions happen

On the cloud cost product, forecast accuracy improved by 25% and decision time fell by 20%. The product worked by putting clear, comparable numbers in front of people at the moment they were choosing between options.

The same principle applies inside an AI team.

  • Show cost next to quality. A prompt change that improves quality by a point and doubles cost should be visible as exactly that.
  • Set a budget per task. "This task should cost under X" is a requirement engineers can design to.
  • Alert on trend, not total. A feature whose cost per task rises 10% a week is a problem long before the invoice looks alarming.

Price it before you ship it

If the feature is sold to customers, work out the unit economics before launch. Ask what a heavy user costs you in a month, and whether the plan they are on covers it. AI features have a real marginal cost per use, which most software does not. A pricing model that ignores this can turn your most engaged customers into your least profitable ones.

Questions to ask your team this week

  • What does one successful task cost us, including retries and review?
  • Which step in the workflow uses a larger model than it needs?
  • Can an engineer see the cost impact of a prompt change before merging it?
  • What does our heaviest user cost us each month?

Cost is not a reason to avoid building with AI. It is a design constraint, like speed or accuracy, and the teams that treat it that way early have far fewer surprises later.

Found this useful? Share it, or get the next one through The AI Product Playbook, my LinkedIn newsletter.

The AI Product Playbook

Get new articles by email

Leave your email and I'll send you each new article on AI product management as it is published. You can also follow the newsletter on LinkedIn.