AI product management
Evals are a product decision
If the product manager does not define "good", the model team will, and they will optimise for whatever is easiest to measure.

When an AI feature underperforms, the first instinct is to ask engineering for a better model. More often the real problem is that nobody wrote down what "good" looks like, so there is nothing to improve against. The team tunes prompts, the demo looks a little better, and nobody can say whether the product did.
Evaluation sounds technical. The hard part is a product question: which outputs are acceptable to this user, in this workflow, at this level of risk? An engineer can build the test harness. Only someone who knows the customer can say what it should test.
Start with the decision, not the metric
Before choosing a metric I ask what the user will do with the output.
| The user will… | A mistake costs… | So measure… |
|---|---|---|
| Act on it directly | A wrong action | Precision: how often it is right when it answers |
| Review it before acting | A missed case | Coverage: how many real cases it catches |
| Use it as a first draft | Their editing time | Edit distance, tone, speed |
The same model can be excellent for one row and unusable for another. A classifier that is right 92% of the time is impressive as a drafting aid and alarming if it files documents on its own. The number means nothing until you know the row.
Three scorecards, reviewed together
I track every AI feature on three scorecards and look at them in the same meeting, every week.
- Quality. Does the output meet the bar, judged against real examples with agreed answers?
- Cost. What does each successful task cost, including retries and human review time?
- Safety. How often does it produce something we would not want a customer to see, and do we catch it before they do?
Looking at them together matters because they trade against each other. The fastest way to improve quality is often a bigger model, which hurts cost. The fastest way to improve safety is to refuse more often, which hurts quality. A team that only watches one scorecard will quietly spend the other two.
I have published the templates I use in the GenAI product metrics kit.
Build the test set from real work
The most useful evaluation sets I have worked with came from the people who do the job. The method is simple and takes about a week.
- Sit with three or four practitioners and collect fifty real cases.
- Ask them to mark the right answer for each, and to say why.
- Ask where they would disagree with each other. Those cases are your grey area.
- Add the awkward cases on purpose: incomplete input, ambiguous requests, and cases where the correct answer is "I don't know".
The explanations are worth as much as the answers. They tell you what the model must never get wrong, which is rarely the same thing as what it gets wrong most often.
Fifty cases sounds small. It is enough to tell a good change from a bad one, and a set people trust beats a large one nobody has read.
Who grades the grader
Many teams now use a second model to score the first. It is practical, and it scales. It also introduces a new component that can be wrong. Before trusting an automated judge, have people score a sample of the same outputs and compare. If the judge and your experts disagree on one case in five, the judge needs work before its numbers go in a status report.
Evals do not replace judgement
A passing score means the feature handled the cases you thought of. Production will send cases you did not. So I treat evals as a gate, not a guarantee, and pair them with two things:
- Monitoring, so drift shows up as a trend and not as a complaint.
- A flag button, so users can mark a bad output in one click.
Every flagged output is a candidate for the test set. Over a few months the set becomes a record of everything the product has learned the hard way.
Questions to ask your team this week
- Can anyone state the quality bar for our AI feature in one sentence?
- Which row of the table above describes our user?
- When did a practitioner last review the test set?
- What happens to an output a user flags?
Owning that loop is part of the product manager's job. Most of the product quality comes from it.
Found this useful? Share it, or get the next one through The AI Product Playbook, my LinkedIn newsletter.


