← Articles
Daily columnWritten fresh this morning

Evidence Quality Belongs in the Product Budget

Evidence Quality Belongs in the Product Budget

19 September 2026

Today's argument

Product leaders should fund the systems that make results trustworthy, rather than treating measurement as a free by-product of shipping.

Most roadmaps fund the change but not the evidence needed to judge it. We allocate engineers to a new onboarding flow, pricing model or AI feature, then assume the existing dashboard will tell us whether it worked. It usually cannot.

This gap is becoming harder to ignore. Product teams are under pressure to demonstrate ROI. At the same time, dashboards keep accumulating metrics, AI suppliers reveal little about how their models behave, and agents are being given access to calendars, inboxes, payment details and operational systems. The decisions are getting more consequential while the evidence behind them remains surprisingly weak.

I think evidence quality should be an explicit part of the product budget.

That means more than adding analytics events before release. Trustworthy evidence requires clear definitions, reliable data collection, a comparison point and an agreed decision that the result can influence. It also requires someone to check whether the system is measuring what the team thinks it is measuring.

Take a redesigned checkout. A team may see conversion improve and declare success. But perhaps payment failures rose, support contacts increased or more customers requested refunds later. The dashboard is not necessarily wrong. It is answering a narrower question than the business needs answered.

The same problem is more serious with AI products. A general benchmark or supplier evaluation cannot tell me whether a model performs reliably inside my product, with my customers, data and consequences. An assistant that drafts a plausible reply has a different risk profile from one that sends the reply, changes an order or initiates a payment. Evaluation has to reflect the actual job and the cost of being wrong.

External evaluators can help, especially when specialist knowledge or independence is required. But outsourcing evaluation does not outsource the product decision. I still need to know what was tested, which cases were excluded, how failures were classified and whether the evaluator’s incentives match mine. A reassuring score without that context is procurement material, not product evidence.

I would make evidence work visible during planning. If we are changing pricing, the scope should include revenue attribution, discount behaviour, churn signals and a defined observation period. If we are introducing an AI support assistant, the scope should include a representative test set, human review criteria, incident logging and a way to compare assisted and unassisted outcomes. If we are entering another language market, translated screens are not enough; we need a way to detect where customers misunderstand the proposition or fail to complete the task.

This work competes for capacity, which is exactly why it belongs in the budget. When it remains implicit, it gets squeezed between implementation and the release date. The team ships, leadership asks whether the change worked, and someone assembles an answer from incomplete data. That answer then drives the next roadmap discussion.

I have found that smaller experiments are useful not simply because they reduce delivery cost, but because they make evidence easier to inspect. A limited release with a specific question often teaches more than a broad launch supported by twenty metrics. The point is not to test everything. It is to know which uncertainty matters enough to resolve before increasing exposure.

Product leaders should also retire measurements. A metric that no longer changes a decision creates noise and maintenance work. Every dashboard tends to grow because adding a metric feels safer than removing one. Over time, teams can explain almost any outcome by selecting a favourable number. Fewer measures with explicit decision rules are more useful than comprehensive reporting that nobody can interpret consistently.

We would not accept a feature whose core workflow was never designed. We should stop accepting launches whose evidence system was never designed either. If a result matters enough to shape investment, staffing or customer exposure, then the quality of that result deserves real product capacity.

This is an automatically generated daily column written in my own voice. The news sources I follow only serve as inspiration for what is topical — nothing here is a summary of, or a quote from, any single article.

Inspired by what was in the air at: lennysnewsletter.com, techcrunch.com, tpgblog.com, mindtheproduct.com, romanpichler.com

  • product-measurement
  • product-investment
  • ai-evaluation
  • decision-making