← Articles
Daily columnWritten fresh this morning

The Model Demo Is Not the Product Review

The Model Demo Is Not the Product Review

22 August 2026

Today's argument

As AI performance shifts from the model to the surrounding system, product reviews must evaluate the whole operating loop rather than celebrate isolated output quality.

A convincing AI demo can now be produced long before a dependable AI product exists. That gap is becoming one of the most expensive sources of confusion in product organizations.

The current discussion points in several directions at once. Teams are questioning whether AI is actually saving time. Others are moving complexity towards the front end as agents take over more of the middle. At the same time, attention is shifting from the underlying model to the harness around it. Add legal challenges over product accuracy and investigations into autonomous systems, and the pattern becomes clear: the output is only a small part of the product.

Yet many product reviews still begin with someone typing a prompt and end when the system produces an impressive answer. This is useful for showing potential. It is almost useless for deciding whether the product is ready.

I would not review a checkout flow by showing only the payment confirmation screen. I would want to see pricing, validation, failed payments, refunds, customer support and reconciliation. AI deserves the same treatment. The generated answer is the confirmation screen. The product is the complete operating loop around it.

That loop includes how context is selected, which tools are available, what permissions apply, how uncertainty appears, where users can intervene, what happens after an error and how the team learns from production behaviour. A stronger model may improve one step while leaving the overall experience unchanged. It may even make failures harder to notice because the response sounds more confident.

This changes how I think teams should run product reviews. Instead of asking for the best output, I want to see a representative task from beginning to end. Show me where the input came from. Show me what the system did not know. Show me the tool calls, the handoff, the exception and the recovery. Then show me the work created for the user or the operations team afterwards.

Consider an AI feature that drafts customer replies. The demo may show that it writes a good response in seconds. The product review should ask whether the right account history was included, whether restricted data was excluded, whether the agent can identify a false claim, how much editing the employee performs and whether a bad reply can be traced and corrected. If the employee spends the saved time checking hidden assumptions, the feature has moved work rather than removed it.

The same applies to planning tools. A generated roadmap can look coherent while combining stale inputs, unresolved dependencies and conflicting commercial commitments. The quality of the prose is not the quality of the plan. What matters is whether the tool exposes the assumptions that require a decision and keeps the resulting commitments connected to delivery.

This is also why broad labels such as model accuracy or answer quality are too weak for a product review. They average away the moments that determine trust. A team needs to know which scenarios are safe, which are merely convenient and which create downstream clean-up. One average score cannot tell you that.

I would structure the review around a small set of recurring scenarios: the normal case, the ambiguous case, the missing-data case, the incorrect-action case and the recovery case. Not as a test theatre with carefully selected prompts, but using realistic inputs and the actual interfaces, permissions and integrations intended for release.

This approach has an organizational benefit too. It makes responsibility visible. Model teams cannot treat interaction problems as design issues. Product teams cannot treat poor context as an engineering detail. Operations cannot be introduced after launch to absorb exceptions nobody designed for. Reviewing the loop forces those dependencies into the same room.

The result may be less impressive than a polished demo. That is exactly the point. Product reviews are not there to create confidence. They are there to find out whether confidence is justified.

As models improve, isolated output quality will become easier to obtain and less useful as evidence. The teams that build dependable products will be the ones that stop reviewing the magic moment and start reviewing everything around it.

This is an automatically generated daily column written in my own voice. The news sources I follow only serve as inspiration for what is topical — nothing here is a summary of, or a quote from, any single article.

Inspired by what was in the air at: lennysnewsletter.com, techcrunch.com, tpgblog.com, mindtheproduct.com, romanpichler.com

  • ai-products
  • product-reviews
  • quality-assurance
  • product-operations