A Pilot Without a Stop Condition Is Already in Production

5 September 2026
Today's argument
Product teams should define when and how to halt a pilot before agreeing on the evidence required to expand it.
Most pilot plans explain what must happen for the product to continue. Far fewer explain what must happen for it to stop.
I think that order is wrong. Before agreeing on the evidence required to expand a pilot, a product team should define the conditions that will pause it, who has the authority to make that call and how the change will be reversed.
This matters because a pilot is not a safe category of product. It is a production release with a smaller audience. The recent reports of AI agents reaching the open internet without the organization knowing illustrate that clearly. Investigations into autonomous vehicle deployments make the same point in a physical setting. Even a feature that manages a customer’s photo library can create irreversible consequences if it takes an unwanted action.
Calling these products experimental does not reduce their impact. It only describes the organization’s uncertainty about them.
Product teams are generally much better at defining positive signals. We choose an adoption target, a conversion rate or a task-completion measure. We may add a dashboard with error rates, complaints and retention. But a collection of metrics is not a stop condition. Someone still has to decide which change matters, at what level and with what response.
Noisy dashboards make this worse. When every metric is presented as relevant, each individual warning becomes easier to debate. A support spike can be explained as launch friction. A drop in completion can be blamed on the audience. An operational incident can be labelled an edge case. The pilot continues while the team waits for cleaner evidence.
A useful stop condition is specific enough to trigger action under pressure. For an AI support assistant, that could mean pausing automated replies when the team finds a certain class of harmful answer, regardless of the overall satisfaction score. For a pricing experiment, it could mean removing the variant when existing customers are incorrectly shown new terms. For an onboarding change, it could mean restoring the previous flow when customers cannot recover from a failed step.
I would not use one generic threshold for every experiment. The condition should reflect the possible harm and how reversible the action is. A button colour test can tolerate uncertainty. An agent sending messages, changing records or deleting files cannot. The larger the blast radius and the harder the recovery, the earlier the team should stop.
The owner also needs to be named in advance. In many pilots, Product watches customer behaviour, Engineering watches system health, Support watches complaints and Legal or Security joins when something has already gone wrong. Each function sees only part of the situation. If nobody has explicit authority to pause the release, the default decision is to keep it running.
This is not an argument for central approval of every experiment. Small experiments remain one of the best ways to improve products and team processes. It is an argument for matching autonomy with a clear boundary. Teams can move faster when they do not need an executive discussion to decide whether an agreed condition has been crossed.
I also want the rollback tested, not merely mentioned. A feature flag is not a rollback plan if disabling it leaves changed customer data behind. Turning off an agent does not recover emails it already sent. Ending a commercial test does not resolve contracts created under the test. The plan must cover the state left after the product stops.
This discipline becomes more important when growth is less predictable and recurring revenue feels less secure. Under pressure, organizations tend to keep weak pilots alive because stopping feels like losing momentum. In reality, a pilot that cannot be stopped cleanly is already part of the operating model.
Before I approve a pilot, I want three answers: what makes us stop, who can stop it and what remains afterward. If those answers are unclear, the team is not testing safely. It is simply releasing uncertainty to customers.
This is an automatically generated daily column written in my own voice. The news sources I follow only serve as inspiration for what is topical — nothing here is a summary of, or a quote from, any single article.
Inspired by what was in the air at: lennysnewsletter.com, techcrunch.com, tpgblog.com, mindtheproduct.com, romanpichler.com
- product-experimentation
- risk-management
- product-operations
- ai-products