dooopSoftware · Strategy · 13 min
How to Decide if an AI Initiative Should Continue
Evaluate real value, operation, and learning to decide whether an AI project should be expanded, corrected, or ended without rewarding enthusiasm.
Published on September 6, 2026
CENTRAL THESIS
A pilot that impresses can still consume capacity without becoming a product.
Continuity requires evidence of value, operation, and learning to expand, correct, or end.
An artificial intelligence pilot can perform well in a demonstration and still not deserve continuation. To decide whether it is time to continue or stop an AI project, leadership needs to separate three types of evidence: value observed in real use, the ability to operate with quality, and learning accumulated for the next decision.
The decision should not reward enthusiasm. It should classify the initiative as expand, correct, or end, with explicit criteria.
The right question is not whether the AI worked, but whether it deserves to continue
The first trap in evaluating an AI initiative is treating technical functioning as a sufficient signal for continuity. If the model responded, if the interface impressed, and if the demonstration reduced a task in a controlled environment, it seems natural to ask for more investment.
But a technical proof answers a limited question: is it possible to do something with artificial intelligence in this context? It does not answer, by itself, whether someone started to decide better, whether the operation can sustain the solution, whether the risk is acceptable, or whether the next investment cycle is clearer.
This distinction matters because AI initiatives often mix three different evaluations:
- technical feasibility: can the solution perform the proposed task under known conditions?
- value for the user: does the solution change behavior, decision, time, quality, or effort in a relevant task?
- operational capability: can the company keep the solution running with monitoring, support, review, and governance proportional to the impact?
When these three questions become one, the meeting tends to become hostage to the demonstration. Those who saw the AI succeed remember the potential. Those who know the operation remember the exceptions. Those who pay the bill want a decision. The problem is that all can be right and yet the initiative is not ready to scale.
The DORA 2025 report presents AI as an amplifier of existing organizational strengths and weaknesses and highlights the importance of the organizational system for return on investment. This is a good lens for the decision: AI should not be evaluated only as a tool but as a capability embedded in process, people, data, leadership, and operation.
If the organization cannot explain which behavior should change, continuity becomes a bet. If it can explain but cannot yet measure, the likely decision is to correct. If it can measure, operate, and learn, expanding may make sense, provided it is not confused with automatic scaling.
Decide if the next step is to expand, correct, or end
A continuity meeting needs to start with the type of decision possible. Without this, the evaluation becomes a dispute between optimism and caution.
There are three useful decisions:
- expand: increase the scope, audience, frequency of use, or integration of the initiative into the product or process. This decision requires consistent signals of value, minimum repeatability, and operational control. It is not enough that there is a case where AI helped. There must be evidence that it helps in a recognizable situation, with identifiable users, under conditions the company can monitor.
- correct: preserve the hypothesis but change design, data, flow, metric, governance, or usage method. This is the appropriate decision when the problem seems relevant but execution is still fragile. Correcting is not postponing out of fear. It is admitting that the initiative still reduces a relevant uncertainty, provided the next cycle has a more precise question.
- end: stop the initiative as a product, feature, or active investment front. This decision is legitimate when evidence of value is low, the cost of operation is disproportionate, the risk is incompatible with the demonstrated benefit, or the remaining learning does not justify further effort.
Ending an AI pilot is not declaring that the technology failed. It can be a mature allocation decision. A worse choice than ending early is maintaining by inertia an initiative that consumes product, engineering, and operation without improving the next decision.
This logic connects to previous work mapping opportunities and prioritizing initiatives. An AI roadmap remains useful only if each experiment has a review point, not just a growing list of possibilities.
Evaluate observed value, not declared intention
The central question is not whether someone liked the AI. It is whether the initiative changed a real task.
Observed value can appear in different forms, as long as the evidence is close to the work the AI intends to improve. Some possible signals:
- recurring adoption by users who have a clear task;
- reduction of rework in a specific step;
- improvement in the quality of an operational decision;
- reduction of time in a defined task;
- increase in perceived consistency in a delivery;
- reduction of doubts or escalations in a known flow.
The decisive word here is "observed." Declared intention helps formulate hypotheses but does not sustain continuity alone. A user may say they would use an intelligent feature and ignore it when it enters the real flow. They may also praise the demonstration and continue preferring the old process because the recommendation arrives late, requires too much verification, or does not fit the decision moment.
Fictional example: a software house tests an AI that summarizes support tickets before service. In the demonstration, the summaries seem clear, organized, and useful. Still, continuity should only be approved if there is evidence that analysts use the summary to serve with less effort, decide better on routing, or reduce avoidable reopenings. These effects should not be presumed. They must be measured in the flow.
If analysts read the summary but continue opening the full history because they do not trust the information selection, the likely decision is to correct. Perhaps it is necessary to indicate the origin of excerpts, separate facts from inferences, or adjust the summary to the ticket type. If analysts ignore the recommendation and the reason is not adjustable, ending may be more responsible than insisting.
This criterion also avoids a common confusion: mistaking a pleasant interface for a capability incorporated into the product. A good experience increases the chance of use, but continuity depends on what changes in operation. This same care appears in discussions of AI maturity: it is not enough to have a tool; it is necessary to know which decision it improves and how the organization sustains that use.
Separate promised productivity from measured productivity
Productivity is one of the most dangerous words in AI projects. It seems objective but is often measured by perception, enthusiasm, or fragile comparison.
The February 2026 update of METR considers new data an unreliable signal of AI's current effect on productivity. The organization points out difficulties such as selecting participants and tasks and challenges measuring time when agents act in parallel. This does not prove that AI does not generate gains. It also does not authorize concluding that it always does. The applicable point for the decision is different: measuring productivity requires careful design.
To evaluate continuity, prefer metrics close to the task. Instead of asking if "the team became more productive," ask what should have changed:
- did the time to classify a request decrease under comparable conditions?
- did the amount of human review needed fall without worsening quality?
- did the analyst make the decision with fewer back-and-forths?
- did service become more consistent among different people?
- did the recommendation reduce rework or just shift effort to verification?
It is also necessary to define an observation window. Measuring after few assisted uses usually captures novelty, not capability. Measuring too late, without a baseline, mixes AI with several other process changes. Compare a defined task, with known quality criteria and enough users to reveal usage variation.
It is not necessary to turn every pilot into academic research. But it is necessary to avoid retrospective estimation without method. "Seems faster" may justify investigation. It should not justify expansion.
The experimentation platform Microsoft ExP is described by Microsoft as a way to incorporate experimentation into the development cycle, validate hypotheses, measure impact, and iterate products. To evaluate an AI initiative, we propose the following criterion: if the company wants to decide with less noise, it needs to design measurement as part of the product, not as a later report.
Measure the cost of keeping intelligence running
An initiative can generate one-time value and still not be ready for scale. This is an uncomfortable but necessary point.
The cost of AI is not only in tool, model, or infrastructure consumption. It appears in activities often invisible in the demonstration:
- curation and updating of data;
- human review of responses, recommendations, or classifications;
- quality monitoring over time;
- handling exceptions and ambiguous cases;
- adjustment of prompts, rules, models, or flows;
- support for internal users or customers;
- security, privacy, and access controls;
- sufficient explainability for the usage context;
- recording incidents, doubts, and exception decisions.
There is no mandatory maturity in building your own model. Mature products can integrate third-party models, provided the company knows which responsibilities remain with it: experience quality, domain suitability, monitoring, support, risk, and decision on when AI should not act alone.
This is the point where many leaders discover that continuity is not just an engineering question. It involves product, operation, service, security, leadership, and, in some cases, customers. If the initiative depends on constant review but no one has time, authority, or criteria to review, the solution has not yet become a capability. It has become a dependency.
The decision to expand should require known cost, even if still estimated within an internal range and accompanied by uncertainties. What does not work is approving scale without knowing who monitors, who corrects, who is responsible for exceptions, and what signal stops use.
This reasoning also connects to the business-connected artificial intelligence strategy. Strategy here means deciding which capabilities the organization accepts to build, maintain, and review.
Turn learning into evidence for continuity
Not every experiment needs to become a product. Every experiment, however, should produce a better decision.
Useful learning is not a collection of loose impressions. It needs to reduce uncertainty. Before approving new investment, it is worth recording what the initiative taught in at least six dimensions:
- tested hypothesis: what assumption was at stake?
- signal found: what appeared in real use?
- risk discovered: which error, dependency, bias, or exception became clearer?
- technical limit: what does the solution still not do well?
- operational limit: which process, data, or responsibility needs to change?
- next decision: what can now be decided with more confidence?
If learning does not change the next decision, it probably does not justify more investment. This phrase is harsh but protects the organization from a common pattern: continuing because "there is still much to learn," without saying which uncertainty will be reduced.
Useful learning eliminates bad paths, strengthens good hypotheses, and makes explicit what is still unknown.
In an AI initiative, this record also helps separate correction from persistence. Correcting makes sense when there is a relevant hypothesis and a diagnosable failure. Persisting without diagnosis only prolongs the cost of doubt.
Use a checklist to decide: expand, correct, or end
AI project continuity becomes more objective when leadership uses a short, commented checklist. It does not replace judgment. It forces the right conversation.
Problem and user
Question: does the initiative solve a specific task for an identifiable user?
Expand when the user, task, and usage moment are clear and recurring. Correct when the problem seems relevant but the user, flow, or usage situation is still poorly defined. End when the initiative depends on a generic promise of innovation without a concrete operational task.
Observed value
Question: is there evidence of real change in behavior, quality, time, or decision?
Expand when improvement appears in real use and repeats in more than one observation cycle. Correct when there is a positive signal but it depends on a restricted group, assisted use, or an exceptional condition. End when the demonstration pleases but does not change execution or anyone’s decision.
Measurement
Question: does the metric used represent the task the AI should improve?
Expand when there is a metric close to operation, with baseline and minimally controlled comparison. Correct when the metric exists but still measures perception, activity, or volume, not outcome. End when the initiative is defended only by opinion, enthusiasm, or estimation without method.
Operation
Question: can the company keep the solution running with acceptable quality?
Expand when there are responsible parties, monitoring routines, exception handling, and known cost. Correct when operation is possible but depends on adjustments in data, process, governance, or support. End when the effort to operate exceeds observed value or requires a capability the company does not intend to build.
Risk
Question: are risks of error, exposure, bias, dependency, or inadequate decision addressed?
Expand when relevant risks have controls proportional to use and decision impact. Correct when risks are known but controls still need to be designed or tested. End when risk is high, poorly understood, or incompatible with demonstrated benefit.
Learning
Question: did the initiative produce learning that improves the next decision?
Expand when learning confirms a capability that can be repeated in similar contexts. Correct when learning points to clear changes for a new test cycle. End when the initiative no longer reduces relevant uncertainty and only consumes effort.
Fictional example: a software house tests an AI that classifies support requests by urgency. Initial accuracy seems good in the demonstration, but continuity should only be approved if classification can reduce reclassifications, speed analysts’ decisions, or improve routing with acceptable monitoring cost. These are effects to measure, not presumed results.
If errors concentrate on critical cases, the likely decision is to correct before expanding, with review of classification criteria, exception handling, and more explicit human supervision. If analysts ignore the recommendation and the reason is not adjustable in the flow, ending may be the most responsible decision.
The concrete decision is to record the initiative as expand, correct, or end, along with the evidence supporting the choice, the person responsible for the next action, and the review point.
The closing should classify the initiative as expand, correct, or end, with the evidence supporting that choice.
If you want to discuss this decision in the context of your company, talk to dooop.
Further reading
- AI in software companies: strategy, delivery, and differentiation
- How to integrate domain knowledge into product strategy
- How to organize an AI experiment portfolio
Sources
NEXT DECISION
Discuss application in your company
Conversation about the software company context
Content by dooop. Registration allows linking this topic to the reader’s journey and tracking interest in the subject.
