Ler original em português

← All content

dooopSoftware · Strategy · 12 min

Responsibilities After Delivering AI Software

After delivery, AI software needs owners for operation, evaluation, and evolution, with clear boundaries to decide on errors and changes.

Published on September 6, 2026

CORE THESIS

After delivery, AI continues to make decisions in use. Without clear owners, every error becomes an authority dispute.

Operation, evaluation, and evolution require separate mandates. Responsibility arises before the incident.

The technical delivery is complete, the AI-powered functionality is in use, and the question changes: who is accountable for the AI’s behavior once it starts to vary, err, partially succeed, and generate operational uncertainty?

Responsibility for AI software should not be hidden in support or backlog. After delivery, three owners must be appointed: operation, evaluation, and evolution. Each decides different matters, with clear limits and minimal review rituals.

Availability Does Not Solve Responsibility

After delivering software without AI, attention usually focuses on availability, fixes, and new demands. In AI software, these remain relevant but are insufficient. The functionality can remain technically available while producing responses inadequate for the context, ambiguous interpretations, or decisions difficult to explain.

The problem is not that AI is unpredictable in every way. The problem is that it introduces a layer of probabilistic judgment dependent on data, context, instructions, actual use, and acceptance criteria. When no one is assigned to monitor this layer, the organization tends to treat every signal as a technical call, customer complaint, or improvement request.

These are different natures.

An infrastructure error demands a response. A poor AI response demands another. A discovery about user behavior may require a product decision, not a rushed fix. This distinction defines who can act, who must evaluate, and when product must decide.

The presentation of the DORA 2025 report describes AI as an amplifier of existing organizational strengths and weaknesses and highlights the importance of the organizational system for return on investment. This formulation is useful here because it shifts the discussion: the question is not only which model was used, but what decision system exists around the software.

If the process was already confusing, AI does not fix the confusion. It makes it more visible.

Separating Operation, Evaluation, and Evolution Avoids a Single Owner for Different Problems

The first decision after delivery is to separate three responsibilities.

The AI operation owner keeps the routine running. They monitor failures, unavailability, noticeable degradation, recurring calls, operational costs, and escalations. Their focus is continuity of use.

The AI evaluation owner decides if the response remains good enough for the context in which it is used. They do not only measure whether the functionality responded. They observe if it responded with acceptable quality, if errors are within agreed limits, and if there are signs that behavior has changed significantly.

The product evolution owner decides what to do with the learning. They turn findings into prioritization, adjustment, experiment, interface change, new business rule, user training, or discontinuation of a capability.

Combining all into one person may seem efficient at first. In simple cases, it might work. But when functionality begins to affect customers’ routines, internal areas, and product decisions, a single owner tends to mix operational urgency with quality judgment and evolution ambition.

This mixture creates poor decisions for good reasons. Operation wants to reduce calls. Evaluation wants to preserve quality. Product wants to learn and evolve. All are legitimate concerns. The mistake is pretending they are the same decision.

For companies still organizing their AI direction, this separation relates to a larger question: what capability does the organization want to build, not just what functionality it wants to launch. This theme appears more broadly in how to create an AI strategy connected to the business and in AI roadmap, but here the focus is the stage after delivering a functionality already in use.

The Operation Owner Ensures the System Remains Usable

The operation owner does not need to judge the AI’s final quality. But they must know when a situation is abnormal and who should be contacted.

In practice, this responsibility includes monitoring signals such as:

  • technical failures and unavailability;
  • increase in calls related to the functionality;
  • degradation perceived by users or internal teams;
  • operational costs incompatible with expected use;
  • human review queues becoming unfeasible;
  • need to limit, pause, or route flows to human support.

The critical point is defining autonomy before the incident. Can the operation owner interrupt a functionality? Reduce its scope? Block a category of response? Trigger evaluation without product prioritization?

Without these answers, operation becomes a messenger. They perceive the problem but lack authority to act. Or, at the opposite extreme, act alone on an issue that required quality judgment or product decision.

A good operational criterion is simple: every AI functionality must have a list of situations where operation can continue, limit, or interrupt use. This list does not need to predict everything. It must cover the most likely scenarios of degradation, doubt, and escalation.

Responsibility after delivery begins when someone can make a decision proportional to the problem, without improvising authority amid crisis.

The Evaluation Owner Decides if the Response Is Still Good Enough

Evaluating AI requires criteria before opinion. “I liked it” and “I didn’t like it” can start a conversation but do not sustain recurring decisions.

The AI evaluation owner must translate quality into observable criteria. In a functionality that classifies requests, for example, criteria may involve adherence to the correct category, need for human review, acceptable error types, unacceptable error types, complaint recurrence, and consistency with the client context.

The most important detail: user feedback is not synonymous with machine learning. A user clicking approve, reject, or commenting on a response can generate data for analysis. This does not mean the model was retrained, the functionality improved, or the next response will be better as a direct consequence.

The update from METR on productivity is a good reminder of methodological caution in another context: the organization itself considers new data an unreliable signal of AI’s current effect on productivity and points out measurement difficulties such as participant and task selection and timing with competing agents. The usefulness of this source here is not to prove AI’s value but to reinforce that measuring AI systems requires care before drawing broad conclusions.

The same caution applies to evaluating a functionality. If the team measures only usage volume, it may confuse adoption with quality. If it measures only complaints, it may ignore silent errors. If it measures only human review, it may miss changes in demand type.

The evaluation owner must have authority to say three things:

  • the response is acceptable to continue in the current flow;
  • the response is doubtful and needs review, restriction, or a larger sample;
  • the response is unacceptable for that use and must return to human decision or be interrupted.

This is a product and operation decision, not a matter of preference.

The Evolution Owner Turns Learning into Product Decisions

Not every finding becomes an improvement. Not every complaint becomes backlog. Not every error requires retraining. The product evolution owner exists to prevent learning from becoming an infinite queue.

After delivery, an AI functionality can point to different paths. A recurring error may require instruction adjustment, interface change, new user question, data review, additional business rule, team training, controlled experiment, or autonomy reduction.

Microsoft describes its ExP platform as a way to incorporate experimentation into the development cycle, validate hypotheses, measure impact, and iterate products. This does not mean experimenting guarantees success. It means hypotheses must be treated as hypotheses, with measurement and subsequent decision.

This point is decisive in AI software. When the team learns something after delivery, it must ask: is this an incident, operational adjustment, quality problem, or product proposal change?

Each answer leads to a different path.

If incident, priority is to restore reliable use. If operational adjustment, improving instructions, escalation, or communication may suffice. If quality, evaluation must review criteria, examples, and limits. If product, leadership must decide if the functionality should change its experience, audience, scope, or promise.

At this point, AI maturity means deciding clearly what will be maintained, limited, tested, or abandoned. Using own or third-party models does not resolve this decision alone. This reasoning connects to the diagnosis addressed in AI maturity, but the post-delivery decision must happen at the operational functionality level.

Fictional Example: Triage Assistant in a B2B Platform

Imagine a fictional example: a B2B platform uses an AI assistant to classify customer requests before routing them to the correct team. The functionality does not make a final business decision. It suggests a category and, in some cases, requests human confirmation.

In the first usage stage, operation notices an increase in reopenings of requests that passed through triage. This is an operational signal. The operation owner does not need to conclude the AI is “bad.” They need to record the pattern, check for technical failure, observe if there is concentration in any flow, and trigger evaluation.

The evaluation owner analyzes samples of reopened requests. They identify a hypothesis: an ambiguous category is confused with another when the customer request mentions two topics in the same text. This analysis alone does not prove the cause is resolved. It delimits an error type and allows discussion if the error is acceptable, doubtful, or unacceptable for that flow.

The evolution owner decides not to treat the case only as a prompt correction. They propose testing a complementary question before automatic classification when the request mentions both topics. The hypothesis is that this question reduces ambiguities without significantly increasing friction. The effect must be measured. It may work, may not, may improve triage and worsen experience. Therefore, it is a product decision, not an automatic reflex.

In this fictional example, roles remain separated:

  • operation notices the increase in reopenings and triggers the process;
  • evaluation identifies the error type and judges its severity;
  • evolution chooses which hypothesis will be tested and what product change makes sense.

The difference seems subtle but changes the conversation. No one needs to pretend every problem is a bug. No one needs to turn every complaint into a demand. No one needs to wait for a crisis to decide who can limit the AI.

Checklist to Appoint Owners After Delivering AI Software

Use this checklist as a management and product tool. It does not replace legal, regulatory, or sectoral analysis when such evaluation is necessary. The goal is to reduce diffuse responsibility after delivery.

Is there an operation owner?

The criterion is clear: the person or area knows how to monitor availability, failures, calls, costs, and escalation routines.

If the answer is no, appoint this responsible before expanding use or adding new cases. AI functionality without an operation owner depends on informal perception, which usually arrives late.

Is there an evaluation owner?

Does the responsible party have documented criteria to say if the AI response is acceptable, doubtful, or unacceptable?

If no criteria exist, define error types, reference examples, and review frequency. Evaluation need not be perfect at birth but must be explicit enough for two people to discuss the same case without relying solely on opinion.

Is there an evolution owner?

Can someone prioritize changes, approve experiments, reject requests, and decide if learning becomes product?

If not, create a decision ritual with product, technology, and business. Without this role, the team accumulates usage signals but the organization does not turn these signals into choices.

Are autonomy limits clear?

Are there described situations where AI must request confirmation, route to a human, limit response, or interrupt an action?

If no limits exist, list risk scenarios before treating the problem as a future improvement. Autonomy without explicit limits often seems efficient until the first difficult case appears.

Is the change criterion defined?

Does the team know how to differentiate incident, operational adjustment, quality improvement, and product proposal change?

If not, create a simple decision matrix. It need not be bureaucratic. It must prevent every signal from becoming backlog and every urgency from becoming product change.

Is there a review cadence?

Is there a minimum frequency to review errors, complaints, usage metrics, and decisions made?

If not, define a short initial cycle with fixed agenda and responsible participants. Cadence is not to produce minutes. It is to prevent responsibility from disappearing after delivery.

When Responsibility Should Return to Human Decision

There are situations where expanding AI autonomy is not the best choice. In these cases, responsibility must return to people with authority to decide.

Responsibility should return to human decision when no clear acceptable error criterion exists, when the impact of a wrong decision is high for the client, when there is insufficient record to review what happened, when human review cannot occur in a timely manner, or when involved areas disagree on what constitutes a good response.

It is also advisable to limit AI when there is not yet sufficient evidence of real use, case variety, and review of ambiguous scenarios. The criterion is not the initial impression of functioning but the capacity to sustain proportional decisions when data limits, doubts, and interpretation conflicts arise.

Defining responsibility for AI software is, fundamentally, deciding who can authorize continuation, limitation, or stoppage. If this answer is not clear, delivery is not yet complete as an organizational capability.

For AI functionality already in use, delivery only becomes capability when these three names exist: who operates, who evaluates, and who decides evolution.

If you want to discuss this decision in your company’s context, talk to dooop.

Further Reading

Sources

To Continue This Reading

NEXT DECISION

Discuss Application in Your Company

Conversation about your software company context

Content by dooop. Registration allows linking this topic to the reader’s journey and tracking interest in the subject.

Conversation about your software company context

We will use your details to deliver this content and contact you about related topics.