Ler original em português

← All content

dooopSoftware · Organization 11 min

How to Assign Responsibility for AI Evaluations

Separate who defines criteria, who executes tests, and who decides consequences to evaluate AI without confusing evidence with approval.

Published on September 6, 2026

MAIN THESIS

Evaluating AI without separating roles mixes the standard, evidence, and consequence.

The minimum matrix defines criteria, execution, and decision before testing.

When evaluating artificial intelligence, reviewing responses is only part of the work. The team needs to separate three responsibilities that often get mixed: who defines the criteria, who performs the evaluation, and who decides what will be done with the result. When these roles are ambiguous, the team may test extensively but cannot explain whether a failure came from the standard, the execution, or the decision to release.

When AI Evaluation Requires Separate Responsibilities

Separation becomes necessary when AI enters a sensitive stage of the product or operation. It does not have to be a high regulatory risk decision to deserve care. It is enough that a wrong response causes rework, confuses a user, hides an incident, or pushes the team toward a poor decision.

Imagine leadership asking to use AI in classifying support requests. One person creates test examples, another reviews some responses, someone from product approves limited use, and weeks later a failure appears. The question comes quickly: was the criterion wrong, was the evaluation poorly executed, or was the release premature?

If no one can answer, there is a lack of decision design before testing.

The DORA 2025 presentation describes AI as an amplifier of existing organizational strengths and weaknesses and highlights the importance of the organizational system for return on investment. This observation does not decide how your company should appoint responsible parties but supports caution in evaluating AI within the team's decision, documentation, and learning system.

Therefore, the question "who evaluates?" is incomplete. A better question is: who defines the expected quality, who produces evidence, and who assumes the consequence of the decision?

Three Roles That Should Not Be Treated as One

In AI evaluations, there are at least three roles that need to appear explicitly. They can be held by different people or combined in small teams, but they should not disappear inside a generic responsibility called "review."

The criteria owner translates objective, risk, and minimum acceptable standard. This person defines what counts as good, acceptable, and unacceptable before execution. They also clarify which errors are tolerable, which require human review, and which prevent use in that context.

The evaluation executor applies the protocol. They gather cases, run tests, record evidence, note divergences, and point out uncertainties. Their role is not to "feel" if the AI is good. It is to produce a reliable basis so that another person or the responsible group can decide.

The decision maker defines what happens next. They can release, limit scope, request a new round, change the criteria, return to solution design, or stop use. This role is responsible for the operational, technical, or product consequence of the decision.

Confusion arises when the same person invents the standard during testing, chooses which evidence to consider, and still releases use because it "seems sufficient." In a small team, combining functions may be inevitable. The problem is not combining. The problem is hiding the conflict.

The practical question is simple: if something goes wrong, can the team say which role failed or which decision needs review?

How to Choose Who Defines the AI Evaluation Criteria

The AI evaluation criteria owner does not need to be the most senior person in the room. They need to understand the impact of the decision the AI will influence.

This role should be close to the affected user, the product objective, and operational risk. In a ticket classification, for example, the person defining the criteria needs to know when a wrong category only causes delay and when it hides a problem that should receive immediate attention.

Good criteria for choosing the criteria owner are:

  • proximity to business or product impact;
  • sufficient knowledge of the affected user, internal client, or operation;
  • ability to distinguish tolerable error from critical error;
  • authority to adjust the standard when unforeseen cases arise;
  • willingness to make trade-offs explicit, not just ask for "more accuracy."

This last point changes the conversation. Asking "the AI needs to be right" is not a criterion. Saying "ambiguous answers should go to human review" already begins to be a criterion. Saying "this category cannot be assigned automatically without evidence in the ticket text" is even better.

The evaluation also needs to be connected to the type of AI-assisted decision. If AI only suggests an internal label, the standard may differ from AI that triggers communication to the client. The same model, interface, and perceived accuracy rate may require different criteria depending on the consequence.

This point relates to a previous maturity decision. Before spreading initiatives, it is worth diagnosing whether the organization can sustain minimum criteria, roles, and records. The article on AI maturity helps look at this starting point without reducing maturity to a tool or proprietary model.

How to Choose Who Executes the Evaluation

The evaluation executor should be chosen for discipline in producing evidence, not for willingness to approve or reject the idea.

Executing an AI evaluation involves following an agreed protocol, recording cases, preserving difficult examples, reproducing results when possible, and separating observation from interpretation. It is work closer to organized investigation than expert opinion.

The person or group executing needs to be able to answer:

  • which cases were evaluated;
  • which criteria were applied;
  • which responses were correct, incorrect, or ambiguous according to the defined standard;
  • which uncertainties appeared during execution;
  • which situations were not covered by the initial criteria.

This role should not change the standard mid-test to save or condemn the solution. If the criterion seems inadequate, the executor records the limitation and recommends a new definition. Changing the standard after seeing the result weakens the evaluation.

Here, documentation ceases to be bureaucracy. DORA treats documentation quality by attributes such as clarity, ease of location, and reliability, and recommends active creation and maintenance of documentation on its documentation quality page. For AI evaluations, the practical application is to keep records that someone outside the round can find, understand, and question.

A good record does not need to be sophisticated. It needs to allow the team to reconstruct the decision. What was the objective? What was the risk? Who defined the criteria? What was tested? What was uncertain? What was decided?

Without this, the evaluation becomes oral memory. And oral memory is fragile when the team changes, when the model changes, or when delivery pressure increases.

How to Choose Who Decides After the Evaluation

The decision maker is not necessarily the person who understands AI best. It is the person responsible for the consequence of the change.

If the evaluation concerns a product feature, the decision may be with product and engineering. If it affects operation, someone responsible for operation needs to be involved. If it changes technical risk, technical leadership must have real power to restrict. The exact design varies, but the authority to decide must appear before testing.

The decision maker must have authorization to choose among alternatives, not just to stamp a recommendation. After an evaluation, the decision maker must be able to choose among:

  • releasing use within the evaluated scope;
  • limiting use to lower-risk cases;
  • requiring human review in ambiguous situations;
  • requesting a new evaluation round;
  • changing the criteria and repeating execution;
  • returning to solution design;
  • stopping use in that context.

This list matters because many evaluations start with an implicit decision: prove the initiative can proceed. When this happens, the executor tends to look for favorable evidence and the decision maker tends to treat limitations as minor pending issues.

A serious evaluation must accept the possibility of not releasing.

This does not mean paralyzing adoption. It means protecting the ability to learn from the test. The DORA page on learning culture relates learning culture to software delivery performance and proposes treating learning as an organizational investment. In AI evaluations, learning requires recording not only what worked but also why the team decided to restrict, repeat, or stop.

There is a relevant difference between evaluation and experiment. An evaluation checks if the solution meets defined criteria for a use. An experiment requires a testable hypothesis, impact measurement, and iteration. Microsoft describes the ExP as a platform to incorporate experimentation into the development cycle, validate hypotheses, measure impact, and iterate products. This does not turn every internal evaluation into a product experiment, nor does it mean any feedback automatically retrains a model.

Fictional Example: Evaluating AI in Ticket Triage

Consider a fictional example. A software company wants to use AI to classify support tickets before they reach the team. The goal is to suggest an initial category, not to respond automatically to the client.

Leadership defines that the first evaluation will not decide a broad deployment. It will only decide if AI can be used as an internal suggestion in low-risk categories. This scope avoids an overly generic evaluation.

The criteria owner is a product person together with a support leader. They define ticket classes, describe accepted examples, and separate critical errors from tolerable errors. A tolerable error may be a classification that changes the queue but remains visible for human triage. A critical error may be classifying a possible production incident as a common question.

The evaluation executor is a quality or engineering person who does not decide release. They apply the protocol on a sample defined by the team, record divergences, mark ambiguous cases, and do not change the standard during execution. When they find tickets that do not fit well into existing classes, they record the gap instead of creating a new category mid-test.

The decision maker is the person responsible for the support flow and internal product change. After analyzing the records, this person does not release AI for all ticket types. The hypothetical decision is to restrict use to low-risk categories, require manual review for any sign of production incident, and request a new round when classes are revised.

No effect of this example should be treated as an actual result. The hypothesis to measure would be whether the suggestion reduces ambiguity in triage without increasing critical errors. For that, the team would still need to define evidence, monitor real use, and review the decision on a scheduled date.

The value of the example is in the separation. One person defined the standard, another produced evidence, another decided the consequence. If something fails, the team can review the right point.

Checklist for Choosing Responsibility for AI Evaluations

Before the first test, fill out a short checklist. It prevents the evaluation from starting as informal conversation and ending as a decision without an owner.

  • What decision will be made with this evaluation? If the test does not change a release, restriction, redesign, or stop decision, it does not yet have sufficient purpose.
  • Who defines what counts as good, acceptable, and unacceptable? Choose someone close to product impact, knowledgeable about the affected user, and authorized to adjust the standard when unforeseen cases arise.
  • Who executes the evaluation without changing the criteria during the test? Choose someone able to follow protocol, record evidence, reproduce cases, and point out uncertainties without turning personal preference into rule.
  • Who decides what happens after the result? Choose the person responsible for the operational, technical, or product consequence of release. Seniority helps but does not replace real connection to the consequence.
  • What evidence will be accepted? Define examples, samples, records, human reviews, or measures before execution. Do not change the standard after seeing the result.
  • What is the stopping rule? Define when the evaluation should be stopped, redone, or escalated. Critical errors, low reliability of records, or ambiguous criteria should trigger review.
  • Where will the decision be documented? Record criteria, execution, result, decision, and justification in an easy-to-find and maintain location. Without reliable documentation, the evaluation becomes oral memory.

This checklist does not replace strategy, AI governance, or product design. It covers a more basic layer: preventing an evaluation from being treated as collective opinion without clear consequence.

If the company is still organizing priorities, it is worth connecting this decision to the AI roadmap and the business-connected artificial intelligence strategy. Responsibility for evaluation should not live isolated from the adoption plan.

The next step is to create a simple matrix before evaluating any relevant AI use: criteria owner, evaluation executor, and consequence decision maker. If these three fields cannot be filled, the team is not yet ready to trust the test result.

If you want to discuss this decision in your company’s context, talk to dooop.

Further Reading

Sources

NEXT DECISION

Discuss Application in the Company

Conversation about the software company context

Content by dooop. Registration allows linking this topic to the reader’s journey and tracking interest in the subject.

Conversation about the software company context

We will use your details to deliver this content and contact you about related topics.