Ler original em português

← All content

dooopSoftware · Learning · 12 min

How to Prioritize the Next Product Experiment

Choose tests by crossing uncertainty that changes real decisions with the ability to measure results without confusing usage, clicks, or apparent execution.

Published on September 6, 2026

CENTRAL THESIS

Easy ideas can occupy the wrong cycle. The best test reduces an uncertainty that changes a real decision.

Cross the relevance of the doubt with the ability to measure. If it cannot be interpreted, prepare before testing.

In many product cycles, good ideas compete for the same test window: an easy change, a recurring demand, a leadership bet, a technical improvement. The criterion to choose among them is to compare two things: which uncertainty, if reduced, would change a real product decision, and whether the team can test it without producing a misleading reading. When a hypothesis is relevant but not yet testable, the next step is to better prepare the question, measurement, and reading criteria.

When Many Ideas Compete for the Same Product Cycle

In a product meeting, it is common for each area to bring different evidence. Support brings recurring complaints. Data shows a drop in a flow step. Engineering sees a possible simplification. Leadership requests visible progress. These inputs serve different roles: some indicate observed problems, others suggest hypotheses, and others just pressure for movement.

Confusion begins when the team treats every idea as an equivalent candidate for testing. "Change the message," "swap the order of recommendations," "add a button," "use another model," "simplify the screen," and "ask for explicit feedback" all enter the same list. The conversation quickly turns into a dispute over preference, effort, or political urgency.

A good experiment is not just a small change. It is a disciplined way to reduce a doubt that blocks a decision.

This applies to digital products in general and becomes even more sensitive in AI features because an apparently successful interaction can hide low-quality results. In an intelligent feature, the user may click, accept a suggestion, or continue the flow without the product reliably solving the task. Therefore, the choice of experiment must come before the desire to "run something quickly."

If you are organizing a broader product learning agenda, it is worth connecting this decision to the guide on learning cycles in AI products. Here, the focus is more specific: deciding which hypothesis enters the next test cycle.

Separate Promising Ideas from Relevant Uncertainty

A promising idea can be interesting, elegant, and even easy to defend. Still, it only becomes a good experiment candidate when it points to relevant uncertainty.

Relevant uncertainty is a doubt whose answer changes a concrete decision. It can decide whether the team launches, changes, stops, expands, reduces, or postpones an initiative. If the answer would not change anything, the hypothesis may generate curiosity but does not deserve to occupy the next product cycle.

A practical way to separate ideas from uncertainties is to ask:

  • What decision would be different if this hypothesis is confirmed?
  • What decision would be different if it is not confirmed?
  • Was the problem observed in behavior, outcome, support, or operation, or did it arise only from internal opinion?
  • Is the risk of keeping this doubt open small, moderate, or high for the experience and product confidence?

Microsoft describes its experimentation platform, ExP, as a way to incorporate experimentation into the development cycle, validate hypotheses, measure impact, and iterate products. This reference helps reinforce a point: experimenting is not just publishing variations. It is connecting hypothesis, measurement, and product decision, as described by Microsoft.

The most common mistake is prioritizing the hypothesis with the greatest narrative appeal. "Users want friendlier answers" may be true, but perhaps the doubt blocking the decision is another: "Do the suggested answers resolve the customer’s request without rework?" The first phrase speaks of perception. The second points to a result that can change the feature’s architecture, data source, or flow design.

Assess Whether the Hypothesis Can Be Tested Without Illusory Reading

After identifying relevant uncertainty, the second dimension is testability. This is not synonymous with low technical effort.

Testability means the team can observe the hypothesis with sufficient quality to avoid self-deception. This depends on success criteria, observable population, possible comparison, guardrail metrics, user segments, and data quality.

A weak hypothesis for testing is one that can be implemented but does not allow confident interpretation of the result. The team changes something, monitors a dashboard, sees a positive usage variation, and declares progress. Only later do they realize the measured behavior did not confirm the result that mattered.

In AI features, this care is even more concrete. Anthropic distinguishes an agent’s execution trajectory from the effective result in the environment. A message saying the task is finished is not enough to prove the expected result happened. The evaluation needs inputs, success criteria, and verifiers, as explained by Anthropic.

Even when the product does not use autonomous agents, this separation is useful. A generated suggestion, a click on "use response," or immediate positive feedback may indicate apparent execution. But the real result may depend on later resolution, absence of rework, less escalation, or confirmation by another event.

Therefore, before authorizing an experiment, the question should not be only "Can we implement it?" It should also be: "Can we know if it worked for the right reason?"

This distinction relates to broader decisions about AI maturity and business-connected AI strategy. The point, for experiment choice, is more immediate: without reading criteria, the team may confuse seductive demonstration with operational value.

Use a Simple Matrix: Uncertainty Relevance Versus Testability

The most useful matrix to choose the next experiment crosses two questions.

The first: Is the uncertainty relevant to a real decision?

The second: Does the team have the ability to test this uncertainty without illusory reading?

When relevance is high and testability is also high, there is a good candidate for the next experiment. The team knows what it needs to learn and can interpret the result.

When relevance is high but testability is low, the hypothesis should not be discarded. It should be prepared. Perhaps success criteria need definition, events instrumentation, segment separation, verifier creation, or support data qualification. In this case, calling preparation an experiment only creates anxiety and noise.

When relevance is low but testability is high, the risk is spending energy on a comfortable change. It runs well, generates numbers, occupies the agenda, and changes no decision. This is a subtle type of waste because it looks like productivity.

When both relevance and testability are low, the idea may stay outside the experimentation cycle. If there is still some interesting signal, it can become qualitative discovery, exploratory analysis, or controlled prototype. But it should not compete with hypotheses that reduce more consequential doubts.

Criteria to Choose the Next Product Experiment

Use this checklist to score each candidate hypothesis from zero to two. The sum does not decide alone but forces the right conversation among product, data, support, and engineering.

  • Does the uncertainty change a concrete decision? Zero: even with the answer, the team would probably do the same. One: the answer helps but does not clearly change priority, scope, or risk. Two: the answer decides whether the team launches, changes, stops, expands, or reduces the initiative.
  • Is the doubt linked to an observable behavior or outcome? Zero: the hypothesis depends on generic perception or broad opinion. One: there are indirect signals, but the main behavior is not well defined. Two: the team can point to which action, outcome, or failure will be observed.
  • Is there a success criterion before implementation? Zero: success would be defined after looking at numbers. One: there is a main metric but no interpretation limit or guardrail metric. Two: there is a main metric, minimum criterion, and regression signals that prevent overly optimistic reading.
  • Can the test distinguish apparent execution from real result? Zero: the system may say it completed the task without proof in the environment. One: there is some verification but it covers only part of the result. Two: there is a verifier, event, review, or evidence confirming whether the expected result happened.
  • Is there volume, segment, or window sufficient to interpret the test? Zero: the test would be read with very few cases or too mixed an audience. One: there is data but segmentation or window may still distort the conclusion. Two: the team knows which group to observe, for how long, and which cuts to review.
  • Is the risk of a wrong conclusion acceptable? Zero: a wrong reading may worsen experience, affect confidence, or increase operational problems. One: the risk exists but can be contained with close monitoring. Two: the experiment has limits, monitoring, and stop criteria proportional to the risk.

How to interpret: hypotheses with high scores tend to be strong candidates for testing, provided there is no immediate operational restriction. Intermediate scores usually need adjustment before entering an experiment. Low scores indicate the idea should perhaps become discovery, data review, support analysis, or controlled prototype.

The checklist does not replace product judgment. It prevents judgment from being hijacked by enthusiasm, technical ease, or pressure for movement.

Applying the Matrix in a Call Center

Fictional example: a call center uses an AI feature to suggest answers to agents. Support reports some answers seem correct but may not resolve the customer’s request. Leadership wants to improve the feature in the next cycle.

Three hypotheses enter the conversation.

The first hypothesis is to change the tone of suggested answers to a more cordial language. It may improve perception but does not directly address the doubt about request resolution. If the observed problem is that the answer seems good yet does not resolve, cordiality may be a lateral improvement. As a priority experiment, relevance is limited for the current decision.

The second hypothesis is to prioritize internally reviewed sources when the request involves exchange policy. It reduces a concrete uncertainty: does the information source influence the chance that the agent forwards an answer that resolves the request? Testability depends on minimum conditions: identifying the exchange request segment, recording which source supported the suggestion, defining what counts as effective resolution, and monitoring regression signals such as increased escalation or later corrections.

The third hypothesis is to add a button for the agent to mark an incorrect answer. This idea may be very useful but may be preparation for future experiments. If the current cycle’s goal is to measure resolution, the button improves signal collection but does not necessarily test outcome improvement. It can enter before the main experiment if the team lacks sufficient evidence to distinguish accepted answer from resolved request.

By the checklist, the hypothesis about prioritizing internally reviewed sources tends to be the best candidate, provided the team can verify effective resolution and review segments. If this verification does not exist, the next cycle should prepare measurement, not publish the change as a complete experiment.

This example also shows why usage metrics are not enough. An increase in suggestion use can coexist with more rework. A decrease in average time can hide worsening in difficult cases. Google SRE recommends choosing monitoring considering data speed, calculations, visualization, and alerts, and reminds that averages can hide problematic behaviors and different audiences need different views, as described in the Google SRE Workbook.

Before Authorizing the Experiment, Define What Would Make You Stop

A product experiment does not only need success criteria. It needs interruption, review, or containment criteria.

This is especially true when the change can affect confidence, support, or operation. If the test only has optimistic metrics, the team tends to seek confirmation. If it also has guardrail metrics, it becomes harder to ignore collateral damage.

Before authorizing the experiment, define which signals would make the team pause or review the reading:

  • regression in a sensitive user segment;
  • drop in a guardrail metric linked to quality, rework, or confidence;
  • inconsistency between recorded event and effective result;
  • data quality problem compromising interpretation;
  • positive average behavior hiding worsening in a specific group.

Microsoft, when addressing experiment monitoring, recommends observing a broad set of metrics and segments to identify regressions and avoid premature interpretations during the test. In post-analysis, it also recommends verifying whether metric changes are compatible with the test design and whether data quality issues compromise interpretation before deciding on launch, as per Microsoft’s texts on monitoring and post-analysis.

This care does not make the process slow by principle. It reduces the risk of launching, expanding, or stopping a change based on fragile reading.

Choose the Test That Changes a Real Decision

The question "Which experiment will we run?" often seems operational. In practice, it reveals how the organization learns.

If the team always chooses the easiest test, it learns little about the most relevant decisions. If it always chooses the most ambitious hypothesis, it may produce results impossible to interpret. If it confuses usage signals with effective results, it risks improving the dashboard and worsening the experience.

The good choice lies at the intersection of a doubt that matters and an honest way to test it. Sometimes this leads to an experiment now. Other times it leads to preparation: defining success, instrumenting events, separating segments, reviewing data quality, or building verifiers.

Here, organizational capacity means being able to formulate hypotheses, measure, segment, decide, and stop when necessary. The value is knowing which doubts to reduce, in what order, and with what evidence standard.

For the next cycle, choose the hypothesis that combines high relevant uncertainty with good testability. If it is relevant but not yet testable, prepare the test before publishing the change. And before starting, answer aloud: if the result is positive, negative, or inconclusive, what will the team do differently?

If you want to discuss this decision in your company’s context, talk to dooop.

Further Reading

Sources

NEXT DECISION

Discuss Application in Your Company

Conversation about the software company context

Content by dooop. Registration allows linking this topic to the reader’s journey and tracking interest in the theme.

Conversation about the software company context

We will use your details to deliver this content and contact you about related topics.