Ler original em português

← All content

dooopSoftware · Product · 5 min

How to Test an AI Feature with Few Users

Use tests with few users to observe tasks, difficulties, and AI limits. Record learnings without presenting general conclusions.

Published on September 6, 2026

MAIN THESIS

A small round can guide a decision without proving a broad effect.

The test question should fit the type of observation the team can perform.

Testing an AI feature with few users can help discover usage difficulties, unclear expectations, and behavioral failures. The test needs to state this purpose and limit its conclusions to what was observed. Instead of trying to prove broad adoption with a small sample, choose a task, observe how it is performed, and record which product decisions the evidence allows.

Choose a Question That Observation Can Answer

Questions like “Will the market adopt it?” or “Did AI increase productivity?” require a different design than a short observation round. For an initial test, look for a concrete doubt: Does the user understand the proposal? Can they verify the response? Do they know how to correct information? Do they find the way when AI does not complete?

Also define the expected state of the task. Presenting a response and completing the work are not the same. The person may accept a text for lack of alternatives, abandon the execution, or correct the result outside the product.

Anthropic on agent evaluations distinguishes trajectory and effective outcome. This separation is useful for observation: follow what happened in the task beyond the conversation or initial impression about the response.

Choose Participants by the Context You Need to Understand

Record why each person participates and what type of use they represent. Domain experience, product familiarity, and task nature can change what the observation reveals. The choice should help examine the defined doubt without being presented as an automatic representation of all users.

Include, when relevant, people who will have different difficulties than the team that built the feature. A demonstration with those who know the internal workings may overlook explanation or expectation problems.

Do not hide who was excluded from the test. The report should indicate contexts not yet observed. This information helps choose the next round and prevents a localized conclusion from becoming a general product promise.

Prepare the Task Without Leading the Person to Success

Describe a usage situation and the goal, avoiding teaching the path you intend to evaluate. If the test examines whether the person finds how to review a response, do not indicate that button before observing their attempt.

Agree on which interventions will be made when the task stalls. A help request can reveal a relevant difficulty. If someone from the team solves the problem during the session, record the intervention instead of counting the result as an independent conclusion.

Also prepare cases that show the feature’s limits. An inadequate response or missing information can help evaluate correction, challenge, and continuity. It is not necessary to expose participants to real risks to examine these behaviors; use controlled conditions appropriate to the test.

Fictional Example: Observing a Classification Suggestion

Imagine a product that suggests categories for internal requests. The team wants to know if people understand the suggestion and can correct an ambiguous case. The session presents requests prepared for the test, including one that could receive more than one category.

The observation records how the person interprets the suggestion, which information they consult, and whether they find the correction alternative. If they accept the category because they did not realize they could change it, that is different from agreeing with the recommendation after reviewing the case.

The result may motivate a change in presentation or editing method. It does not allow stating that all users will behave the same or that the change will increase a business metric. Those would be questions for later investigation.

Record Observation, Interpretation, and Decision Separately

A useful research note distinguishes what the person did, the explanation offered, and the team’s hypothesis. “Looked for another path and asked for help” is observation. “Does not trust AI” is a broad interpretation that would require more evidence.

Microsoft ExP describes experimentation as hypothesis validation, measurement, and iteration. The qualitative round can help formulate these hypotheses. It does not need to be called a controlled experiment if its design does not allow that classification.

Organize findings by task and difficulty. Preserve cases that contradict the predominant interpretation. An exception may indicate a need for segmentation, a limit of the experience, or just a condition still poorly understood.

Decide the Next Step Without Expanding the Conclusion

When closing the round, choose between adjusting the experience, correcting a behavior, observing another context, or maintaining the solution for later evaluation. Relate each decision to the case that motivated it.

Record what has not yet been demonstrated. The absence of a failure in sessions does not prove it will not occur. A declared preference does not prove recurring use. Avoid turning small counts into an appearance of precision that the design does not support.

A test with few users will be useful when it reduces a specific doubt and guides a next action. Start with the question that can be observed, describe the limits, and maintain the difference between learning about the experience and proving a general effect.

If you want to discuss this decision in your company’s context, talk to dooop.

Further Reading

Sources

To continue this reading

NEXT DECISION

Discussing the application in the company

Conversation about the software company context

Content by dooop. Registration allows relating this topic to the reader’s journey and tracking interest in the subject.

Conversation about the software company context

We will use your details to deliver this content and contact you about related topics.