Analytics data allows us to understand what users do on a website, but it does not always explain why they struggle to complete a task. We may identify a drop-off during checkout, few clicks on an element, or a low conversion rate, but we still need to understand what is happening behind those numbers.
User testing allows us to directly observe how someone uses a website while trying to complete specific tasks. Within a CRO process, it can help us uncover usability issues, questions, or points of friction that we can then compare with other sources of information and turn into optimization hypotheses.
Contents
What is user testing?
User testing is a research technique in which people who are representative of the target audience use a website, app, or prototype while completing a series of tasks. During the process, we observe how they interact with the interface, what decisions they make, and where they encounter difficulties.
For example, in an online store, we might ask a participant to find a particular product, check when it would be delivered, and proceed through the checkout process. The goal would not be to see whether they can follow a set of instructions, but to observe whether they can complete those actions in a reasonably intuitive way.
This can reveal issues that are not always obvious when we know the website well. A user may fail to notice an option that seems clearly visible to us, interpret a message differently, or follow a completely different path from the one we expected.
Although we can also ask participants questions, the value of the test does not come only from hearing their opinions. What they do during the session can be even more relevant than what they say. Someone may describe a page as easy to use while still needing several attempts to find the information they are looking for.
In this article, I will focus mainly on user tests designed to evaluate the usability of a website or a specific process.
What is user testing useful for in CRO?
Within a CRO process, user testing is particularly useful for understanding the possible causes behind a problem. Quantitative data can show us where a drop-off occurs, but watching someone try to complete the process provides a different kind of context.
Imagine that we identify a high abandonment rate at one stage of the checkout. We can review metrics, events, or funnels to measure the problem, but those data will not always tell us whether users are unsure about shipping costs, cannot find a payment method, misinterpret a field, or receive a message they do not understand.
A user test can help us identify some of these points of friction. If several participants stop at the same point, try to interact with an element that does not behave as expected, or ask the same question, we have a signal worth investigating.
User testing can also be useful when there is not yet enough quantitative data available. For example, before launching a new process, during a redesign, or while working with a prototype. In these cases, we can identify important issues before they begin affecting real users.
This does not mean that a user test automatically allows us to measure the impact of a problem on conversion. We normally work with a small number of participants and in a testing situation that does not exactly reproduce a real visit. Its main value lies in helping us discover and understand potential problems, rather than determining how many website users are affected by them.
Types of user testing
User testing can be conducted in person or remotely, but one of the most important distinctions is whether or not someone moderates the session. Both approaches can be useful, although they offer different levels of interaction.
Moderated user testing
In a moderated test, a researcher accompanies the participant throughout the session. They can introduce the tasks, observe what happens, and ask questions when they need to better understand an action or comment.
This makes it possible to explore situations that we had not anticipated beforehand. If a user pauses for several seconds on a page, for example, we can ask what they are thinking or what they expected to find at that point.
The presence of a moderator can also introduce bias. A gesture, an additional explanation, or a question phrased in a particular way may give the participant clues. For this reason, it is important to intervene only as much as necessary to understand the experience without helping them complete the tasks.
Unmoderated user testing
In an unmoderated test, the participant completes the session independently by following instructions prepared in advance. The screen is usually recorded and, in some cases, participants are also asked to explain what they are thinking out loud.
This format makes it easier to run tests with participants in different locations and removes the need to be present during every session. It can also reduce some of the influence that a moderator may introduce during the interaction.
The main limitation is that we cannot react to what happens. If a user misunderstands a task or an interesting situation arises, we cannot ask a follow-up question at that moment. This makes preparing the instructions particularly important.
Both moderated and unmoderated tests can be conducted remotely. A moderated session, for example, can take place over a video call while the participant shares their screen. The most suitable approach will depend on the type of product, the participants we need, and the resources available.
How to prepare a user test
A large part of the value of a user test depends on how it has been prepared. If the objective is too broad, we choose participants who are not closely related to our audience, or we create tasks that tell people exactly what to do, the results will be much harder to interpret correctly.
Before starting the sessions, it is therefore important to define what we want to learn, which users we need, and which situations we want to observe.
Define the objective of the test
The first step is to clarify what we want to investigate. An objective such as “analyze the usability of the website” is too broad and will probably lead us to observe many things without knowing afterwards which ones are actually relevant.
It can be more useful to focus on a specific part of the experience. For example, checking whether users understand the delivery options, can easily find the returns policy, or can complete the registration process without assistance.
Having a clear objective also helps determine which tasks to include and which participants we need. A test does not necessarily have to focus on a single question, but the different tasks should respond to objectives that have been defined beforehand.
In CRO, we can also start from problems that we have already identified through other sources. If the data shows an unusually high drop-off at one part of the funnel, the test can focus on understanding what happens during that step.
Select the participants
Participants should be sufficiently similar to the people who would actually use the product or complete the task we want to analyze. It is not always necessary to replicate every audience segment exactly, but we should include the characteristics that could have a meaningful effect on behavior.
For example, if we want to analyze the process of signing up for a financial product aimed at people with no previous experience, using only participants who work in the financial sector could produce an unrepresentative view. They would probably understand concepts and processes that create uncertainty for regular users.
The same can happen if we use colleagues who already know the website well. Although they can help identify obvious issues at a very early stage, their prior knowledge means they will not navigate the site in the same way as someone visiting it for the first time.
If there are groups of users with clearly different needs, it may also be necessary to treat them separately. The same flow may work perfectly for returning users while being confusing for people who have never used the service before.
Prepare the tasks and questions
The tasks should represent situations that a user could naturally encounter. Instead of specifying every step they need to follow, it is better to give them an objective and observe how they try to achieve it.
For example, if we want to find out whether the filters in an online store are easy to use, it would not make much sense to ask the participant to “open the filters, select a size, and sort the products by price.” In that case, we would mainly be testing whether they can follow our instructions.
A more natural scenario might be that they need to find an item in their size within a certain budget. From there, we can observe whether they use the filters, the search function, category navigation, or any other path.
We also need to be careful with the questions we ask. Asking “Was it difficult to find the delivery information?” already introduces the possibility that there was a problem. A more neutral question about how they found the process allows participants to describe the experience in their own words.
This is particularly important because participants often want to be helpful and complete the test correctly. If our wording gives them any indication of the answer we expect, they may consciously or unconsciously adjust their behavior.
Run a pilot test
Before running all the sessions, it is advisable to test the study with one or two participants. This initial test helps us check whether the tasks are understood correctly and whether they actually generate the information we expected to obtain.
Sometimes an instruction that seems clear while we are preparing the study is interpreted differently by participants. We may also discover that a task is too easy, too difficult, or that we are providing information that influences the path they take.
Fixing these issues after all the sessions have already been completed can be difficult. A small pilot test allows us to adjust the instructions, the order of the tasks, or the questions before using the test with the remaining participants.
How many users are needed?
There is no single number of participants that works for every test. The sample size will depend on the objective, the type of research, and the diversity of users we need to represent.
For qualitative usability testing, around five participants per user group is commonly used as a reference. The first participants tend to reveal many of the most obvious problems and, as we add people with similar profiles, we are more likely to start seeing the same behaviors repeated. This does not mean that five users guarantee that we have found every problem, or that this reference can be applied automatically to every study.
If there are groups of users with clearly different needs or behaviors, we will probably need participants from each of them. This reference also changes completely when we want to obtain quantitative results: a small sample can help uncover points of friction, but it cannot tell us what percentage of users experience a problem or allow us to statistically compare two versions.
For this reason, within CRO it is more useful to use these tests to identify potential problems and patterns than to measure their exact frequency. If three out of five participants encounter the same difficulty, we have an interesting signal to investigate, but we cannot conclude that 60% of real users will experience the same problem.
How to conduct a user test
At the beginning of the session, it is useful to briefly explain what will happen and make it clear that we are not assessing the participant’s abilities. The goal is not to find out whether they know how to use the website correctly, but to observe how the experience works while they try to complete the tasks.
During the test, we can ask participants to explain what they are thinking as they navigate. This can help us understand why they make certain decisions, what they expect to find when they click, or what doubts they have before moving forward.
However, we should not interpret every comment literally either. Users do not always know exactly why they performed an action and may try to rationalize decisions that were actually quite intuitive. This is why it is useful to combine what they say with what we actually observe during the session.
When a difficulty appears, the moderator should avoid stepping in too quickly. If someone cannot find a button for a few seconds, immediately helping them would remove precisely the information we want to obtain. We should also avoid confirming whether an action is right or wrong, as even a change in tone or a gesture can influence what they do next.
Maintaining a relatively neutral position allows us to observe how they try to solve the problem on their own. If we need to better understand something that has just happened, we can use open-ended questions about what they expected to find, what they were looking for, or what led them to make a particular decision.
How to analyze the results
Once the sessions are complete, the goal should not be to collect every comment made by participants, but to identify patterns that help us understand where problems may exist.
We can review which tasks were completed, where difficulties appeared, which paths users followed, and which elements caused uncertainty. It is also useful to pay attention to behaviors that occur repeatedly, even if participants do not mention them directly.
For example, none of the users may say that the returns information is difficult to find, but several of them may visit different pages before locating it. That behavior may be more useful than a general comment such as “the website is easy to use.”
We can also record data such as task success or the time required to complete each task. These can help us compare sessions and identify differences, but we need to be careful when interpreting them with a small sample. In a qualitative test, these metrics mainly support our understanding of behavior rather than estimate the performance of the entire user base.
Based on the sessions, we can group similar issues and assess which ones appear to be more relevant. A point of friction that prevents someone from completing a purchase will normally be more important than a minor uncertainty that does not affect the journey. A problem that repeatedly appears across different participants also deserves particular attention.
This does not mean that individual cases should be ignored. A single user may uncover an important issue that nobody else has identified. The important distinction is not to automatically assume that this behavior represents the entire audience. The test results should help us decide what deserves further investigation rather than turn every observation into a general conclusion.
How to integrate user testing into a CRO process
User testing works particularly well when combined with other sources of information. It can help us better understand a problem that has already been identified or uncover new points of friction that we can investigate in more detail.
For example, a funnel may show us that there is a significant drop-off during checkout. Several user testing sessions may then reveal that some participants struggle to understand the delivery options. Together, these sources begin to provide a more complete explanation of the problem.
We can also compare these findings with session recordings, surveys, customer support feedback, or analytics data. If different sources point to the same friction, we have stronger reasons to prioritize it.
From there, we can formulate a hypothesis. If we believe users are abandoning the process because the delivery options are not clear enough, we can propose a change that makes this information easier to understand and define which behavior we expect to improve.
The solutions suggested by participants themselves are not necessarily the ones we should implement. A user may say they would add a button, move an element, or completely change a process. That feedback can help us understand the underlying problem, but finding the best solution still requires further analysis.
Depending on the change, we may validate it through an A/B test, measure its performance after implementation, or test the experience again with users. In this way, user testing becomes another source of information within the CRO process: it helps us understand problems and build hypotheses, while other techniques allow us to determine whether the changes ultimately produce the expected effect.
In conclusion, user testing allows us to directly observe how real people try to complete tasks on a website and identify points of friction that are not always clearly visible in quantitative data. To make these tests useful, it is important to define what we want to investigate, select suitable participants, and create tasks that do not influence their behavior.
Within a CRO process, the results should not be interpreted as conclusions that are representative of all users. Their main value lies in uncovering problems, providing context, and generating hypotheses that we can then compare with other sources of information and validate using the most appropriate method.

