Why a Medical Device Can Look Ready After Formative Testing—and Still Fail Human Factors Validation
What late-stage testing reveals when early studies were too helpful, too fragmented, or too forgiving.

A realistic usability study places the participant, moderator, device, and test conditions under the same scrutiny.
Four formative studies. Four rounds of improvement.
Participants complete the tasks, and the team hears the words every development group wants to hear: clean, clear, intuitive.
By the last study, the difficult usability problems seem resolved. The interface looks ready. Then human factors validation begins.
Representative users work with a device that reflects the final design, under realistic conditions, without the moderator influencing how they perform the tasks. A clinician misses a critical setting. Another user selects the wrong option, notices the mistake, and recovers. A third cannot tell whether setup is complete after an interruption.
The room gets quiet. People question the participants, the scenarios, or the scoring. Sometimes the protocol truly does need attention. More often, validation is revealing something the formative program made easier to overlook.
THE SIX WARNING SIGNS

Moderator assistance

Training that is still fresh

Screens tested without the workflow

Recovered errors counted as success

Changes made after the final formative

Participants unlike the real users
“The goal is not reassuring results. It is the truth while the design can still change.”
Formative testing should expose risk—not create confidence.

The Moderator Became Part of the Interface
Interaction during formative testing is not automatically a problem. Early in development, moderators may ask participants to think aloud, pause a task, or explain what they expected to happen. That is often how the design team learns.
The problem starts when the team stops separating insight gained through discussion from performance achieved independently.
A participant hesitates. The moderator asks, “What are you looking for?” Someone taps the wrong control, and a team member explains that the screen is still being refined. The participant becomes stuck, so the moderator gently redirects the conversation. None of this feels like assistance. It feels like productive research.
But the performance record has to be honest. Did the user complete the task independently? After a neutral probe? After a hint? Or only after the moderator explained the design intent? Those are very different outcomes. During validation, participants should use the device as independently and naturally as possible. When the moderator has quietly become part of the interface, validation may be the first time the interface is truly tested on its own.

The Training Was Still Doing the Work
Another common source of false confidence is testing people immediately after training. Ten minutes earlier, the trainer demonstrated the workflow, explained the terminology, and pointed out the important controls. Of course the participants perform well. The information is still easy to retrieve.
That may not resemble actual use. Some devices are used hours or days after training. Others are used infrequently, during a stressful event, when the user cannot remember every instruction.
Validation training should approximate what actual users receive, and testing should not occur immediately afterward. Depending on the device, an hour may be reasonable. In other situations, one or more days may better represent the expected conditions of use.
There is no universal waiting period. The right question is whether the training, delivery method, and delay reflect the way the device will really be used.

The Team Tested Screens Instead of the Workflow
Formative testing often focuses on individual parts of an interface: the alarm screen, a revised setup sequence, or a new confirmation message. That is useful early in development because the team can compare alternatives without building the entire system.
Users, however, do not experience a medical device as a collection of isolated screens. They experience a workflow.
Problems often appear in the spaces between the parts. A user forgets information shown several screens earlier. A setting from the previous procedure remains active. Two screens look almost identical even though the device is in a different state. A task is interrupted, and the user cannot tell where to resume. A handoff occurs, and the next person does not know what has already been completed.
Every individual screen may look understandable when tested alone. The complete workflow can still break down. Late-stage formative work should place critical tasks inside realistic scenarios and examine transitions, interruptions, persistent settings, recovery, and handoffs. That is where the seams become visible.

A Recovery Was Counted as Success
A nurse reaches a confirmation screen, hesitates, and selects the wrong option. She notices the problem, goes back, and corrects it. Everyone relaxes. She recovered.
The observation note may say: minor difficulty; participant resolved independently. Technically, the task was completed. From a risk perspective, something important still happened.
FDA describes a close call as a situation in which a user has difficulty or makes a use error that could result in harm, but then recovers before the harm occurs. Close calls should be recorded and investigated. Repeated attempts and apparent confusion can also indicate a potential use problem.
The question is not simply, “Did the user finish?” The team also needs to understand what caused the initial error, whether the user would always recognize it in time, and whether recovery would still be possible under pressure or distraction.
A finding does not become unimportant because only one participant encountered it. When the potential harm is serious, one close call can be enough to demand careful root-cause analysis.

The Device Changed After the Last Study
The final formative study used version 0.8. Validation will use version 1.2.
Between those versions, engineering corrected defects. Marketing added a screen. Terminology changed. A confirmation step moved to a different point in the workflow. Each change seemed small, and none appeared important enough to justify another study.
Taken together, the team created an interface that no representative user had ever used.
Human factors validation should evaluate a user interface that represents the final design, including the device, labeling, instructions, and training. When changes occur after the last formative study, the team should assess whether they affect user behavior, critical tasks, or use-related risk.
Not every change requires another full study. But every relevant change deserves a documented impact assessment. Otherwise, the team is not confirming the design that performed well in formative testing. It is testing a new design for the first time during validation.

The Participants Were More Capable Than the Real Users
Formative participants are sometimes more experienced, more motivated, or more familiar with the technology than the intended user population. Company employees may know how the system is supposed to work. Clinical advisors may be more technically confident than typical users. Frequent device users may rely on shortcuts that a novice would not know.
These participants can provide valuable feedback. They can also make a difficult interface appear easier than it is.
Validation participants should represent the intended user groups and the range of characteristics that could affect safe operation. Depending on the device, those characteristics may include age, education or literacy, sensory or physical limitations, clinical specialty, experience, and familiarity with similar products.
Representative does not mean intentionally recruiting incapable users. It means avoiding a test group that is more prepared, more careful, or more knowledgeable than the people who will actually use the device.
Treat the Last Formative as a Dress Rehearsal
The answer is not necessarily more formative studies. It may be one honest late-stage evaluation.
A useful late-stage rehearsal should include:

A device that closely represents the final user interface

Representative users from each intended user group

Realistic scenarios built around critical tasks

The training actual users receive, with a realistic delay

No coaching—record errors, close calls, hesitation, workarounds, and confusion
Teams sometimes resist this approach because a realistic rehearsal might uncover problems. That is exactly why it is valuable. A finding during formative work can lead to another design iteration. The same finding during validation can lead to redesign, protocol revision, additional documentation, and repeat testing.Validation will eventually reveal how the device performs without the design team in the room. The only real choice is whether the team learns that lesson early or late.
Find the problems before validation finds them.
Preparing for Human Factors Validation?
A focused late-stage UX and human factors review can identify where moderator help, training assumptions, incomplete workflows, post-formative changes, or unrepresentative participants may be hiding risk.
Areteworks helps medical device teams connect formative findings to the user interface, critical tasks, and use-related risk analysis before the validation protocol is locked.
| Areteworks offers a free 30-minute medical device UX review. Contact Areteworks |
About the author
Steven Liu has designed FDA-regulated medical device interfaces for more than 25 years, including systems used in critical care, imaging, therapy, and scientific instrumentation. Areteworks specializes in user research, workflow design, interface development, formative usability evaluation, and human factors documentation for medical devices and scientific instruments.