Category:
Intelligent User Interface UX Design
Duration: Duration icon 11 min read
Created on: Created icon Sep 26, 2026

Agentic AI in healthcare: workflows clinicians can approve

A clinician’s signature used to cover work the clinician had done, or at least watched being done. Agentic AI in healthcare breaks that link. The software takes a goal from a care team, such as referring a patient or filing a prior authorization, does the reading, checking and drafting on its own, and hands over finished work to sign.

Whether that signature means anything depends on what the clinician can see. The simplest rule is to pause the agent for a clinician only when a step needs medical judgment or is missing medical information, and send administrative checks to staff. At each stop, the reviewer sees only their own steps, with the evidence beside each suggestion, which fits FDA’s January 2026 decision support guidance.

How agentic AI in healthcare changes what a signature means

Older clinical software gives one answer, such as an alert or a risk score, and then waits for a person. An agent does several things in a row: it collects information, applies rules, fills the gaps it can and hands over finished work. So when a clinician signs, the signature now covers steps they never saw happen.

In this article, a “step” means something the agent does in the real world, such as looking up a record, checking an insurer’s rules or sending a message. How the AI model was built is a separate question for technical documents. Vendors sell agentic AI in healthcare as AI workflow assistants or healthcare automation, and neither label tells you where the clinician signs.

Where the clinician signs matters now because physicians already use AI widely. The American Medical Association’s 2026 physician survey of nearly 1,700 doctors found that 81% use AI professionally, more than double the 2023 figure. It also found that 85% want to be consulted or directly involved in decisions about AI adoption. Those doctors will judge an agent partly by what it asks them to sign.

One referral, two reviewers

Picture a family doctor in a busy clinic who asks for a cardiology referral at the end of a visit. The next patient is already waiting. An agent takes the request from there, and two people check different parts of its work: a referral coordinator and the doctor who made the request.

  1. It reads the visit note and quotes the doctor’s own reason for the referral, with a link to the note.
  2. It finds cardiology practices that accept the patient’s insurance and offer the specialty the doctor asked for.
  3. It gathers recent lab results and the medication list, and notices that a heart ultrasound report mentioned in the note is not in the record. It stops here.
  4. It suggests how urgent the appointment should be, once the missing report is found or the doctor says to go ahead without it.
  5. It sends the referral to the specialist’s office after both people have signed off.

The coordinator checks the practice match in step 2, because a wrong match delays care. The doctor handles steps 3 and 4: first the missing report, then the urgency level, shown next to the note and results it is based on. Nobody is asked to approve the whole run with a single click at the end.

Step 5 is the point of no return. Once the referral reaches the specialist’s office, pulling it back takes a phone call, and the patient’s information has already been shared. Giving each person only the steps they answer for is what we call step-level approval. It is also where visual agent design for clinical work has to start.

Too few stops or too many: how review breaks down

Clinician review of an agent fails in two opposite ways: too few stops invite rubber-stamping, and too many teach people to click through. In the referral example, the dangerous version is a polished draft that hides the missing ultrasound report. The design job is to place a few stops only where medical judgment is actually needed.

The first failure is called automation bias. FDA’s 2026 guidance describes it as “the propensity of humans to over-rely on a suggestion from an automated system.” Between patients, a doctor is thinking about the person just seen, the person waiting and an unfinished note. A referral that looks complete but silently lacks the heart ultrasound is easy to sign.

Alert fatigue is the second failure. AHRQ’s Patient Safety Network describes it as busy clinicians becoming desensitized to safety alerts, and notes that most alerts from the systems doctors use to place orders do not matter clinically. An agent that asks for confirmation at every step would teach doctors to ignore it, which defeats the purpose of the extra stops.

Administrative steps should still pause when the agent is unsure, below a confidence level that the workflow’s owner has set and tested on real cases. Showing that confidence clearly is its own design problem, covered in Fuselab’s article on AI transparency in UX.

What FDA expects from agentic AI in healthcare

FDA’s Clinical Decision Support Software guidance, dated January 29, 2026, explains when decision support software for clinicians stays outside medical device rules. One condition is that the clinician can independently review the basis for each recommendation. For an agent, that becomes a screen requirement: every clinical suggestion shows what it is based on at the moment of signing.

The guidance adds two points that matter for agents. Software for urgent, time-critical decisions does not meet this condition, because FDA says a clinician “is unlikely to have sufficient time to independently review the basis of the recommendations.” And software that acts on its own suggestion, instead of handing it to a person, raises a harder question: whether it replaces the clinician’s judgment.

For other clinical functions, the guidance lists what a clinician should be able to see before relying on a recommendation. The table below translates that list into plain design terms for an agent’s review screen. It is an interpretation, not a legal checklist, and each row is something a design team can test with practicing clinicians.

What the clinician should be able to see What that means on the review screen
What the software is for, and which clinicians and patients it was built for Label each suggestion with the role and patient group it was built for, and flag it when a patient falls outside that group
Which information it used, and where that information came from List the inputs behind each suggestion, with source and date, one tap away
How the software was built and tested Keep this in reference material linked from the screen, not on the signing screen itself
What is missing or unusual in this patient's data Stop and name any missing input before showing a suggestion that depends on it

Some vendors have started to build part of this into their products. AWS describes Amazon Connect Health as letting clinicians “tap any AI output and view the underlying evidence immediately,” which partly addresses the second row of the table. Deciding which steps stop, and naming missing information before a suggestion appears, is still up to the team designing each workflow.

Our reading is that this condition is easier to meet one medical decision at a time than one finished workflow at a time. A single sign-off at the end asks the doctor to check the missing report, the urgency level and the insurance match all at once. Splitting them keeps each piece small enough to check properly.

FDA guidance explains how the agency reads the law and is not itself binding, and rules in this area are still changing, so check the current version before relying on it. This is a design reading, not legal or clinical advice. It does not describe the regulatory status of Health Monitor, ClyHealth, Radiology Queue or any other Fuselab project.

Agentic AI use cases in healthcare, grouped by who reviews them

Agentic AI use cases in healthcare fall into four groups, based on what each step decides and who checks it: administrative work such as prior authorization, clinical drafting such as discharge summaries, research support, and patient-facing tools such as symptom intake. Because one agent run can cross several groups, approval rules belong on individual steps, not on whole products.

Group Examples Who reviews
Administrative work Prior authorization packets, insurer and practice matching, scheduling, medical coding, claims paperwork Staff at first, then sampled checks when audits support it. Any medical-necessity statement goes to a clinician.
Clinical drafting Urgency levels, draft orders, discharge summaries, visit notes drafted from the conversation, draft replies to patient portal messages The responsible clinician, one decision at a time
Research support Finding eligible patients for studies, sorting published research, measuring how common a disease is A researcher, who checks that results can be reproduced
Patient-facing tools Symptom intake and self-triage assistants Decided per tool after a regulatory review, since the FDA condition above covers only software for clinicians

The referral example crossed two of these groups: the practice match was administrative, and the urgency level was clinical. That mix is normal in agentic AI in healthcare, and a single approval rule for the whole agent would have sent both checks to the same person. The next two sections look at products from Fuselab’s healthcare UX design work, built for the kinds of clinicians who would do this reviewing.

Health Monitor: keep the data dense, make the gap stand out

Health Monitor is an EHR redesign we managed for doctors and nurses in primary care, emergency and critical care. It suggests a rule for agent review screens: keep the dense patient record clinicians trust, and make the few items behind a suggestion stand out within it. A missing input needs its own look, so it does not blend in with abnormal results.

Learning time mattered too, because working in an ER “does not provide a lot of free time for learning new technologies.” The research also found that “the more dense the data is the better”: clinicians felt more confident with more information in view. That argues against a stripped-down approval screen, although it says nothing about how accurately people review.

Health Monitor was not an agent product, and the project measured nothing about agent review. Many ER decisions are also the time-critical kind that FDA’s guidance excludes. A future agent supporting decisions that cannot wait should not assume it escapes medical device rules, however clear its review screen is.

ClyHealth: showing the reasoning before the provider approves

ClyHealth, an AI-powered healthcare platform Fuselab designed, is the nearest example in this article of the review screen described above. Its supplement module generates a daily protocol for each patient, and providers review the reasoning before approving it. The project’s design principle is stated plainly: “a recommendation without visible justification does not get followed.”

The AI assistant reads the patient’s complete medical record and bases its answers on that patient’s own test results, lab history and genetic data. That matters for review. A provider can only check a suggestion against evidence that belongs to the patient in front of them, and generic guidance gives the reviewer nothing specific to verify.

ClyHealth’s dashboard is built in three layers. Critical patient indicators are visible on first load, clinical detail is one action away, and full biomarker breakdowns sit deeper still. That structure carries over well to an agent’s review screen: the suggestion and its main reason up front, the supporting inputs one tap away, and the full data available without crowding the moment of signing.

Still, ClyHealth recommends and a provider approves, which makes it the simpler case. Agentic AI in healthcare adds steps the provider never watched, and each clinical step would need the same visible reasoning. Whether providers review a chain of agent steps as carefully as a single recommendation is still an open question, and neither project described here measured it.

What design cannot fix

Design cannot make up for weak evidence. A scoping review of agentic AI in healthcare published in npj Digital Medicine in March 2026 found only seven eligible studies. Just one was a trial involving patients, and six of the seven had not yet been implemented in real-world settings. Treat vendor outcome claims with skepticism, this article makes none.

Nor can design settle whether software counts as a medical device, since that depends on what the software is meant to do. On Radiology Queue, a scan review platform we managed, x-ray specialists record each approval or rejection themselves. A clear review step like that is good design, but it does not answer the regulatory question, which this article does not assess for any project.

Review volume is the last limit. No screen can help a doctor asked to review a full shift of agent suggestions, however well each one is laid out. The first fixes there are organizational: move administrative review to staff, and limit how many agent workflows reach one clinician, so each review still gets real attention.

What to settle before design starts

Planning agentic AI in healthcare comes down to three decisions made before design starts: who reviews each step, where the evidence for each suggestion lives, and how the agent records what it did. The workflow owners make the first one, sorting every step by type and sending clinical and borderline steps, such as a medical-necessity statement, for regulatory review.

Data comes next. Datamonitor Healthcare, the dashboard we designed for Informa’s pharmaceutical users, needed its data cleaned and structured before the interface could work. An agent’s evidence has the same need: it must be findable and linkable, whether it sits in structured fields or free-text notes. Name the systems the agent reads and writes, and the standard, such as HL7 v2 or FHIR.

The last requirement is a record. The Qgen Health Lab oncology research platform had to let researchers repeat experiments and reproduce their results, so grouping and documenting each experiment was central. An agent needs the same kind of record of its own actions, one that a second person can follow step by step and question later.

Once those decisions are made, the review screen’s structure is largely set before anyone opens a design file. What remains is making sure that at every stop the clinician can see, in a few seconds, what they are being asked to approve and what it is based on.

Frequently asked questions

Is agentic AI in healthcare the same as generative AI?

Agentic AI in healthcare and generative AI overlap, but they are not the same thing. Generative AI produces text, images or other content, while agentic AI takes steps toward a goal, such as routing a referral, often using a generative model along the way. Either one needs clinician sign-off when its output goes out under a clinician’s name.

Is a clinical AI agent a medical device?

A clinical AI agent may or may not be a medical device, depending on what each of its functions is meant to do. US law excludes some software from device rules, including administrative tools and certain decision support for clinicians. Software that analyzes medical images or signals usually remains a device, so regulatory counsel should review each function before launch.

Does a HIPAA certification prove a vendor handles patient data safely?

HIPAA certification has no official standing: HHS states that no standard requires it and that it does not recognize private certifications. Better questions are which HIPAA Security Rule protections the vendor uses and how access controls cover what the agent does. A vendor handling patient data for a provider also needs a business associate agreement, the contract HIPAA requires, covering the agent’s work.

What are the best AI agents for healthcare teams?

The best AI agent for a healthcare team depends on the workflow more than the brand. For administrative work, look for insurer rule coverage, audit logs and spot-check controls; for clinical drafting, look for step-level stops, clearly named missing information and visible evidence behind each suggestion. Treat a vendor that cannot show a working review screen on a realistic case as not ready.

Can an AI agent act without clinician approval?

An AI agent can handle some administrative steps, such as scheduling, with staff spot checks once audits show that is enough. Those steps should still pause when the agent is unsure. In clinician-facing workflows, steps that involve medical judgment, such as an urgency level or a draft order, should always stop for the responsible clinician.

How do you test whether clinicians review agent suggestions properly?

Testing clinician review starts with early usability testing (formative evaluation), where practicing clinicians work through realistic cases that include a missing input or a wrong suggestion. This finds design problems during development but does not prove a design is safe. After launch, audit a sample of signed suggestions against the record to see what reviewers missed.

How is agentic AI different from traditional clinical decision support?

Agentic AI and traditional clinical decision support differ in how much happens before a clinician looks. Traditional decision support flags a single point, like a drug interaction, at the moment of ordering. Agentic AI completes several actions first, so its design has to decide which of those steps pause for the clinician and what evidence travels with each one.

Author

Vlad Bobu

Project Manager Health

10

Years of experience

6

Years in Fuselab

Vlad brings over 10 years of experience at the intersection of healthcare and technology. Trained as a biomedical engineer, he now owns projects end-to-end at Fuselab Creative: from proposals and contracts to UX research, delivery, and client success with a focus on medical device interfaces, AI-driven EHR systems, and preventive medicine platforms built to FDA guidelines and UX benchmarks. Fascinated by how the brain processes information and how culture shapes user decisions, Vladimir applies a deeply human-centered lens to every digital product.