Five Failure States to Test in Healthcare AI Citation UX
Five evidence failures need five different product states.
A healthcare AI citation can fail in at least five materially different ways. The interface may have no supporting passage, lose access to a known source, block one user from restricted evidence, discover a newer edition or encounter sources that disagree. Showing the same generic error for all five leaves the user with the wrong explanation and the product team with no useful diagnosis.
This problem can begin before a retrieval model is live. The data point that frames this drill comes from Pharos Production's aesthetic-medicine citation UX case: an audit found six of seven citations in the design mockups were wrong, including four that were invented. Those figures describe one design-stage audit, not the error rate of a deployed medical AI system.
The useful buyer question is narrower than "Does the demo have citations?" Ask whether the prototype shows an actionable state when evidence changes. This is an acceptance tool, not clinical validation or regulatory classification.
Start with a state model, not a generic fallback
A citation marker on a successful answer is the easiest screen to design. The hard part is deciding what the product does when the evidence path breaks. Each state needs a trigger, truthful explanation, next action and owner.
Treat the displayed claim as the unit of control. One answer may contain four claims, and only one may lose support. Keep the source edition, visible claim, supporting passage, review decision and interface version separate. A green badge that survives every change is decoration, not a review boundary.The five-state acceptance map. Each failure produces a different explanation and next action.
State 1: no supporting passage exists
The source set may contain the same topic words without evidence for the claim. This unsupported state should be triggered by the absence of a passage that establishes the exact wording, not by a low model-confidence score. Withhold the affected claim, name the searched source set and offer nearby covered material or a content request. Never cite the nearest topical paragraph simply to fill the slot.
Acceptance check: ask the team to submit a question about a known absent topic. The prototype passes only if it avoids a plausible answer, identifies the searched scope and offers a usable next step. Record who receives the resulting content request.
State 2: the source is temporarily unavailable
An outage is different from absence. The passage may exist, but the source service cannot supply it now. Saying "not found" would turn an operational failure into a false content statement. Show source unavailable, name the source class and explain whether a retry is possible. If a saved answer remains visible, preserve the edition and retrieval time that supported it.
Acceptance check: interrupt the source connection during the demo. The user-facing state, retry behavior and operational event should agree. A polished toast with no trace for the support team fails the test.
State 3: the evidence exists, but this user cannot open it
Permission failures create their own risk. Confirming that restricted information exists can disclose something, while claiming it does not exist misleads the user. An access restricted state should explain that the evidence cannot be displayed in the current context, then offer an approved access process or authorized review route without leaking the passage.
Acceptance check: run the same claim with a reviewer account and a normal learner account. Verify that the authorized reviewer can inspect the passage, the learner receives the intended boundary and neither route changes the underlying review record.
State 4: a newer source edition invalidates the old review
A citation can remain reachable after its source has changed. Page locations move, wording changes and an earlier approval may no longer cover the displayed claim. In this source superseded state, keep the original context when permitted, label the newer edition and route dependent claims back to review. Silent replacement destroys the history needed to explain why a decision changed.
Acceptance check: swap the test corpus for a new edition. The prototype should invalidate the relevant review status without erasing the earlier record. A badge that remains approved after the source hash or edition changes is a release blocker.
State 5: valid sources disagree
A handbook, a study and a regulator's record answer different questions. When valid sources conflict, a sources differ state should keep each under its own class label and show the relevant passages side by side. A regulatory record can establish regulatory status. It does not rewrite what a study observed, and the study does not establish an approval.
Acceptance check: load a fixture in which two source classes reach different conclusions. The interface passes if the disagreement remains visible and the user can inspect both bases. It fails if the system chooses a winner without a declared rule.
Why reviewability belongs in the interface
The U.S. Food and Drug Administration's Clinical Decision Support Software guidance, issued January 29, 2026, uses the phrase "to independently review the basis for such recommendations". That criterion belongs to a specific US regulatory analysis. It is useful here as a design constraint, not as a claim that a prototype qualifies for any regulatory category.
Independent review requires more than a source list. The intended reviewer must reach the passage, see its edition and record a decision without privileged help. A tiny reference drawer, hover-only explanation or answer-wide approval button may make a formally visible source impractical to evaluate. Observe the reviewer completing the task instead of narrating the route.
Run the five states in twenty minutes
Use one synthetic or otherwise approved source fixture. Do not upload patient information to make a demo feel realistic. Freeze the visible claim, expected passage and interface revision before the session.
- Minutes 0-4: supported baseline. Open the passage behind one claim and record the source edition plus the observed wording.
- Minutes 4-8: remove support. Ask a known absent question and verify that the claim is withheld rather than patched with nearby text.
- Minutes 8-12: change access. Repeat the review with a restricted user and confirm that the boundary neither leaks evidence nor claims it is absent.
- Minutes 12-16: replace the edition. Introduce a changed source and confirm that the old review status expires for dependent claims.
- Minutes 16-20: create disagreement. Present two source classes with conflicting conclusions and check that both remain inspectable.
Keep an Evidence-State Decision Record for every step: scenario ID, exact claim, source class and edition, injected condition, expected behavior, observed behavior, owner and release decision. A screen recording is context, not a substitute.
Disqualifiers in a supplier demonstration
Several behaviors should stop acceptance until the team can explain and correct them:
- every failure becomes the same "I cannot answer" message;
- the presenter needs administrator access to reveal the evidence;
- an approved label survives a source or claim change;
- conflicting source classes are blended into one conclusion;
- the team fixes a failed example live without preserving the failed version;
- nobody owns the content, access or operational action created by the state.
These findings do not prove the entire product is unsafe. They show that the current prototype cannot support the narrow review task it appears to promise. Turn each failed state into a bounded delivery item with an owner and completion condition.
Use the result to scope the next build
The failed state identifies the missing work. Unsupported claims point to corpus coverage or claim boundaries. Unavailable sources expose operating behavior. Restricted evidence needs access design. Superseded sources require versioning and review invalidation. Conflicts require a visible provenance policy.
Use the full AI citation UX case with answer anatomy and honesty-state screens to compare this five-state acceptance model with a documented aesthetic-medicine design process. The case includes the source audit, the corrected citation method and the limits that remained open at the end of discovery.
Run the drill before approving a prototype or asking for a fixed implementation estimate. A failed state should end in a specific next task, not a broad promise that the AI will be accurate.
About the author
Dmytro Nasyrov is the founder and CTO of Pharos Production. He works with founders and product teams on healthcare AI, evidence-bound RAG systems and complex software delivery.
This article presents a product-design and procurement test. It does not provide medical advice, clinical validation or a legal determination about any healthcare software product
