Explainability Bench
A learning signal is evidence to be interrogated, never an instruction. Work the four questions before you respond to anything on this platform.
Interrogate the signal
Worked example uses a synthetic acute-care record. Open each question to see what to inspect.
“Pattern consistent with clinical deterioration over the next 12 hours in this synthetic record.” It is a statement about a pattern, not about a diagnosis or a required action.
- Output type: ranked learning signal, not a validated score
- Time horizon the signal was framed around
- Population the training example described
Reading a pattern statement as a clinical conclusion.
Where meaning is made
The signal is not the intervention. The nurse's assessment and action are the intervention.
- Step 1 — Review contextTimeline
Synthetic vitals, symptoms, labs, medications, notes and treatment history assembled into one reviewable story.
- Step 2 — Learning signalPrioritised, never automatic
A transparent rule fires and shows contributing factors in context. Nothing is closed, ordered or actioned by the system.
- Step 3 — Explainability panelInputs, gaps, limits
Horizon, contributing signals, missing data, known limitations and links to institutional guidance.
- Step 4 — Nurse actionMeaning is made here
Validate, assess, reassess, communicate, escalate per local policy and document clinical reasoning.
Nurse authority
Same discipline, different context
Each setting states its own boundary. None of these are decision-support claims.
Deterioration pattern recognition with reviewable worklists to support timely escalation.
A single screening score is never sufficient on its own.
Symptom-burden trends across long, complex longitudinal records.
Never infer an adverse drug event or recommend a medication change.
Outreach priorities and missed-care patterns to support equity-focused prevention.
Social-determinant variables require documented equity review before use.
High-risk transitions and medication complexity flagged for education and follow-up.
Flags support teaching and handover only, not discharge decisions.
What goes wrong
Performance gaps by race, ethnicity, age, language or geography that overall accuracy hides.
Control: Measured subgroup performance, published in the Transparency Passport.
The plausible suggestion is accepted because it is on screen, not because it fits this person.
Control: Mandatory rationale, and disagree / need-more-data as first-class choices.
Volume erodes attention until real signals are dismissed with everything else.
Control: No alarm styling, no sounds, ranked review lists instead of interruptions.
Populations and care patterns move; an artifact quietly stops describing them.
Control: Scheduled re-review dates and append-only change history per registry record.
Fairness is measured, not assumed
Overall accuracy tells you nothing about how an artifact behaves for a subgroup. Where subgroup performance has not been measured, this platform says so rather than implying it is fine.