The agent on the telephone

Oxford NHS Dora

Three weeks after cataract surgery, thousands of NHS patients take a phone call from Dora, an autonomous clinical assistant built by Oxford spin-out Ufonia. It asks about the five symptoms that matter and recommends discharge or a clinician callback. It acts alone, inside a boundary drawn by clinicians and licensed by published evidence.

Worked case · public-source reconstruction Autonomy: execute · the follow-up call

Why this case matters

Autonomy was earned, not assumed.

The interesting choice is not giving an AI a phone line. It is the order of operations: a validation study in which an ophthalmologist supervised every call, peer-reviewed publication, a regulatory mark, and only then routine autonomy. Dora executes because the design proved it could, on the record, before it was allowed to act alone.

Why execute rather than prepare? Assign the rung to the action, not to the agent. Dora's consequential action here is the follow-up call itself: it places the call, holds the conversation, judges which symptoms matter and routes the outcome, with no person in the loop while it does so. That is execution. Confirming discharge is a separate action that sits lower, and it stays with a clinician. An agent that merely drafted the call for a person to make would sit at prepare; Dora makes the call. See the full autonomy ladder →

Validated
Peer-reviewed, every call clinician-supervised
Status
Default post-cataract follow-up at Oxford University Hospitals
Regulatory
UKCA-marked medical device
The rule
Any flag means clinician callback within 48 hours

The case

Full autonomy over five questions. None beyond them.

Cataract surgery is one of the most common operations in the world, and most follow-up appointments confirm what everyone expected: the patient is fine. Those appointments still consume clinic slots, and the few patients who are not fine must be caught reliably. That is the workflow Dora handles. About three weeks after an uncomplicated operation, it telephones the patient and holds a natural conversation about the five symptom domains that matter: redness, pain, vision, new floaters, and flashing lights.

A clear call ends in a discharge recommendation, which a clinician reviews and signs off. Any flagged symptom routes to a clinician callback within 48 hours. In the validation study, run with an ophthalmologist supervising every call, Dora caught 93.75 percent of the patients who needed review, agreed with the supervising clinicians symptom by symptom, and completed 96.5 percent of calls without human help. In routine deployment across two NHS trusts, covering 1,636 consultations, only 0.3 percent of patients judged clear later needed an unplanned change in care.

Oxford University Hospitals went on to make Dora the default follow-up pathway for thousands of cataract patients and extended it toward pre-operative assessment; a third trust reports more than 2,100 calls and hundreds of nursing hours returned to patients. The design is why the autonomy holds: the loop is five questions wide, two outcomes deep, and a clinician on every discharge.

Why it matters

The execute rung, made safe by subtraction.

Execute-level agents usually alarm people because their loop is open: they can decide, act, and keep going. Dora's loop is closed. It conducts a real clinical conversation end to end, but it cannot prescribe, cannot book theatre time, and cannot do anything with a worrying answer except hand the patient to a human, fast. The autonomy is real, and so is the subtraction that makes it safe.

Read through the Agent Operating Model: cognition is the conversation; control is the protocol, the 48-hour callback rule, and clinician sign-off on every discharge; reach is a telephone and a structured summary. Power is the product of the three, and here each one was chosen, published, and regulated before the agent was trusted alone with a patient.

The filled canvas

Turning the story into nine design decisions.

The entries below are intentionally compact. A canvas should not read like a requirements document; it should make the load-bearing choices visible enough for a team to argue about them.

Worked example

Oxford NHS Dora

Agentic AI Design Canvas
designtheagent.com

North Star

What business outcome should this agent's work ultimately contribute to?

Return scarce ophthalmology capacity to the patients who need it by safely discharging routine cataract patients after surgery.

Target Workflow & Agent Role

Which workflow is it part of, and what contribution is the agent responsible for making within it?

Workflow: follow-up about three weeks after uncomplicated cataract surgery. Role: conduct the routine post-operative review conversation and identify the patients who need a clinician to look again.

01

Users & Stakeholders

Who uses it directly, and who else is affected by what it does?

Direct users: post-operative patients on the telephone, and the ophthalmology team receiving flags and summaries. Affected: waiting-list patients who gain the freed capacity, and the trusts.

02

Success Measures & Standards

What measures will show that the workflow improved, and what performance standards must the agent meet?

Measures: sensitivity to patients needing review, and unplanned care after a clear call. Standards, published first: 93.75% sensitivity, 96.5% of calls unaided.

03

Autonomy by Action

For each consequential action in its role, how independently may the agent act?

Autonomy: execute. It conducts the call and classifies the case end to end, with nobody in the loop while it does so. Confirming discharge sits lower, at prepare: a clinician reviews the summary first.

04

Tools & Action Channels

Which tools, systems or channels may it use to get information, make changes, communicate or take action?

Telephony: it places real calls and converses in natural language, then writes a structured summary into the clinical record. It cannot prescribe, book theatre time, or act beyond routing a callback.

05

Rules & Boundaries

What must it always do, what must it never do, and when must it stop, escalate or hand off?

Only uncomplicated cataract cases. Any flagged symptom routes to a clinician callback within 48 hours. Emotionally complex situations are excluded; those belong to humans.

06

Context & Knowledge

What information may it use, and which sources should take priority when they disagree?

The cataract pathway and the five symptom domains that matter after surgery: redness, pain, vision, new floaters, and flashing lights. Restricted to uncomplicated surgeries by design.

07

Memory & Learning

What should it remember across interactions, what must it forget, and how should it learn and improve over time?

Each call becomes a structured record in the patient's pathway. Improvement is audited in public: a supervised validation study first, then a published evaluation of deployment.

08

Ownership & Oversight

Who owns the workflow outcome, who is accountable for how the agent operates, and how will it be reviewed?

Clinicians review every summary before discharge and call back every flagged patient; the trusts own the pathway; named investigators published the results that license the autonomy.

09
Job · what it's for Authority · what it may do Knowledge · what it may know Accountability · who answers

What to notice

Four lessons travel beyond the NHS.

Cell 04

Execute is safe when the loop is small.

Full autonomy over a five-question call; zero autonomy beyond it. The rung is high precisely because the task is narrow.

Cell 03 · the standard

The bar was set by evidence, not a demo.

Published sensitivity, clinician agreement, and a regulatory mark came before routine use. The standards were numbers on the record, not adjectives on a slide.

Cell 06

The hand-off is the design.

Any flagged symptom routes to a clinician, and emotionally complex situations are excluded on purpose. The boundary says what the agent must not try to handle.

Cell 03 · the measure

The measure is the freed clinic.

The value is counted in clinician time returned to the patients who need it, and in missed problems that stay near zero, not in calls handled.

Coach audit

We ran this canvas through the coach.

Not as an outside fact-checker, but as a clinical-operations lead preparing a governance review. Dora is the most autonomous agent in these cases and the most carefully bounded, and the coach still found a load-bearing gap in it. Here is what came back.

Coach review · round one

Load-bearing gaps

What holds together

  • The autonomy is licensed by evidence: published sensitivity and agreement bars came before routine use, not after.
  • The rung is split by action. Dora conducts and classifies the call itself; the discharge decision stays with a clinician who reviews every summary.
  • The scope is drawn tightly: uncomplicated cataract cases only, with emotionally complex situations excluded by design.

What pulls apart

  • Rules & Boundaries (Cell 06). Every flagged symptom is routed to a callback within 48 hours, but some post-cataract symptoms may be sight-threatening emergencies that cannot safely wait two days. Split the rule into urgency tiers.
  • Ownership & Oversight (Cell 09). Clinicians review summaries and the trust owns the pathway, but no one is named who can inspect Dora's calls, change its standards, or stop the system.
  • Context & Knowledge (Cell 07). The canvas does not say what Dora does when a patient raises a symptom outside the five named domains, or which source governs if the pathway guidance is updated.
  • Memory & Learning (Cell 08). The structured record that persists is described, but not whether the call audio is kept, for how long, or what is deleted.
  • The pre-operative extension. The “uncomplicated cataract only” boundary and the five post-op symptom domains do not describe pre-op risk, so reusing them would leave the new work ungoverned.
“Should some flags trigger urgent, same-day action rather than a 48-hour callback, and how would Dora tell them apart on the call?”The coach, on Rules & Boundaries (Cell 06)

Sources

What we can say publicly.

Only public-source claims are treated as facts. The performance figures come from the two peer-reviewed evaluations; operational details beyond them should be read as design implications unless the trusts or Ufonia confirm them.

More worked cases