Categories:
Research
healthcare-ai medical-ai diagnostics nhs evaluation

The Best Medical AI Doesn't Replace Doctors — It Reads the Tests They Already Ordered

Feature image for The Best Medical AI Doesn't Replace Doctors — It Reads the Tests They Already Ordered

The Best Medical AI Doesn’t Replace Doctors — It Reads the Tests They Already Ordered

When Imperial College London announced an AI that reads routine ECGs in under two seconds, the early headlines grabbed the shiniest numbers available: detection rates “up to 81% and 90%.” That framing survived about a week. The British Heart Foundation’s own September 2 update shows the fuller picture, and it’s messier: 77% and 81% for reduced pumping function across two patient cohorts, 80% and 90% for aortic stenosis. Still strong. Just not one clean number, and not the same number for everything.

The messy version is the more interesting story. This tool isn’t a diagnostic miracle and was never meant to be. It’s a triage layer that squeezes new signal out of a test clinicians already order millions of times a year, and that turns out to be where a lot of real medical AI value lives.

What the tool actually does

The model was trained on 10.6 million ECGs paired with clinical reports. Given a routine ECG, it flags signs of two conditions that normally require an echocardiogram to catch: reduced heart pumping function and aortic valve disease. It doesn’t confirm either one. It doesn’t rule anything out. It points at patients and says, in effect, “this one should get a closer look.”

That’s a modest job description, and it’s the whole point. The ECG is one of the cheapest, most common tests in medicine. If software can read existing traces and notice the fingerprints of hidden disease before symptoms land someone in A&E, you’ve added a screening capability at close to zero marginal cost. No new machines, no extra appointments, and nobody in the clinic has to change how they work.

Imperial is now testing the model with 590 NHS patients across London and Bristol, and routine use is optimistically two years out, pending approval.

The numbers, honestly

This is where the story earns some trust, and where most coverage lost it.

Early Guardian reporting said detection rates “up to 81% and 90%.” The British Heart Foundation update gives you the per-cohort reality: for reduced pumping function, 77% in a cohort of 5,442 patients and 81% in a cohort of 61,520. For aortic stenosis, 90% and 80%. The headlines picked the best number from each condition and let them stand side by side.

None of this means the tool is weak. It means the tool performs differently across conditions and populations, which is exactly what an honest evaluation looks like. The BHF also notes the model can’t confirm or exclude disease on its own. It raises a flag; the diagnosis still happens downstream.

I’d argue the per-cohort numbers make the tool more credible. A single uniform figure would be the suspicious outcome.

Why the boring middle wins

Healthcare AI stories tend to get written in one of two registers: miracle cure or doom. This one lives in the boring middle, which is where the compounding value usually hides.

Think about the economics. An echocardiogram costs money, takes time, and requires a specialist. An ECG is nearly free and already happening. The software replaces no one, but it changes what an existing test can tell you, and that changes who gets escalated. Hidden heart failure gets caught earlier and more cheaply. Multiply that across every routine ECG in the system and the aggregate effect is large, even though each individual decision is small.

This pattern isn’t unique to cardiology. It’s the same shape as the best automation wins everywhere: take a process that already runs and make it smarter.

How to read AI performance claims

The gap between “up to 90%” and the actual cohort data is a lesson that travels well beyond medicine. Next time you’re evaluating an AI tool, run three checks:

  1. Find the denominator. “90% detection” means nothing until you know the cohort size, who was in it, and how they were selected. 90% on 5,442 carefully chosen patients and 90% on everyone who walked into a clinic are different claims.
  2. Hunt for the per-slice numbers. If a vendor quotes one number, ask how it varies by condition or population. Uniform headline numbers are usually an average or a maximum wearing a costume.
  3. Ask what happens after detection. Detection is not diagnosis. A flag that routes patients to further testing has a completely different risk profile than a tool that closes the case. The Imperial team is explicit that their model triages. Vendors who blur that line are doing it on purpose.

None of this takes long. Most of the time it takes one follow-up question, and the answer tells you whether you’re looking at an evaluation or a press release.

What to do now

If you work in or near healthcare: watch this trial, because the triage-not-diagnosis pattern is the template you’ll see repeated across medical AI over the next few years. The tools that ship first and cause the least trouble will be the ones that read existing data and escalate, rather than the ones that issue verdicts.

If you build or buy AI tools anywhere: steal the evaluation discipline. Demand per-cohort numbers. Know whether you’re buying a detection layer or a decision layer. And treat any “up to X%” claim as an opening bid until the cohort data says otherwise.

A two-second ECG read makes for good copy. What lasts is the pattern behind it: cheap software, pointed at the pile of data everyone was already ignoring, handing the judgment to a human.


Source: Imperial College London research reported by The Rundown AI and the British Heart Foundation (September 2, 2026 update).

Related Articles