Why I'm Not Ready to Let AI Interpret an Ultrasound (Yet)
Aug 27, 2026
AI in veterinary ultrasound is developing quickly, but how much should we actually trust it to interpret what we're seeing? As AI becomes more common in veterinary medicine, I think we need to be clear about where it genuinely helps and where a knowledgeable human still needs to stay firmly in the loop.
You've probably seen the story doing the rounds: an Oregon veterinary hospital is suing an AI diagnostics company, alleging that its tool misread a tissue sample from an 11-year-old dog as inflammatory when it was actually cancerous and that the dog died after a second, more invasive surgery that the lawsuit says wouldn't have been needed if the AI had got it right the first time.
The company had marketed the tool as "the world's most capable veterinary AI analyzer." Oregon vet sues AI company saying misdiagnosis of cancer tumor led to 11-year-old dog’s death | The Independent
It's a genuinely awful story, and I understand why it's spreading fast through vet circles. But I don't think the right lesson from it is "AI in veterinary medicine is dangerous, avoid it." I think the right lesson is narrower, and more useful, than that.
What You'll Get in This Post
- Why "AI vs human" is the wrong comparison to be making about this case
- What I actually think AI is good for in a veterinary business — with a real example
- Why FOVU Report still isn't AI, and why that's deliberate
- What I've learned trying AI tools for image interpretation myself
- Why a slightly hedged ultrasound note is a feature, not a weakness — and where under-confidence actually costs you
No Test Is 100% Sensitive and 100% Specific
Here's the thing that's been missing from most of the commentary I've seen on this case: no diagnostic test — human or AI — is ever 100% sensitive and 100% specific. Cytology read by an experienced pathologist can miss things or over-call things. So can a vet at a microscope on a busy afternoon. So can an algorithm. It’s also dependent on the sample they are presented.
What we don't actually know from this single case is whether the AI tool's sensitivity and specificity for this type of sample is worse than an experienced human's would have been, better, or roughly the same with a different failure pattern. One bad outcome doesn't tell you that and it's worth being honest about that, rather than reaching for the most emotionally satisfying conclusion.
That's not a defence of the company. If they marketed a tool as more capable than it was and didn't disclose known limitations, that's a serious problem, and a completely different one from "AI got a tricky read wrong." But it does mean the useful question isn't "AI or human" it's "how do we actually find out which does what, and where, and how well, before we hand over the decision?"
That question is harder to answer than it should be, because AI companies are often not very transparent about how their tools actually work, what they were trained and validated on, or whether anything has changed under the hood between one version and the next. If a model gets quietly updated, retrained, or "improved" behind an API, how would you even know whether the tool that misread a sample in October is the same tool giving you an answer today? That opacity is part of what makes proper testing — the kind we'd expect for any other diagnostic tool or treatment — so difficult to insist on.
The Real Question: Has Anyone Actually Tested This?
In human medicine, before a diagnostic tool goes anywhere near a patient, it goes through trials. Sensitivity, specificity, comparison against the existing gold standard, in defined populations, with published results. We don't take a manufacturer's word for how well something performs.
Veterinary AI, as a field, is nowhere near that standard yet. Tools are being marketed on capability claims that most of us have no way to independently verify. That's the actual problem this case points to, not that AI touched a diagnosis, but that it may have been deployed with less scrutiny than we'd ever accept for a new drug or a new piece of diagnostic kit.
So I don't think the answer is "no AI." I think the answer is: treat every AI tool the way you'd treat any other diagnostic aid. Find out what it's actually been tested against, apply it to a specific, narrow task, and check whether it outperforms what you were already doing before you trust it with anything that changes a treatment plan.
How I Actually Use AI at FOVU
I'm not anti-AI. We use it at FOVU but in a specific, bounded way, and only after checking it does the job better than what came before.
FOVU.ai How FOVU Club is Transforming Learning with AI! is a good example. FOVU Club has years of case recordings, teaching videos and Q&A sessions in it — a genuinely huge archive.
Finding the right five-minute clip on, say, "identifying a small volume of free fluid" used to mean scrolling through hours of content or guessing at search terms. FOVU.ai answers the question and links you straight to the exact point in the video where it's demonstrated so you can check the source for yourself rather than just trusting the answer.
That's a narrow, well-defined task: search and retrieval with a citation, not diagnosis. It's a clear improvement on a generic search bar, and it's easy to verify when it's right.
FOVU Report doesn't currently use AI at all — and I know a lot of people assume it does. It's a teaching tool that walks you through a fixed structure, built on SPEEDS™, designed to make sure you think through and document a scan properly every time. We are looking at AI for very specific, limited jobs within it. Tidying up repetitive phrasing that the algorithm naturally produces, smoothing paragraph flow. Not interpretation. Not deciding what a finding means. That line matters to me, and I'm not moving it until I have a much better reason to.
What I've Learned Trying AI Interpretation Tools Myself
I've tested some of the AI tools available for ultrasound image interpretation. My honest experience so far: they're unreliable unless I already know enough to guide them heavily and sense-check what comes back. Which rather defeats the point for anyone who doesn't already have that knowledge.
That worries me more than the headline of the Oregon case does. If a vet who's still building confidence starts leaning on an interpretation tool that sounds authoritative but hasn't been through proper reliability testing, you get a specific and dangerous failure mode: overconfident diagnoses with no human backup checking them. Not because AI is inherently reckless, but because a fluent, confident-sounding output is very easy to mistake for a correct one — especially if you don't yet have the experience to spot when it's wrong.
The Value of a Slightly Uncertain Note
When I was building FOVU Report, one thing came up again and again in conversations with vets: they were keen not to sound overconfident in their ultrasound notes. Phrases like "no definitive mass identified, though repeat imaging may be warranted if clinical signs persist" rather than a flat, confident "normal."
I think that instinct is a genuinely good one. A note that honestly reflects uncertainty helps whoever reads it next — a colleague, a specialist, the same vet in six months. They understand exactly what was and wasn't seen, and how much weight to put on it. That's human, and it's valuable, and it's not something I'd want an AI tool to smooth away in the name of sounding more polished or more confident.
But there's a flip side to that same lack of confidence, and it's one I see a lot: sometimes it doesn't just soften how something is written, it means the thing doesn't get written down at all. A structure that wasn't assessed with full confidence sometimes just gets left out of the report entirely, for fear of putting something down that turns out to be wrong. That's a real loss — the next reader doesn't even know it was looked at. This is one of the things FOVU Report is specifically built to help with: it walks you through the process so that everything you assessed gets included, with the right level of confidence attached to it, rather than quietly dropped because you weren't sure.
AI Is a Tool, Not the Answer
Walking round the London Vet Show last year, I remember being genuinely taken aback — overwhelmed isn't quite the right word, but it's close — by how many companies had "AI" in their business name. Not "AI-powered scanner" or "AI-assisted triage," just AI, as if the technology itself were the product and the value.
I think that's backwards, and it's part of why cases like this Oregon one happen. AI is a tool. It's not the answer to everything on its own. Point it at a specific, well-defined task and test whether it does that task better than what you had before, and you can get something genuinely useful — that's FOVU.ai finding the right video clip. Treat it as a broad solution to a broad problem — "AI will diagnose this," "AI will run this business" — and you tend to get a broad, unreliable answer dressed up in confident language.
Where I Land on This
AI isn't a blanket cure-all, and it shouldn't be treated like one by companies marketing it or by vets adopting it. But the answer isn't to avoid it either. It's to test it properly, apply it to specific tasks it's actually good at, and keep the line clear between "this helps me find or organise information" and "this is telling me what to do about a patient."
The other half of this, for me, is protecting the human side of the equation. If we let AI quietly absorb the interpretive skill of scanning, we lose something that takes years, mistakes, and proper education to build and that's not something you can get back quickly if the tool turns out to be wrong, or unavailable, or you're the vet on call with no signal and a dog that can't wait.
A quick disclosure, since it's relevant to everything above: I used AI in a specific, bounded way to help me write this article more quickly. The ideas, the argument and the edits are mine, but I used it as a drafting tool once I'd worked out what I wanted to say. That's actually a fair example of how I think AI should be used in a business like mine: for a defined task, started with a human idea, and checked and edited by a human before it goes anywhere near you.
Takeaway
No test is perfect, human or AI, and this case doesn't tell us which was more at fault but it does tell us that veterinary AI needs the kind of rigorous, transparent testing we'd expect for any other diagnostic tool. Use AI where it's been shown to help with a specific task. Keep a knowledgeable human in the loop for anything that shapes a diagnosis. And don't be afraid of a note that admits some uncertainty — it's doing its job.
_____________________________________________________________________________________________________
Author bio: Dr. Camilla Edwards (DVM, CertAVP, MRCVS) is a peripatetic veterinary ultrasonographer and founder of FOVU. She helps first-opinion vets build confidence in scanning and reporting so they can deliver better care with less stress.