Yes, But What Does It Actually Mean? | No. 7
Base rate neglect, the “99% accurate” test, and why a positive result often isn’t what you think
Here is a question that has caught out more doctors than anyone in the medical profession would like to admit.
There’s a disease that affects 1 in 1,000 people. There’s a test for it, and the test is 99% accurate. You take it. It comes back positive. How worried should you be – what are the odds you actually have the disease?
The intuitive answer, the one most people land on and quite a few clinicians have too, is somewhere around 99%. The test is 99% accurate, you tested positive, so surely you’re almost certainly ill.
The real answer is about 9%.
That is not a typo, and it isn’t a trick. It’s one of the most useful pieces of arithmetic you’ll ever do, and it comes apart the moment you stop thinking about percentages and start counting actual people.
Counting the people
Take 1,000 people, because that’s the number the disease is measured against.
One of them, on average, has the disease. The test is 99% accurate, so it correctly catches that person. So far, so reassuring.
Now look at the other 999, who are perfectly healthy. The test is 99% accurate for them too, which means it gets it wrong 1% of the time. One per cent of 999 is about 10 people. Ten healthy people who will be told, incorrectly, that they’ve tested positive.
So how many positive results are there in total? Eleven. Ten false, one real. And you are somewhere in that group of eleven, with no way from the result alone to know which kind you are. Your odds of being the genuine case are one in eleven. Around 9%.
The number feels wrong right up until you’ve counted the people, and then it’s simply undeniable.
What just happened
The thing almost everyone forgets is how rare the disease was to begin with. That figure – 1 in 1,000 – is the base rate, and it turns out to matter enormously. When something is genuinely rare, the small percentage of false positives among the very large healthy group can easily outnumber the true positives among the tiny sick one. The test being “99% accurate” doesn’t save you, because 1% of a big number is still bigger than 99% of a tiny one.
This is base rate neglect: fixing on the headline accuracy and quietly ignoring how common or rare the underlying thing actually is. And once you notice it, you realise that “99% accurate” and “a positive result means you probably have it” are not remotely the same statement. The first is about the test. The second depends entirely on the base rate, and the rarer the condition, the wider the gap between the two.
Why we fall for it
Part of the reason is that the human mind is genuinely bad at holding two different numbers in view at once. We’re handed the 99% and it’s vivid, concrete, and right in front of us. The base rate is abstract, easily forgotten, and often not mentioned at all. So we anchor on the number we were given and quietly drop the one we weren’t.
It doesn’t help that “99% accurate” is designed to sound reassuring. It’s the number that goes on the marketing. The base rate is the number nobody puts on the box, and it’s the one that decides what the result actually means.
Closer to home
Which brings us to a version of this now playing out in classrooms and exam boards across the country.
AI writing detectors promise to spot work generated by tools like ChatGPT, and they’ll happily quote an accuracy figure to prove it. Say a tool is 99% accurate at identifying AI-written text. An institution runs 10,000 student essays through it. Sounds like a sensible safeguard.
But apply the same counting exercise. Suppose, encouragingly, that the overwhelming majority of your students wrote their essays themselves, and only a small fraction used AI dishonestly. The honest work is now the big group, the cheating is the rare thing, and that is precisely the setup where false positives pile up. Even at 99% accuracy, running 10,000 honest-majority essays through the tool produces a substantial stack of students wrongly flagged for misconduct. Not because they did anything, but because they had the misfortune to be the 1%.
The difference from the medical case is that here, a false positive isn’t an unnecessary fright cleared up by a second test. It’s an accusation of cheating against someone who didn’t. And a great many institutions are leaning on exactly these tools without ever doing the sum that tells them how many of the flags will be false.

What to ask instead
You don’t need to run the numbers every time. You mostly need to remember that the base rate exists, and a few questions keep you honest:
How common is the thing actually being tested for? Before you trust a positive result, find out how rare or frequent the underlying condition is. That number changes everything.
Accurate compared to what? “99% accurate” is close to meaningless until you know how it behaves on the rare cases specifically, not just on average.
If I picture 1,000 cases, how many flags are real? Counting real people, as we did above, cuts through the intuition far better than staring at a percentage.
And the one that matters most when there are consequences attached: what happens to the person on the wrong end of a false positive? Because when the answer is “they get accused of something they didn’t do”, the base rate stops being a curiosity and becomes a duty of care.
The point
Base rate neglect isn’t about tests being bad. A 99% accurate test is a genuinely good test. It’s about the fact that a good test, pointed at a rare enough thing, can still produce more wrong answers than right ones, and that the accuracy figure alone will never tell you when.
So the next time you’re handed an impressive-sounding percentage and asked to act on a positive result, ask how rare the thing was in the first place. Because yes, the test might be 99% accurate. But what does it actually mean?
Read more in this series here
