The Average Is Nobody's Result
DEV Community

The Average Is Nobody's Result

In 2013 four researchers went back to a completed mammography study and asked it a question it had not been designed to answer. The original study had put 50 radiologists in front of 180 mammograms, twice. Once unaided, once with computer-aided detection marking suspicious regions. The finding was a null. On average, computer aid changed nothing measurable, and the profession moved on. Andrey Povyakalo and his colleagues at City University London reanalysed the data by splitting the readers instead of pooling them. What they found was that the tool had done two large things at once, in opposite directions. For the 44 least discriminating radiologists, on 45 relatively easy cancers, computer aid was associated with "a 0.016 increase in sensitivity (95% confidence interval [CI], 0.003-0.028)". For the 6 most discriminating radiologists, on the 15 hardest cancers, "with CAD, sensitivity decreased by 0.145 (95% CI, 0.034-0.257)". The weakest readers got slightly better at the cases that were already easy. The best readers got substantially worse at the cases that were hard, which is to say the cases where a radiologist is the only thing standing between a patient and a missed cancer. Averaged together, those two effects cancelled, and the study reported that nothing had happened. Their own summary of it: "despite the original study detecting no significant average effect, CAD helped the less discriminating readers but hindered the more discriminating readers." The average was true of neither group. Averages behave like this everywhere, including on the dashboard you looked at this morning. Take the last tool your organisation rolled out and measured. Somebody produced a figure: review throughput up nine per cent, tickets resolved up fourteen, defect-escape rate down a fifth. That figure was computed across everybody who touched the work, which makes it a mixture. There were as many different effects in it as there were people, weighted by how much work each of them happened to do that quarter. A mixture behaves in ways an effect does not. It can be positive while the effect on a third of your team is negative. It can be zero while two large things are happening. And because the weights are your staffing, the mixture is a property of who was on shift as much as of the software. Hire four juniors and your measured number moves without anything about the tool changing, which is also why the vendor's benchmark never reproduces in your shop. Say what that nine per cent still is, though, because the argument is easy to overshoot. It is a real answer to a real question: across the actual mixture of people and work you had last quarter, this is what happened. That is worth knowing, and it may well be enough to justify keeping the tool. What it cannot tell you is how to deploy it. Suppose the nine per cent is three senior engineers on migrations and nine juniors on routine diffs. It is equally consistent with the tool lifting the juniors and doing nothing for the seniors, with the reverse, and with a gain for everybody. Those three worlds ask different things of you: which group to train, which to watch, and whether the number survives your next dozen hires. One number is compatible with all three and cannot tell them apart. I have made a version of this argument before, structurally, in Where Your Metrics Fold: a scalar reading is a lossy projection, and two situations that demand opposite decisions can share one perfectly accurate number. This piece is the empirical half. Here is a documented case where the projection folded, and here is how often anybody bothers to check. Bound the 2013 case honestly before it carries any weight. It is a post-hoc reanalysis: the strata were cut using thresholds derived from the same regression that produced the estimates, the authors call their own method "exploratory", and the headline decline rests on six radiologists. The contrast is also conditional on two things at once, reader ability and case difficulty, rather than being a simple comparison between stronger and weaker readers. One case establishes one thing, and it is enough: an average can conceal two opposite effects, undetected, in a published null result, in exactly the class of tool everyone is now buying. Whether it usually does, nobody can say. So how often does anyone look? I could not find a count, so I made one. I chose a field unusually favourable to finding operator-level analysis. Adenoma detection rate in colonoscopy, ADR in the literature below, is a hard, standardised, patient-relevant endpoint, and computer-aided detection has been trialled against it more heavily than anywhere else in medicine. If a field splits its results by operator anywhere, it splits them here. The corpus is a frozen PubMed query: 259 records, 255 with abstracts. The question asked of each one was narrow and fixed before any reading began. Does this abstract report the assistant's effect separately for two or more groups of operators, defined by something about the operator, such as baseline detection rate, experience, training year, sex, or individual identity? Describing your cohort as experienced does not count. Splitting by patient age does not count. Twenty-one of the 255 report the effect split by the operator. The other 234 abstracts report an average over operators and no operator-level effect. Counting only primary studies, by PubMed's own publication-type labels rather than my judgement, it is 21 of 157. Twenty-one studies did the split. If the answer were obvious, they would agree. Across the 21, seven report the larger effect in the weaker operators, four in the stronger, three find no interaction, three find nothing that survives stratification, and one finds the two groups diverging over time. Both directions appear in randomised trials, on the same endpoint, in the same procedure. The honest reading of that spread is that the literature has no stable answer about which operators benefit most. A reader who files it away as "the effect is probably small" has reached a different conclusion, and the two license different decisions. This is the three-worlds problem from a few hundred words ago, except now it is not hypothetical: one of the most heavily trialled areas of AI assistance in medicine contains randomised trials pointing in opposite directions about which operators benefit, and it has not resolved them. An average is an honest description of a mixture. It just cannot tell you the thing that decides your next move, which is whether the tool is doing the same thing to everyone. The number you have been reporting is a summary of who happened to be on shift. Whether it describes any of them is a question you have not asked. If you go and run the split, I would like to know what came back. Estimates pointing the same way across your groups, estimates pointing in opposite directions, or cells too thin to read: all three are results, and at the moment none of them is written down anywhere. The colonoscopy literature has 21 attempts at this question and no answer. Yours would be the twenty-second, and it would be about a system you actually control. Which number are you judged by that you have never once split by the person who produced it? Top comments (0)

Comments

No comments yet. Start the discussion.