What one score can’t tell you
1 July 20265 min readSillyMaths
A practice paper comes back marked: 62 out of 100.
Is that good? It depends what the paper was. Is it progress? It depends what last month’s number was, and whether the two papers are even comparable. Should you be worried? The number has no opinion. A score compresses a hundred separate decisions — a hundred small moments of reading, choosing and computing — into a single figure, and in the compression it throws away almost everything you would need in order to do something.
The uncomfortable truth is that 62 is not a finding. It is a summary of findings you were never shown.
Two children, one score
Here is the cleanest way to see what the number hides. Elliot and Zara, both ten, sit the same fractions paper: 45 questions in three parts, and each part asks the same things in a different way. Both children score 34.
Same paper, same day, same number. Now look one level down.
| Part | Elliot | Zara |
|---|---|---|
| 1 | 14/15 | 12/15 |
| 2 | 5/15 | 11/15 |
| 3 | 15/15 | 11/15 |
| Total | 34 | 34 |
These are not similar children. They are opposite children.
Elliot is close to flawless in two of the three parts — one mark lost in fifteen, then none at all — and in the third he loses two marks in every three. The same maths, dressed differently, and one particular dressing defeats him every time. Zara is the opposite shape: she drops marks everywhere, including on questions Elliot would call easy, and she gets no worse as the paper goes on. Whatever is going wrong for her, it is not the difficulty.
Zoom in once more
Part scores are better than one number, but the real information lives one level deeper — in which questions were missed, grouped by the mistake they were built to catch. That is what the design is for: every mistake on the paper is asked more than once, so a wrong answer is never on its own.
Elliot's eleven wrong answers are not eleven separate events. They cluster hard. Three of the questions asked him to take a fraction of what was left rather than of the original amount — the mistake we call working with what's left — and he got all three wrong. Three more ran backwards, giving the part and asking for the whole, and he missed all three of those too. Two wrong turns account for six of the eleven marks he lost. Elliot does not have a fractions problem. He has two specific reading problems that fractions questions happen to be very good at exposing.
Zara’s wrong answers are spread thinly across the whole paper, which looks like noise until you group them. Then one row stands out: every single question that asked her to add or subtract fractions with different denominators was missed, the same way each time — tops added to tops, bottoms to bottoms — the mistake we call adding tops and bottoms. Her equivalent-fractions work is shaky in the same mechanical way. But notice what her second and third parts are quietly saying: when a question’s set-up doesn’t depend on that one broken mechanic, her reasoning carries her through problems that defeated Elliot.
So the prescriptions are opposite. Elliot needs no more basics — he needs slow, deliberate work on multi-step reading, with one question asked out loud again and again: a third of what? Handing Zara harder papers would be the worst possible use of her time; she needs a quiet week rebuilding how fractions are added, and then the marks come back across all three parts at once.
The score said the same thing about both children. The pattern said opposite things.
Now imagine both families acting on the number alone. Both children get “more practice papers”. Elliot grinds through foundations he mastered a year ago and meets his two real problems only occasionally, by luck. Zara meets her broken mechanic constantly and reinforces it. Both scores drift. Both children conclude they are bad at fractions. Neither is.
The number on the day versus the pattern across months
There is a second thing the single score cannot do, and it may matter more: it cannot be compared with itself.
Say the March paper came back 61 and the May paper 63. Progress? Noise? A harder paper generously marked, an easier paper on a tired day? There is no way to tell, because a one-number summary carries no memory of what kind of marks were lost. Two points of movement is well inside the wobble of a ten-year-old’s ordinary week.
A named mistake, tracked across months, behaves completely differently. If you know Elliot has now met “of the rest” six times since March and has caught it the last three times, that is not noise — that is a wrong turn closing down, visible regardless of what the totals did. If instead he has met it six times and fallen for it six times, that is the single most useful sentence anyone can tell you about his preparation, and no pair of scores would ever have surfaced it. One test has luck in it. A pattern across tests does not.
What to look for in any practice material
None of this requires any particular product. It requires specificity, and you can audit any practice material — ours or anyone’s — against three questions:
- Does each question isolate one skill? If a question can be failed five different ways, a wrong answer tells you almost nothing. If it was built to catch one specific wrong turn, a wrong answer is a diagnosis.
- Does the marking record which, not just how many? A total forgets everything. Keep the wrong question numbers, and group them by what the question was testing — that grouping is where every insight in this article came from.
- Can two tests be compared? The categories need to be the same in April and June, or “is this improving?” stays unanswerable. A shared vocabulary across papers is what turns a pile of marked PDFs into a record.
For what it is worth, this is exactly how SillyMaths is put together: every question in every pack is tagged with the one mistake it exists to catch, and the same tags run through the whole library, so the pattern survives from one test to the next — the routine is on how it works.