Narcissus knew one face perfectly, and no other face at all
The seer Tiresias was asked, when Narcissus was born, whether the boy would live a long life. He gave the strangest prophecy in Greek myth: yes — if he never knows himself. Years later Narcissus knelt at a still pool, met his own reflection, and could not leave. He studied that one face until he knew every eyelash. The world kept walking past behind him, and he never turned around again.
A machine can fail in exactly this shape. Let it study its own pile of examples too hard, for too long, with too much room to spare, and it stops learning the world and starts memorising the reflection. On everything it studied it becomes flawless. On everything else it is lost — and the terrible part is that its scores will not tell you, because it grades itself at the pool.
Written for ages 12+, and for any adult who has heard a machine is 99% accurate and wondered why it still gets their name wrong. Nothing here assumes you know anything. The real engineering words are all at the bottom.
the claim
A perfect score can mean nothing at all
Here is the least intuitive fact in this whole series, and everything else in this piece hangs on it: a machine can answer every question it studied without having learned anything. Not most questions. Every question. Word for word, fleck for fleck, flawless — and still be helpless in front of the first question it has not met.
You already know the human version. There are two ways to pass a maths test: learn how the sums work, or get hold of the answer key and memorise it. From the outside, on test day, the two students look identical. The difference between them is invisible — right up until somebody changes the questions.
Machines are spectacular memorisers. The pile of dials you have met in every piece of this series can, given enough dials, store the whole answer key without ever noticing there was a rule. So the real question in machine learning is never did it get them right? It is which student is this? — and there is exactly one honest way to find out.
step one
Two students, one answer key
Take twelve flashcards. Each has a number on the front and a number on the back, and the secret is boring on purpose: the back is the front, times three. Two of the cards were written in a hurry and have a small mistake on them — a seven that should have been a six. Keep that detail; it matters more than anything else on this page.
Now let two students study. One of them looks for the pattern, finds times three, and stops. The other memorises all twelve backs, mistakes included, because to a memoriser a mistake does not look like a mistake. It looks like one more fact.
Quiz them on the same twelve cards and something upside-down happens: the memoriser wins. Perfect score — the smudged cards too, reproduced exactly, wrong answer and all. The learner misses those two, because the learner answers with the rule and the card disagrees with the rule. Studying the pool harder beats understanding it, as long as the questions come from the pool.
quiz both students. then change the questions
Notice what the memoriser's perfect score was actually measuring: how well the quiz matched the pool. Nothing else. Hold onto the smudged cards, too — a machine that reproduces its pile's mistakes with total confidence has not learned the subject. It has learned the pile, mistakes included.
step two
Room enough to memorise
Why would a machine ever pick the memoriser's strategy? It does not pick. It slides — and what it slides on is room.
Training, remember, is nudging dials until the answers come out right. A machine with a few dials is cramped. It cannot possibly store every example, so its only way to score well is to find something short that explains a lot: a rule. Give it thousands of dials, though, and a second strategy quietly opens up — there is now enough room to bend itself around every single example individually, the true ones and the smudged ones alike, without ever finding the pattern that would have explained them all at once.
You can watch this happen with a curve. Scatter a handful of measurements that follow some real, smooth pattern — plus the little accidents every real measurement carries. Fit them with a cramped machine and you get the smooth shape underneath. Fit them with a roomy one and you get a curve that visits every measurement personally, wiggling violently in between — a portrait of the accidents, not the pattern. It scores better than the honest curve on the points it studied. On fresh points, it is a disaster.
give it more dials and watch the rule dissolve into memory
Do not take away that big machines are bad — the machines behind modern speech and language are enormous, and need to be. Take away that room converts studying into memorising unless something pushes back. The rest of this piece is about the pushing back.
step three
Studying past the point of learning
Room is one road to the pool. Time is the other — and this one you can watch happen live, which makes it the most useful picture in practical machine learning.
Before training starts, split the flashcards. Most go in the study pile. A handful go in a drawer, and the machine never, ever sees them — they exist so that you can ask it fresh questions whenever you like. Now train, and after every pass through the study pile, score it twice: once on cards it is studying, once on the drawer.
Two curves appear. Early on they fall together — the machine is genuinely learning, and real learning helps on the drawer nearly as much as on the pile, because what it is picking up is the rule. The drawer sits a little above the pile from the very first reading and always will; that small standing gap is not the problem. Then, somewhere, the curves separate. The study score keeps improving, pass after pass, looking like progress. The drawer score stalls — then turns and gets worse. That turn is the moment learning ran out and memorising began: everything gained after it is detail about the reflection, purchased by forgetting the world.
keep studying, and watch for the moment the curves part
Engineers run this split on every serious training run, watching the drawer curve the way a pilot watches fuel — the drawer scores arrive every half hour, and nobody looks at anything else first. The two curves in this instrument are not one of those runs: they are drawn from a formula, so the shape is legible and identical every time the page opens. The shape is the honest part. A real run has a wobble these do not.
step four
The exam that leaked
By now the fix sounds simple: keep an exam the machine has never seen, and believe only that. It is the right fix. It has one failure mode, and the failure mode is the quietest scandal in modern machine learning: the exam leaks.
Nobody has to cheat on purpose. The piles these machines study are scooped from the internet by the millions, and the internet also contains the exams — the famous test questions, the benchmark sentences, the very recordings a field agreed to grade itself on. Scoop wide enough and some of the exam ends up in the study pile by accident. The machine memorises it along with everything else, the way it memorised the smudged cards: not knowing it is an exam, only knowing it has seen it before.
Then the score comes back, and it is wonderful, and it is not a lie so much as a wrong answer to a different question. It measured how much of the exam was at the pool. Every leaked question gets answered by reflection; only the clean remainder measures learning. The dishonest part is that the report does not say which was which — one shiny number, exactly like the innkeeper's ledger in the last piece, and inflated for exactly the same reason: the machine was graded on its own guests.
leak a little of the exam into the study pile and watch the score detach from the truth
This is not hypothetical for us either. This very month, mid-training, we found study cards sitting in one of our own exam drawers — nearly the whole English exam had leaked. We stopped the run the same day, rebuilt the drawer from cards the machine had provably never seen, and added a gate that checks every future pile against every exam before a single dial moves. The score got worse and became true. That trade is the job.
step five
What actually cures it
Overfitting has real cures, and after four instruments you can see why each one works. They all do the same thing from a different side: they make memorising a worse strategy than learning.
Give it more cards, and more varied ones — memorising a hundred cards is easy, memorising ten million is expensive, and long before the pile is exhausted the cheapest way to score well becomes finding the rule. Give it less room — a cramped machine cannot store the pile, so it is forced to summarise. And stop at the turn — hand the finished job to the version of the machine from the drawer curve's lowest moment, before the memorising hours were spent.
Notice the last piece's lesson running the other way. Procrustes' bed could not be fixed by more data, because more of the same road was more of the same lie. Narcissus can be — more varied cards genuinely cure memorising. Two diseases, opposite prescriptions, and telling them apart is much of an engineer's actual week: one is the pile is wrong for the world, the other is the machine is too devoted to the pile.
a machine that has memorised, and three honest medicines. try them in any combination
There are subtler medicines in the working engineer's cabinet — nudging the dials toward smallness, blurring the cards a little so no single fleck is worth memorising, asking one machine to check another. Every one of them cashes out the same way: memorising gets more expensive, learning gets comparatively cheap.
the trouble
A whole field can kneel at one pool
One more turn of the screw, because the pool does not only catch machines. It catches the people building them.
A benchmark — a public exam a whole field agrees to compete on — starts life as a drawer: clean, held out, trusted. But every published result is a glance at it. Thousands of researchers, tuning thousands of machines, each keeping whichever version scored best on that exam — nobody memorising anything, everybody choosing by the reflection. Year by year the numbers climb, and some of the climb is learning, and some of it is the field as a whole quietly fitting itself to one specific set of questions. The exam wears out. It stops measuring the world and starts measuring devotion to itself — and the only remedy is the one you already know, at a larger size: somebody has to build a fresh drawer.
Tiresias' prophecy turns out to be an engineering specification. He will live long if he never knows himself. A machine — or a field — stays honest exactly as long as there is a question it has never met. The moment everything has been seen, scores stop meaning anything, and nobody can tell the learner from the memoriser ever again. Keeping something unseen is not a chore on the checklist. It is the whole mechanism by which the truth stays measurable.
the last thing
Echo loved him. Of course she did.
The myth has one more gift, and it is the pairing. Echo — the same Echo from piece two, cursed so she could only repeat what she had just heard — loved Narcissus and could not tell him. He heard only his own words come back. She could only repeat; he could only recognise. Two ways of seeming to know things, and neither is knowing: giving back what you were given, and recognising what you have already seen.
Everything this series calls learning lives in the narrow country between those two failures — answering a question nobody in your pile ever asked, about a face you have never met, and being right anyway. That is the entire miracle these machines occasionally manage, and it is why the honest test is never how well do you know your pile? The pool answers that question perfectly, forever, and it means nothing.
So carry the one habit this piece was built to give you. Next time anybody — a company, a paper, an excited adult — tells you a machine scored some marvellous number, ask the only question Narcissus could not survive: had it seen the questions before? If nobody can prove it had not, you have learned the score measures the reflection. And you, unlike him, get to stand up and turn around.
the glossary
The words the engineers use
| the pool | the training set |
| studying the reflection instead of the world | overfitting |
| memorising the answer key | memorisation — fitting the training data instead of the pattern |
| learning the rule ×3 | generalisation |
| the smudged cards | label noise — and memorising them is fitting noise |
| missing the smudged cards on purpose | a well-fit model disagrees with noisy labels |
| room to bend around every example | model capacity (parameter count is only a rough proxy for it) |
| the wiggle through every point | a high-variance fit — interpolating the noise |
| the drawer of cards it never sees | the held-out validation and test sets |
| the two curves | training loss versus validation loss |
| the moment the curves part | the onset of overfitting |
| stopping at the teal curve's lowest point | early stopping — keeping the best checkpoint, not the last |
| exam questions found at the pool | data leakage — train/test contamination |
| the score it reports vs the truth | inflated benchmark performance |
| checking every pile against every exam | contamination gates — deduplicating training data against evaluations |
| ten times the cards | more, and more varied, training data |
| a quarter of the room | reducing capacity; regularisation generally |
| blurring the cards a little | data augmentation, dropout — making single examples not worth memorising |
| a whole field kneeling at one pool | benchmark overfitting — a community tuning to a public test set |
| building a fresh drawer | new held-out benchmarks |
| Tiresias' condition — never knows himself | scores only mean anything on data the model has never seen |
| Echo repeating, Narcissus recognising | parroting and memorising — the two impostors of learning |
One sentence to take away
Perfect on everything it studied proves only that it studied; the single question that measures learning is one the machine has never seen — so guard a drawer of those, stop training when that drawer says stop, and trust no score that cannot prove the exam stayed dry.
Every honest result in this field survived that sentence. Every scandal in it, somewhere underneath, is a machine — or a field — that was graded at its own pool.
the sources
Where these ideas come from
Tiresias' warning has a technical literature. Each note says what the paper actually established.
- L. Prechelt, “Early Stopping — But When?”, in Neural Networks: Tricks of the Trade (1998). The oldest honest fix, studied properly across 1,296 training runs: watch a held-out drawer, and stop when it says stop — not when the training score looks flattering.
- C. Zhang, S. Bengio, M. Hardt, B. Recht & O. Vinyals, “Understanding deep learning requires rethinking generalization”, ICLR (2017). The unsettling experiment: standard networks can perfectly memorise completely random labels — so the fact that they usually generalise cannot be explained by the textbook story, and memorising is always within reach.
- M. Belkin, D. Hsu, S. Ma & S. Mandal, “Reconciling modern machine-learning practice and the classical bias–variance trade-off”, PNAS (2019). Double descent: past the point of exactly fitting the pile, test error can start falling again — the U-shaped curve you learn first turns out to be only the left half of the picture.
- P. Nakkiran et al., “Deep Double Descent”, ICLR (2020). The same shape found everywhere in modern practice — across model size, training time, even data amount — including regimes where more data briefly makes things worse.
- C. Xu, S. Guan, D. Greene & M.-T. Kechadi, “Benchmark Data Contamination of Large Language Models: A Survey” (2024). The leak, at internet scale: how public exams seep into training piles, how researchers try to detect it, and what it does to reported scores — Narcissus grading himself at the pool, as a research field.