Narcissus knew one face perfectly, and no other face at all

The seer Tiresias was asked, when Narcissus was born, whether the boy would live a long life. He gave the strangest prophecy in Greek myth: yes — if he never knows himself. Years later Narcissus knelt at a still pool, met his own reflection, and could not leave. He studied that one face until he knew every eyelash. The world kept walking past behind him, and he never turned around again.

A machine can fail in exactly this shape. Let it study its own pile of examples too hard, for too long, with too much room to spare, and it stops learning the world and starts memorising the reflection. On everything it studied it becomes flawless. On everything else it is lost — and the terrible part is that its scores will not tell you, because it grades itself at the pool.

Written for ages 12+, and for any adult who has heard a machine is 99% accurate and wondered why it still gets their name wrong. Nothing here assumes you know anything. The real engineering words are all at the bottom.

The pool perfectly still
one face, studied forever · the world, walking past unrecognised

the claim

A perfect score can mean nothing at all

Here is the least intuitive fact in this whole series, and everything else in this piece hangs on it: a machine can answer every question it studied without having learned anything. Not most questions. Every question. Word for word, fleck for fleck, flawless — and still be helpless in front of the first question it has not met.

You already know the human version. There are two ways to pass a maths test: learn how the sums work, or get hold of the answer key and memorise it. From the outside, on test day, the two students look identical. The difference between them is invisible — right up until somebody changes the questions.

Machines are spectacular memorisers. The pile of dials you have met in every piece of this series can, given enough dials, store the whole answer key without ever noticing there was a rule. So the real question in machine learning is never did it get them right? It is which student is this? — and there is exactly one honest way to find out.

step one

Two students, one answer key

Take twelve flashcards. Each has a number on the front and a number on the back, and the secret is boring on purpose: the back is the front, times three. Two of the cards were written in a hurry and have a small mistake on them — a seven that should have been a six. Keep that detail; it matters more than anything else on this page.

Now let two students study. One of them looks for the pattern, finds times three, and stops. The other memorises all twelve backs, mistakes included, because to a memoriser a mistake does not look like a mistake. It looks like one more fact.

Quiz them on the same twelve cards and something upside-down happens: the memoriser wins. Perfect score — the smudged cards too, reproduced exactly, wrong answer and all. The learner misses those two, because the learner answers with the rule and the card disagrees with the rule. Studying the pool harder beats understanding it, as long as the questions come from the pool.

quiz both students. then change the questions

Instrument 01 · The answer key both students ready
front → back · the rule is ×3 · two cards are smudged
the memoriser the rule-learner
On the study cards the memoriser is perfect and the learner drops the two smudged ones — being wrong about the smudges is what being right about the rule looks like. On new cards the memoriser has nothing to look up, and answers from the nearest card it remembers.

Notice what the memoriser's perfect score was actually measuring: how well the quiz matched the pool. Nothing else. Hold onto the smudged cards, too — a machine that reproduces its pile's mistakes with total confidence has not learned the subject. It has learned the pile, mistakes included.

step two

Room enough to memorise

Why would a machine ever pick the memoriser's strategy? It does not pick. It slides — and what it slides on is room.

Training, remember, is nudging dials until the answers come out right. A machine with a few dials is cramped. It cannot possibly store every example, so its only way to score well is to find something short that explains a lot: a rule. Give it thousands of dials, though, and a second strategy quietly opens up — there is now enough room to bend itself around every single example individually, the true ones and the smudged ones alike, without ever finding the pattern that would have explained them all at once.

You can watch this happen with a curve. Scatter a handful of measurements that follow some real, smooth pattern — plus the little accidents every real measurement carries. Fit them with a cramped machine and you get the smooth shape underneath. Fit them with a roomy one and you get a curve that visits every measurement personally, wiggling violently in between — a portrait of the accidents, not the pattern. It scores better than the honest curve on the points it studied. On fresh points, it is a disaster.

give it more dials and watch the rule dissolve into memory

Instrument 02 · The wiggle 3 dials
dashed: the real pattern · dots: what got measured · solid: what the machine believes
how much room it has3 dials
error on studied points error on fresh points
Slide right and watch the two numbers part company: the studied error falls toward zero while the fresh error climbs. The gap between those two numbers is the whole subject. A machine that has memorised looks better and better by the first number and worse and worse by the one that matters.

Do not take away that big machines are bad — the machines behind modern speech and language are enormous, and need to be. Take away that room converts studying into memorising unless something pushes back. The rest of this piece is about the pushing back.

step three

Studying past the point of learning

Room is one road to the pool. Time is the other — and this one you can watch happen live, which makes it the most useful picture in practical machine learning.

Before training starts, split the flashcards. Most go in the study pile. A handful go in a drawer, and the machine never, ever sees them — they exist so that you can ask it fresh questions whenever you like. Now train, and after every pass through the study pile, score it twice: once on cards it is studying, once on the drawer.

Two curves appear. Early on they fall together — the machine is genuinely learning, and real learning helps on the drawer nearly as much as on the pile, because what it is picking up is the rule. The drawer sits a little above the pile from the very first reading and always will; that small standing gap is not the problem. Then, somewhere, the curves separate. The study score keeps improving, pass after pass, looking like progress. The drawer score stalls — then turns and gets worse. That turn is the moment learning ran out and memorising began: everything gained after it is detail about the reflection, purchased by forgetting the world.

keep studying, and watch for the moment the curves part

Instrument 03 · The two curves not started
lower is better · gold: the study pile · teal: the drawer it has never seen
error on the study pile error on the drawer
The gold curve never once turns around — by the pool's own light, every extra hour looked like progress. Only the drawer can tell you when to stop, and the honest move is exactly what it looks like: stop at the teal curve's lowest point, and keep the machine from that moment, not the one that studied longest.

Engineers run this split on every serious training run, watching the drawer curve the way a pilot watches fuel — the drawer scores arrive every half hour, and nobody looks at anything else first. The two curves in this instrument are not one of those runs: they are drawn from a formula, so the shape is legible and identical every time the page opens. The shape is the honest part. A real run has a wobble these do not.

step four

The exam that leaked

By now the fix sounds simple: keep an exam the machine has never seen, and believe only that. It is the right fix. It has one failure mode, and the failure mode is the quietest scandal in modern machine learning: the exam leaks.

Nobody has to cheat on purpose. The piles these machines study are scooped from the internet by the millions, and the internet also contains the exams — the famous test questions, the benchmark sentences, the very recordings a field agreed to grade itself on. Scoop wide enough and some of the exam ends up in the study pile by accident. The machine memorises it along with everything else, the way it memorised the smudged cards: not knowing it is an exam, only knowing it has seen it before.

Then the score comes back, and it is wonderful, and it is not a lie so much as a wrong answer to a different question. It measured how much of the exam was at the pool. Every leaked question gets answered by reflection; only the clean remainder measures learning. The dishonest part is that the report does not say which was which — one shiny number, exactly like the innkeeper's ledger in the last piece, and inflated for exactly the same reason: the machine was graded on its own guests.

leak a little of the exam into the study pile and watch the score detach from the truth

Instrument 04 · The leak the exam is clean
every square is an exam question · gold: it leaked · rings: answered wrong
share of the exam that leaked into the study pile0%
the score it reports the truth, on unseen questions
The right-hand number is the machine's actual ability, and the slider cannot touch it. Everything the slider adds goes into the left number only — which is the number that gets put in the announcement, because it is the number anybody can measure.

This is not hypothetical for us either. This very month, mid-training, we found study cards sitting in one of our own exam drawers — nearly the whole English exam had leaked. We stopped the run the same day, rebuilt the drawer from cards the machine had provably never seen, and added a gate that checks every future pile against every exam before a single dial moves. The score got worse and became true. That trade is the job.

step five

What actually cures it

Overfitting has real cures, and after four instruments you can see why each one works. They all do the same thing from a different side: they make memorising a worse strategy than learning.

Give it more cards, and more varied ones — memorising a hundred cards is easy, memorising ten million is expensive, and long before the pile is exhausted the cheapest way to score well becomes finding the rule. Give it less room — a cramped machine cannot store the pile, so it is forced to summarise. And stop at the turn — hand the finished job to the version of the machine from the drawer curve's lowest moment, before the memorising hours were spent.

Notice the last piece's lesson running the other way. Procrustes' bed could not be fixed by more data, because more of the same road was more of the same lie. Narcissus can be — more varied cards genuinely cure memorising. Two diseases, opposite prescriptions, and telling them apart is much of an engineer's actual week: one is the pile is wrong for the world, the other is the machine is too devoted to the pile.

a machine that has memorised, and three honest medicines. try them in any combination

Instrument 05 · The medicines as found: roomy, over-studied
the same machine as instrument 02 · watch the fresh-question bar
error on studied points error on fresh points
Every cure makes the first number a little worse. That is not a side effect — it is the receipt. A machine slightly imperfect on its own pile has stopped polishing the reflection, and the fresh-question bar is where the refund arrives.

There are subtler medicines in the working engineer's cabinet — nudging the dials toward smallness, blurring the cards a little so no single fleck is worth memorising, asking one machine to check another. Every one of them cashes out the same way: memorising gets more expensive, learning gets comparatively cheap.

the trouble

A whole field can kneel at one pool

One more turn of the screw, because the pool does not only catch machines. It catches the people building them.

A benchmark — a public exam a whole field agrees to compete on — starts life as a drawer: clean, held out, trusted. But every published result is a glance at it. Thousands of researchers, tuning thousands of machines, each keeping whichever version scored best on that exam — nobody memorising anything, everybody choosing by the reflection. Year by year the numbers climb, and some of the climb is learning, and some of it is the field as a whole quietly fitting itself to one specific set of questions. The exam wears out. It stops measuring the world and starts measuring devotion to itself — and the only remedy is the one you already know, at a larger size: somebody has to build a fresh drawer.

Tiresias' prophecy turns out to be an engineering specification. He will live long if he never knows himself. A machine — or a field — stays honest exactly as long as there is a question it has never met. The moment everything has been seen, scores stop meaning anything, and nobody can tell the learner from the memoriser ever again. Keeping something unseen is not a chore on the checklist. It is the whole mechanism by which the truth stays measurable.

the last thing

Echo loved him. Of course she did.

The myth has one more gift, and it is the pairing. Echo — the same Echo from piece two, cursed so she could only repeat what she had just heard — loved Narcissus and could not tell him. He heard only his own words come back. She could only repeat; he could only recognise. Two ways of seeming to know things, and neither is knowing: giving back what you were given, and recognising what you have already seen.

Everything this series calls learning lives in the narrow country between those two failures — answering a question nobody in your pile ever asked, about a face you have never met, and being right anyway. That is the entire miracle these machines occasionally manage, and it is why the honest test is never how well do you know your pile? The pool answers that question perfectly, forever, and it means nothing.

So carry the one habit this piece was built to give you. Next time anybody — a company, a paper, an excited adult — tells you a machine scored some marvellous number, ask the only question Narcissus could not survive: had it seen the questions before? If nobody can prove it had not, you have learned the score measures the reflection. And you, unlike him, get to stand up and turn around.

the glossary

The words the engineers use

the poolthe training set
studying the reflection instead of the worldoverfitting
memorising the answer keymemorisation — fitting the training data instead of the pattern
learning the rule ×3generalisation
the smudged cardslabel noise — and memorising them is fitting noise
missing the smudged cards on purposea well-fit model disagrees with noisy labels
room to bend around every examplemodel capacity (parameter count is only a rough proxy for it)
the wiggle through every pointa high-variance fit — interpolating the noise
the drawer of cards it never seesthe held-out validation and test sets
the two curvestraining loss versus validation loss
the moment the curves partthe onset of overfitting
stopping at the teal curve's lowest pointearly stopping — keeping the best checkpoint, not the last
exam questions found at the pooldata leakage — train/test contamination
the score it reports vs the truthinflated benchmark performance
checking every pile against every examcontamination gates — deduplicating training data against evaluations
ten times the cardsmore, and more varied, training data
a quarter of the roomreducing capacity; regularisation generally
blurring the cards a littledata augmentation, dropout — making single examples not worth memorising
a whole field kneeling at one poolbenchmark overfitting — a community tuning to a public test set
building a fresh drawernew held-out benchmarks
Tiresias' condition — never knows himselfscores only mean anything on data the model has never seen
Echo repeating, Narcissus recognisingparroting and memorising — the two impostors of learning

One sentence to take away

Perfect on everything it studied proves only that it studied; the single question that measures learning is one the machine has never seen — so guard a drawer of those, stop training when that drawer says stop, and trust no score that cannot prove the exam stayed dry.

Every honest result in this field survived that sentence. Every scandal in it, somewhere underneath, is a machine — or a field — that was graded at its own pool.

the sources

Where these ideas come from

Tiresias' warning has a technical literature. Each note says what the paper actually established.

  1. L. Prechelt, “Early Stopping — But When?”, in Neural Networks: Tricks of the Trade (1998). The oldest honest fix, studied properly across 1,296 training runs: watch a held-out drawer, and stop when it says stop — not when the training score looks flattering.
  2. C. Zhang, S. Bengio, M. Hardt, B. Recht & O. Vinyals, “Understanding deep learning requires rethinking generalization”, ICLR (2017). The unsettling experiment: standard networks can perfectly memorise completely random labels — so the fact that they usually generalise cannot be explained by the textbook story, and memorising is always within reach.
  3. M. Belkin, D. Hsu, S. Ma & S. Mandal, “Reconciling modern machine-learning practice and the classical bias–variance trade-off”, PNAS (2019). Double descent: past the point of exactly fitting the pile, test error can start falling again — the U-shaped curve you learn first turns out to be only the left half of the picture.
  4. P. Nakkiran et al., “Deep Double Descent”, ICLR (2020). The same shape found everywhere in modern practice — across model size, training time, even data amount — including regimes where more data briefly makes things worse.
  5. C. Xu, S. Guan, D. Greene & M.-T. Kechadi, “Benchmark Data Contamination of Large Language Models: A Survey” (2024). The leak, at internet scale: how public exams seep into training piles, how researchers try to detect it, and what it does to reported scores — Narcissus grading himself at the pool, as a research field.