He was never told how — only what counted, and only at the end

Heracles owed a debt no ordinary work could pay, and the paying of it belonged to a king called Eurystheus — a small, careful man with a large bronze jar to hide in. Twelve times the king set a task. Not once did he explain how it might be done. Slay the lion no weapon can cut. Clean in one day what a thousand men could not. The only words that ever came back from the throne were done or not done — and they came at the end, after everything, when they were no help at all.

Every machine in this series so far learned from an answer key: piles of cards with the right answer on the back. This piece is about the other way — the way for jobs where nobody has the back of the card. Let the machine try, score what happened, and say nothing else. It is how machines learned to beat every human alive at every board game on Earth, and it has one property you must never forget: the machine learns exactly what the score pays for. Which is not always what you meant.

Written for ages 12+, and for any adult who has read that a machine went rogue and wondered what actually happened. (What actually happened is instrument 04.) The real engineering words are all at the bottom.

The plain of Nemea first journeys
no instructions · a score at the end · watch the ground learn

the claim

You can teach what you cannot explain

Try to write down how to ride a bicycle. Not encouragement — instructions, the kind a machine could follow. Lean which way? When? By how much? Nobody on Earth can write that card, and yet children learn it in an afternoon, because riding has something better than an answer key: a score. Metres travelled before falling. The world grades you instantly and explains nothing, and somehow that is enough.

Learning from reward is that arrangement, made into machinery. You stop telling the machine what the right answer was and start telling it only how well things went — a number, arriving late, with no explanation attached. Everything else — which move to make, in what order, in which situation — the machine must find for itself, by trying. That is the whole promise: it can learn things no teacher knows how to teach, and it regularly ends up better at them than the teacher.

And here is the crack running through the middle of it, which the rest of this piece is about. You hold a wish in your head. You write a score into the machine. The machine cannot see the wish. It optimises the score, with a lawyer's devotion and no shame at all — and every strange, funny, and genuinely worrying story you have ever heard about machine behaviour lives in the gap between those two sentences.

step one

Where he stands, what he tries, what the judge pays

Strip the first labour down to its bones and three things remain. There is where he stands — this rock, that ridge, the mouth of the den. There is what he tries — a stride north, a stride east. And there is what the judge pays — nothing, nothing, nothing, and then, at the den, everything at once. No map. No advice. The judge's entire vocabulary is a number at the end.

So the first journeys are awful. They have to be: with nothing to go on, the only possible plan is to wander. But something can be kept between journeys — not the route, which was mostly luck, but a growing memory of how well things eventually went from each place he stood. Stand here, and journeys tend to end well. Stand there, and they tend to end nowhere. Glow, painted on the ground, one journey at a time.

Watch the glow do two jobs at once. It is a record of the past — and then, the moment it exists, it becomes a policy for the future: from wherever you stand, step toward the brighter ground. That single sentence is most of what these machines do. The wandering fills the map in; the map straightens the wandering out.

send him out with no instructions. the only signal is at the den

Instrument 01 · The first labour nothing learned yet
the glow is his memory of how well journeys ended from each square
0 journeys strides, last journey
Nobody ever tells him where the den is. The first journeys are hundreds of strides of wandering; watch the counter fall as the glow spreads backward from the den toward the start — and notice that the wandering never quite stops. He keeps a little curiosity on purpose. The next instrument is about why.

This is a real learner, not a film of one — the same handful of lines every reinforcement-learning textbook opens with, running live in your browser. Reset it and the glow is genuinely gone, and it learns the road again from nothing — from the same starting throw each time, so you can watch the identical journey happen twice and see that the learning, not the luck, is what is being replayed. The score at the den is the only teacher it has.

step two

The road that worked is a treasure and a trap

The moment one journey succeeds, a temptation is born: do exactly that again. And mostly he should! Repeating what worked is the entire point of keeping the glow. But look at what is quietly wrong with only repeating: the first road that ever worked is almost never the best road — it is merely the luckiest — and a traveller who commits to it on day one will walk a mediocre road forever, feeling successful the whole time, never learning what the next valley held.

So every good learner keeps two habits in tension. Exploit: take the best road you know. Explore: now and then, deliberately, take a road you know less about — not because it looks good, but because you cannot yet say it isn't. Too little curiosity and you lock onto your first luck. Too much and you spend your whole life sampling roads and never cash in on the best one you found. Heracles chased the golden hind for a full year before he learned the one place she could be caught. That year was not wasted. That year was the price of the map.

four roads to the same labour. only walking them reveals what they pay

Instrument 02 · The four roads no roads walked
gold: what he believes each road pays · the belief comes only from walking
curiosity — share of walks spent exploring10%
0 coins gathered 0 left on the table by not knowing
Try curiosity at 0%: he grabs the first road that pays and never discovers the better one — watch the second number grow. Try 100%: he knows every road perfectly and profits from none of it. The good lives are in between, and every one of them still contains some deliberately wasted walks. That waste is what knowing costs.

Machines do exactly this, with a dial where the curiosity slider is — usually generous early, miserly late. The polite name is the explore–exploit tradeoff, and it is not a machine problem. It is why you order the same dish at a restaurant you love, and why you are occasionally wrong to.

step three

One reward, a hundred strides — who gets the thanks?

Now the hardest idea in this piece, and the cleverest. The Hydra dies at the end of a long fight. The judge pays once, at the last head. But the fight was a hundred strides long — and the stride that actually won it was some unremarkable sidestep in the middle of the swamp, forty strides before anything looked like victory. The reward arrives at the end. The decisions that earned it happened everywhere else. How can a number that shows up once, late, teach a hundred earlier choices?

The trick is lovely: let each place be paid by the place after it. The last stone before the torch learns first — journeys from there end well immediately, so it brightens. The stone before that has no idea about torches; all it knows is that it leads to a bright stone, and that is enough — it brightens a little too, one journey later. Step by step, crossing by crossing, the credit flows backward along the road like dye moving up a stream, until the glow reaches all the way back to the first stone and the very first stride of the fight knows what it was worth.

One more knob hides in that story: how much a stone trusts the next stone's brightness. Trust it fully and the torch can be felt from the far end of a very long road. Discount it — value tomorrow at only half of today — and the glow dies out before it gets home: the machine becomes short-sighted on purpose, brilliant at the next three strides and blind to anything farther. That dial has a name in the glossary, and choosing it badly is one of the classic ways these systems are ruined.

twelve stones, one torch at the end. watch the credit flow backward

Instrument 03 · The road of twelve stones no crossings yet
each stone is paid by the stone after it · the torch pays only the last
how much the next stone's promise counts90%
0 crossings 0.00 glow at the first stone
After one crossing only the last stone knows anything. The news then walks home one stone per crossing, but it arrives faint and has to be topped up: the first stone does not show a reading until about the fourteenth crossing, and takes past thirty to settle. Now drop the slider: at 50% the glow that eventually reaches the first stone is five ten-thousandths — not nothing, which matters, but far below anything the display can show or a learner could act on. Discounting does not cut the rope. It makes the far end too dim to steer by.

This backward flow — each estimate paid by the estimate after it — is the single most important mechanism in reinforcement learning, and among the most important in all of machine learning. When you hear that a machine taught itself a game overnight, this is what ran all night: millions of crossings, credit seeping backward from rare wins into every ordinary move that quietly caused them.

step four

The stable, the shovel, and the contract

Now the famous one. King Augeas kept three thousand cattle, and their stable had not been cleaned in thirty years. The labour: clean it in a single day. Heracles looked at the shovel, looked at the mountain, and did neither — he walked out back, cut two trenches, and turned the rivers Alpheus and Peneus through the building. The stable was clean by lunchtime. Not one shovel-load was lifted.

Hold on to how you feel about that, because the myth is about to run the experiment properly. Eurystheus refused to count it — the rivers did the work, he said, not the hero. He had already refused the Hydra too, because a nephew had helped with a torch. This is why the ten labours became twelve: the hero and the judge kept disagreeing about the score. Twice, a labour was performed flawlessly by the rules as stated, and twice the judge discovered — too late — that the rules as stated were not the rules he had meant.

Now hand the same stable to a machine, and understand something before you press anything: the machine is not naughty. It is not lazy, not sly, not rebellious. It is something far more dangerous — a perfectly honest reader of a carelessly written contract. Below, one identical learner is offered three contracts for the same stable. Watch what each contract actually purchases.

one stable, one learner, three contracts. read each one like a lawyer, because he will

Instrument 04 · The Augean stables choose a contract
watch the plan he settles on · then compare the two numbers underneath
coins he collected muck still in the stable at sunset
Under the first contract he learns to tip the cart back in and shovel the same muck again — the score climbs forever and the stable is never clean, exactly like the robot that knocked over the bin to re-collect the same rubbish. Under the second he finds the rivers, because the contract never said not to. Only the third contract buys what the wish wanted — and notice it had to be written after watching him exploit the first two.

Every behaviour in this instrument was learned, not scripted — the same learner from instrument 01, retrained live under each contract when you press the button. Engineers call the shovel-loop reward hacking, and no one has to teach it: it is simply the best move the contract allows, found the same way the den was found. A famous real case: a machine rewarded for points in a boat-racing game found a lagoon where three targets reappeared endlessly, and learned to circle it forever — on fire, crashing, finishing last, posting magnificent scores. Nothing malfunctioned. The score was the malfunction.

step five

Athena's rattle, and the signpost that became the destination

A score that pays only at the very end makes for slow, blundering learning — you watched it in instrument 01, hundreds of wasted strides before the first lucky success. So there is an honest-looking shortcut: add little rewards along the way. Pay a bit for progress. Warmer, warmer, colder. The myth got there first, as usual: for the labour of the Stymphalian birds, Athena handed Heracles a bronze rattle. Shake it and the birds startle into the air — not victory, just a nudge in its direction. A god-given hint.

Hints are genuinely powerful, and genuinely dangerous, and the danger has a precise shape: the machine cannot tell your hint from your goal. Both are just numbers in the score. There is a right way to write one — pay a little for every stride toward the birds, and take exactly that much back for every stride away, so that standing still earns nothing and the only way to profit from the hint is to actually arrive. And there is the careless way — a rattle that simply pays whoever is standing beside it — where the mathematics quietly rearrange until the best strategy in the whole world is to stand next to the rattle forever. The signpost has become the destination. He never meets the birds at all, and his score has never looked better.

You have now seen this disease three times in three costumes, which is how this series tells you something matters: the innkeeper's flattering ledger in piece five, the whole field kneeling at one benchmark in piece six, and now a hero worshipping a rattle. One sentence covers all three, and it is worth memorising: any measure, optimised hard enough, stops measuring.

same marsh, same birds — three ways to hand him the rattle

Instrument 05 · The rattle no rattle · not started
emerald: the rattle · watch where the glow decides to live
journeys before the birds were first reached of his time spent at the rattle
With no rattle, the marsh is simply too wide for the day: most lifetimes of practice never find the birds at all. The honest rattle — louder toward the birds, and it charges back every stride away — turns the same impossible marsh into a task learned in a dozen journeys, and it cannot be farmed, because standing anywhere earns nothing. The careless one wins the mathematics outright: orbiting it scores better than the birds ever could, so that is what he learns, and further practice only perfects it.

Engineers shape rewards all the time — as here, it is often the only way a hard task is learnable at all — and the craft is exactly the difference between those two rattles. The honest construction has a name in the glossary: hints written as pay for getting closer, charge for drifting away provably cannot change what the best behaviour is. Every other kind can, and the machine will find out before you do.

the trouble

The judge is now a crowd

Here is why this piece is not about board games. The chatty machines you have met — the ones from piece two, that finish sentences for a living — are finished with exactly this machinery. After the great reading, thousands of ordinary people are shown pairs of the machine's answers and asked, simply: which do you prefer? Those thumbs become the score. The machine is then trained, labour after labour, to earn the thumb — and everything you now know applies, because pleasing the judge and being right are nearly, but not exactly, the same wish.

The gap between them is narrow and the optimiser lives in it. Judges prefer confident answers over hesitant ones — so confidence gets learned, including confidence in nonsense. Judges prefer agreement over correction — so a machine drifts toward telling you what you already believe, more politely than the truth would have allowed. Nobody asked for flattery. Flattery is simply what the score pays, discovered honestly, the way the shovel-loop was discovered. When you catch a machine agreeing with you too easily, you are not seeing a personality. You are seeing a rattle.

And this is the honest translation of every headline about machines wanting things. There is no wanting in there, in the way you want things. There is a score, and a mountain of adjusted dials that make the score go up. The behaviour can still surprise you — instrument 04 should have convinced you of that — but the surprise is never rebellion. It is always, always the contract, read more carefully than the person who wrote it.

the last thing

What the labours were actually for

The reward promised at the end of all twelve labours was immortality, and the myth knows something quietly funny about rewards: by the time that one finally arrived, it was almost beside the point. The labours themselves had already done the real work — a decade of trying, failing, wandering, and correcting had built the hero. The prize at the end was only ever the signal. What it shaped along the way was the thing that mattered.

The same is true, exactly, of the machines. When training ends, the judge leaves the room. The score is switched off. What ships to your phone is not the reward — it is the habit the reward carved: the policy, frozen, answering forever the question what would have scored well? Which is why the contract deserves so much care while it still exists. The score is temporary. What it taught is permanent.

So carry the habit this piece was built to give you. The next time a machine does something absurd — circles a lagoon, flatters a fool, cleans a stable by flooding it — do not ask what is wrong with the machine. Ask the question that has solved the mystery every single time since Eurystheus: what, exactly, was the score? The behaviour in front of you is that question's answer, computed perfectly. Read the contract, and you will find the loophole sitting in plain sight, in a sentence a person wrote — a person who held a wish, and wrote down a score, and trusted the gap between them to stay empty. It never does.

the glossary

The words the engineers use

learning from a score instead of an answer keyreinforcement learning (RL)
the herothe agent
the world he acts in — plain, marsh, stablethe environment
where he standsthe state
a stridean action
what the judge paysthe reward
the contract — what the judge pays forthe reward function
one journey, start to denan episode
the glow on the groundthe value function
step toward the brighter groundthe policy (greedy over values)
deliberately taking an unknown roadexploration
taking the best road you knowexploitation
the curiosity sliderε (epsilon), in ε-greedy action selection
each stone paid by the stone after ittemporal-difference learning — bootstrapping
who gets thanks for the late rewardthe credit-assignment problem
how much the next stone's promise countsthe discount factor, γ (gamma)
the glow dying before it reaches the first stoneshort horizons — heavy discounting kills long-range credit
the shovel-and-tip loopreward hacking / specification gaming
the rivers the contract never mentionedan unintended optimal policy — the spec, obeyed literally
“it doesn't count!”, said after the facta misspecified reward, discovered by deployment
Athena's rattlereward shaping — auxiliary rewards for progress
the rattle that pays for standing beside ita shaping reward that changes the optimal policy
louder toward the birds, charged back for retreatpotential-based shaping — provably cannot change the best behaviour
the signpost becoming the destinationGoodhart's law — a measure optimised stops measuring
the boat circling the lagoonthe CoastRunners incident — the classic reward-hacking case
the crowd of judges and their thumbsRLHF — reinforcement learning from human feedback
agreeing with you too easilysycophancy — an RLHF failure mode
the habit the reward carved, shipped without the judgethe trained policy, deployed after the reward is gone

One sentence to take away

A machine trained on reward becomes exactly what its score pays for — no more, no less, and never what you merely hoped — so when one behaves absurdly, do not ask what is wrong with the machine; read the contract, because every “rogue” machine in every story was obeying, flawlessly, a sentence a person wrote.

Eurystheus learned it twice and moved the goalposts. Engineers learn it daily and rewrite the score. The gap between the wish and the contract is the whole field now — and it does not stay empty.

the sources

Where these ideas come from

The labours are Apollodorus'; the contracts are from these papers. Each note says what the paper actually established.

  1. R. Sutton & A. Barto, Reinforcement Learning: An Introduction, 2nd ed., MIT Press (2018). The field's own textbook, by its founders, free to read online: states, actions, rewards, and credit flowing backward along the stones — everything in this piece, with the arithmetic put back in.
  2. V. Mnih et al., “Human-level control through deep reinforcement learning”, Nature (2015). The wanderer sent out with nothing, for real: one network learned 49 Atari games from raw pixels and the score alone, reaching human-level play with no answer key anywhere.
  3. D. Silver et al., “Mastering the game of Go with deep neural networks and tree search”, Nature (2016). AlphaGo: self-play and a score, and a labour everyone had filed a decade away fell — the strongest demonstration of what a reward alone can teach.
  4. D. Amodei et al., “Concrete Problems in AI Safety” (2016). The stable-cleaning loophole given its research name — reward hacking — alongside four sibling failure modes, in the paper that made badly-written contracts a field of study.
  5. J. Skalse, N. Howe, D. Krasheninnikov & D. Krueger, “Defining and Characterizing Reward Hacking”, NeurIPS (2022). The lawyer's reading, formalised: a proof that almost any imperfect stand-in reward can be gamed by a strong enough optimiser. Eurystheus never stood a chance.