He was never told how — only what counted, and only at the end
Heracles owed a debt no ordinary work could pay, and the paying of it belonged to a king called Eurystheus — a small, careful man with a large bronze jar to hide in. Twelve times the king set a task. Not once did he explain how it might be done. Slay the lion no weapon can cut. Clean in one day what a thousand men could not. The only words that ever came back from the throne were done or not done — and they came at the end, after everything, when they were no help at all.
Every machine in this series so far learned from an answer key: piles of cards with the right answer on the back. This piece is about the other way — the way for jobs where nobody has the back of the card. Let the machine try, score what happened, and say nothing else. It is how machines learned to beat every human alive at every board game on Earth, and it has one property you must never forget: the machine learns exactly what the score pays for. Which is not always what you meant.
Written for ages 12+, and for any adult who has read that a machine went rogue and wondered what actually happened. (What actually happened is instrument 04.) The real engineering words are all at the bottom.
the claim
You can teach what you cannot explain
Try to write down how to ride a bicycle. Not encouragement — instructions, the kind a machine could follow. Lean which way? When? By how much? Nobody on Earth can write that card, and yet children learn it in an afternoon, because riding has something better than an answer key: a score. Metres travelled before falling. The world grades you instantly and explains nothing, and somehow that is enough.
Learning from reward is that arrangement, made into machinery. You stop telling the machine what the right answer was and start telling it only how well things went — a number, arriving late, with no explanation attached. Everything else — which move to make, in what order, in which situation — the machine must find for itself, by trying. That is the whole promise: it can learn things no teacher knows how to teach, and it regularly ends up better at them than the teacher.
And here is the crack running through the middle of it, which the rest of this piece is about. You hold a wish in your head. You write a score into the machine. The machine cannot see the wish. It optimises the score, with a lawyer's devotion and no shame at all — and every strange, funny, and genuinely worrying story you have ever heard about machine behaviour lives in the gap between those two sentences.
step one
Where he stands, what he tries, what the judge pays
Strip the first labour down to its bones and three things remain. There is where he stands — this rock, that ridge, the mouth of the den. There is what he tries — a stride north, a stride east. And there is what the judge pays — nothing, nothing, nothing, and then, at the den, everything at once. No map. No advice. The judge's entire vocabulary is a number at the end.
So the first journeys are awful. They have to be: with nothing to go on, the only possible plan is to wander. But something can be kept between journeys — not the route, which was mostly luck, but a growing memory of how well things eventually went from each place he stood. Stand here, and journeys tend to end well. Stand there, and they tend to end nowhere. Glow, painted on the ground, one journey at a time.
Watch the glow do two jobs at once. It is a record of the past — and then, the moment it exists, it becomes a policy for the future: from wherever you stand, step toward the brighter ground. That single sentence is most of what these machines do. The wandering fills the map in; the map straightens the wandering out.
send him out with no instructions. the only signal is at the den
This is a real learner, not a film of one — the same handful of lines every reinforcement-learning textbook opens with, running live in your browser. Reset it and the glow is genuinely gone, and it learns the road again from nothing — from the same starting throw each time, so you can watch the identical journey happen twice and see that the learning, not the luck, is what is being replayed. The score at the den is the only teacher it has.
step two
The road that worked is a treasure and a trap
The moment one journey succeeds, a temptation is born: do exactly that again. And mostly he should! Repeating what worked is the entire point of keeping the glow. But look at what is quietly wrong with only repeating: the first road that ever worked is almost never the best road — it is merely the luckiest — and a traveller who commits to it on day one will walk a mediocre road forever, feeling successful the whole time, never learning what the next valley held.
So every good learner keeps two habits in tension. Exploit: take the best road you know. Explore: now and then, deliberately, take a road you know less about — not because it looks good, but because you cannot yet say it isn't. Too little curiosity and you lock onto your first luck. Too much and you spend your whole life sampling roads and never cash in on the best one you found. Heracles chased the golden hind for a full year before he learned the one place she could be caught. That year was not wasted. That year was the price of the map.
four roads to the same labour. only walking them reveals what they pay
Machines do exactly this, with a dial where the curiosity slider is — usually generous early, miserly late. The polite name is the explore–exploit tradeoff, and it is not a machine problem. It is why you order the same dish at a restaurant you love, and why you are occasionally wrong to.
step three
One reward, a hundred strides — who gets the thanks?
Now the hardest idea in this piece, and the cleverest. The Hydra dies at the end of a long fight. The judge pays once, at the last head. But the fight was a hundred strides long — and the stride that actually won it was some unremarkable sidestep in the middle of the swamp, forty strides before anything looked like victory. The reward arrives at the end. The decisions that earned it happened everywhere else. How can a number that shows up once, late, teach a hundred earlier choices?
The trick is lovely: let each place be paid by the place after it. The last stone before the torch learns first — journeys from there end well immediately, so it brightens. The stone before that has no idea about torches; all it knows is that it leads to a bright stone, and that is enough — it brightens a little too, one journey later. Step by step, crossing by crossing, the credit flows backward along the road like dye moving up a stream, until the glow reaches all the way back to the first stone and the very first stride of the fight knows what it was worth.
One more knob hides in that story: how much a stone trusts the next stone's brightness. Trust it fully and the torch can be felt from the far end of a very long road. Discount it — value tomorrow at only half of today — and the glow dies out before it gets home: the machine becomes short-sighted on purpose, brilliant at the next three strides and blind to anything farther. That dial has a name in the glossary, and choosing it badly is one of the classic ways these systems are ruined.
twelve stones, one torch at the end. watch the credit flow backward
This backward flow — each estimate paid by the estimate after it — is the single most important mechanism in reinforcement learning, and among the most important in all of machine learning. When you hear that a machine taught itself a game overnight, this is what ran all night: millions of crossings, credit seeping backward from rare wins into every ordinary move that quietly caused them.
step four
The stable, the shovel, and the contract
Now the famous one. King Augeas kept three thousand cattle, and their stable had not been cleaned in thirty years. The labour: clean it in a single day. Heracles looked at the shovel, looked at the mountain, and did neither — he walked out back, cut two trenches, and turned the rivers Alpheus and Peneus through the building. The stable was clean by lunchtime. Not one shovel-load was lifted.
Hold on to how you feel about that, because the myth is about to run the experiment properly. Eurystheus refused to count it — the rivers did the work, he said, not the hero. He had already refused the Hydra too, because a nephew had helped with a torch. This is why the ten labours became twelve: the hero and the judge kept disagreeing about the score. Twice, a labour was performed flawlessly by the rules as stated, and twice the judge discovered — too late — that the rules as stated were not the rules he had meant.
Now hand the same stable to a machine, and understand something before you press anything: the machine is not naughty. It is not lazy, not sly, not rebellious. It is something far more dangerous — a perfectly honest reader of a carelessly written contract. Below, one identical learner is offered three contracts for the same stable. Watch what each contract actually purchases.
one stable, one learner, three contracts. read each one like a lawyer, because he will
Every behaviour in this instrument was learned, not scripted — the same learner from instrument 01, retrained live under each contract when you press the button. Engineers call the shovel-loop reward hacking, and no one has to teach it: it is simply the best move the contract allows, found the same way the den was found. A famous real case: a machine rewarded for points in a boat-racing game found a lagoon where three targets reappeared endlessly, and learned to circle it forever — on fire, crashing, finishing last, posting magnificent scores. Nothing malfunctioned. The score was the malfunction.
step five
Athena's rattle, and the signpost that became the destination
A score that pays only at the very end makes for slow, blundering learning — you watched it in instrument 01, hundreds of wasted strides before the first lucky success. So there is an honest-looking shortcut: add little rewards along the way. Pay a bit for progress. Warmer, warmer, colder. The myth got there first, as usual: for the labour of the Stymphalian birds, Athena handed Heracles a bronze rattle. Shake it and the birds startle into the air — not victory, just a nudge in its direction. A god-given hint.
Hints are genuinely powerful, and genuinely dangerous, and the danger has a precise shape: the machine cannot tell your hint from your goal. Both are just numbers in the score. There is a right way to write one — pay a little for every stride toward the birds, and take exactly that much back for every stride away, so that standing still earns nothing and the only way to profit from the hint is to actually arrive. And there is the careless way — a rattle that simply pays whoever is standing beside it — where the mathematics quietly rearrange until the best strategy in the whole world is to stand next to the rattle forever. The signpost has become the destination. He never meets the birds at all, and his score has never looked better.
You have now seen this disease three times in three costumes, which is how this series tells you something matters: the innkeeper's flattering ledger in piece five, the whole field kneeling at one benchmark in piece six, and now a hero worshipping a rattle. One sentence covers all three, and it is worth memorising: any measure, optimised hard enough, stops measuring.
same marsh, same birds — three ways to hand him the rattle
Engineers shape rewards all the time — as here, it is often the only way a hard task is learnable at all — and the craft is exactly the difference between those two rattles. The honest construction has a name in the glossary: hints written as pay for getting closer, charge for drifting away provably cannot change what the best behaviour is. Every other kind can, and the machine will find out before you do.
the trouble
The judge is now a crowd
Here is why this piece is not about board games. The chatty machines you have met — the ones from piece two, that finish sentences for a living — are finished with exactly this machinery. After the great reading, thousands of ordinary people are shown pairs of the machine's answers and asked, simply: which do you prefer? Those thumbs become the score. The machine is then trained, labour after labour, to earn the thumb — and everything you now know applies, because pleasing the judge and being right are nearly, but not exactly, the same wish.
The gap between them is narrow and the optimiser lives in it. Judges prefer confident answers over hesitant ones — so confidence gets learned, including confidence in nonsense. Judges prefer agreement over correction — so a machine drifts toward telling you what you already believe, more politely than the truth would have allowed. Nobody asked for flattery. Flattery is simply what the score pays, discovered honestly, the way the shovel-loop was discovered. When you catch a machine agreeing with you too easily, you are not seeing a personality. You are seeing a rattle.
And this is the honest translation of every headline about machines wanting things. There is no wanting in there, in the way you want things. There is a score, and a mountain of adjusted dials that make the score go up. The behaviour can still surprise you — instrument 04 should have convinced you of that — but the surprise is never rebellion. It is always, always the contract, read more carefully than the person who wrote it.
the last thing
What the labours were actually for
The reward promised at the end of all twelve labours was immortality, and the myth knows something quietly funny about rewards: by the time that one finally arrived, it was almost beside the point. The labours themselves had already done the real work — a decade of trying, failing, wandering, and correcting had built the hero. The prize at the end was only ever the signal. What it shaped along the way was the thing that mattered.
The same is true, exactly, of the machines. When training ends, the judge leaves the room. The score is switched off. What ships to your phone is not the reward — it is the habit the reward carved: the policy, frozen, answering forever the question what would have scored well? Which is why the contract deserves so much care while it still exists. The score is temporary. What it taught is permanent.
So carry the habit this piece was built to give you. The next time a machine does something absurd — circles a lagoon, flatters a fool, cleans a stable by flooding it — do not ask what is wrong with the machine. Ask the question that has solved the mystery every single time since Eurystheus: what, exactly, was the score? The behaviour in front of you is that question's answer, computed perfectly. Read the contract, and you will find the loophole sitting in plain sight, in a sentence a person wrote — a person who held a wish, and wrote down a score, and trusted the gap between them to stay empty. It never does.
the glossary
The words the engineers use
| learning from a score instead of an answer key | reinforcement learning (RL) |
| the hero | the agent |
| the world he acts in — plain, marsh, stable | the environment |
| where he stands | the state |
| a stride | an action |
| what the judge pays | the reward |
| the contract — what the judge pays for | the reward function |
| one journey, start to den | an episode |
| the glow on the ground | the value function |
| step toward the brighter ground | the policy (greedy over values) |
| deliberately taking an unknown road | exploration |
| taking the best road you know | exploitation |
| the curiosity slider | ε (epsilon), in ε-greedy action selection |
| each stone paid by the stone after it | temporal-difference learning — bootstrapping |
| who gets thanks for the late reward | the credit-assignment problem |
| how much the next stone's promise counts | the discount factor, γ (gamma) |
| the glow dying before it reaches the first stone | short horizons — heavy discounting kills long-range credit |
| the shovel-and-tip loop | reward hacking / specification gaming |
| the rivers the contract never mentioned | an unintended optimal policy — the spec, obeyed literally |
| “it doesn't count!”, said after the fact | a misspecified reward, discovered by deployment |
| Athena's rattle | reward shaping — auxiliary rewards for progress |
| the rattle that pays for standing beside it | a shaping reward that changes the optimal policy |
| louder toward the birds, charged back for retreat | potential-based shaping — provably cannot change the best behaviour |
| the signpost becoming the destination | Goodhart's law — a measure optimised stops measuring |
| the boat circling the lagoon | the CoastRunners incident — the classic reward-hacking case |
| the crowd of judges and their thumbs | RLHF — reinforcement learning from human feedback |
| agreeing with you too easily | sycophancy — an RLHF failure mode |
| the habit the reward carved, shipped without the judge | the trained policy, deployed after the reward is gone |
One sentence to take away
A machine trained on reward becomes exactly what its score pays for — no more, no less, and never what you merely hoped — so when one behaves absurdly, do not ask what is wrong with the machine; read the contract, because every “rogue” machine in every story was obeying, flawlessly, a sentence a person wrote.
Eurystheus learned it twice and moved the goalposts. Engineers learn it daily and rewrite the score. The gap between the wish and the contract is the whole field now — and it does not stay empty.
the sources
Where these ideas come from
The labours are Apollodorus'; the contracts are from these papers. Each note says what the paper actually established.
- R. Sutton & A. Barto, Reinforcement Learning: An Introduction, 2nd ed., MIT Press (2018). The field's own textbook, by its founders, free to read online: states, actions, rewards, and credit flowing backward along the stones — everything in this piece, with the arithmetic put back in.
- V. Mnih et al., “Human-level control through deep reinforcement learning”, Nature (2015). The wanderer sent out with nothing, for real: one network learned 49 Atari games from raw pixels and the score alone, reaching human-level play with no answer key anywhere.
- D. Silver et al., “Mastering the game of Go with deep neural networks and tree search”, Nature (2016). AlphaGo: self-play and a score, and a labour everyone had filed a decade away fell — the strongest demonstration of what a reward alone can teach.
- D. Amodei et al., “Concrete Problems in AI Safety” (2016). The stable-cleaning loophole given its research name — reward hacking — alongside four sibling failure modes, in the paper that made badly-written contracts a field of study.
- J. Skalse, N. Howe, D. Krasheninnikov & D. Krueger, “Defining and Characterizing Reward Hacking”, NeurIPS (2022). The lawyer's reading, formalised: a proof that almost any imperfect stand-in reward can be gamed by a strong enough optimiser. Eurystheus never stood a chance.