Theseus walked into the Labyrinth. Your phone does it every time you speak to it.
A computer chops one second of talking into about a hundred tiny slivers of sound. The word it has to write down might only be three letters long. Nobody ever tells it which slivers made which letter — it has to find its own way through.
The Greeks already told both versions of this story. One hero solves it by walking. One solves it by standing perfectly still and staring. Only one of them is affordable.
Written for ages 12+, and for any adult who has nodded along to the phrase "speech recognition" without ever being told what is actually happening. Every chapter has a Deeper into the maze box for anyone who wants the harder version.
one · Pan
Pan chased a nymph and got a musical instrument
Syrinx ran from Pan, begged the river for help, and was turned into a stand of reeds. Pan, being Pan, cut the reeds and made a set of pipes out of her. When he blows across them the air inside shivers — and that shivering is the entire subject of this chapter.
Sound is not a thing that flies through the air. Nothing travels from your mouth to someone's ear. What happens is that you shove the air next to your lips, that air shoves the air next to it, and a squeeze ripples outward while every single air particle just jiggles in place and stays home. Like a stadium wave: the wave crosses the whole stadium, nobody changes seats.
A microphone is a tiny drum skin that gets pushed and pulled by those squeezes. It writes down how hard it is being pushed, over and over and over. That is all a microphone can do, and it is all any speech system has ever had to work with — one wiggly line.
Deeper into the maze
The wiggly line gets written down as actual numbers — usually 16,000 of them every second. Each number answers one question: how hard is the drum skin being pushed right now?
Why 16,000? To capture a wiggle you need at least two measurements per wiggle, one for the up and one for the down. Human speech carries almost nothing above 8,000 wiggles per second, and 8,000 × 2 = 16,000. That is the whole reason for the number, and it has a name: the Nyquist limit.
1 second of speech = 16,000 numbers = a list, nothing more
two · Atropos
Atropos and her shears
Three sisters ran every life. Clotho spun the thread, Lachesis measured out its length, and Atropos — the one nobody argued with — cut it. She is here because the first thing any speech system does to your voice is take shears to it.
Sixteen thousand numbers a second is far too many to think about at once, and a single number on its own tells you nothing whatsoever. So the line gets cut into short lengths. Each length is a sliver — a tiny window of sound, roughly as long as one small part of one letter.
Here is the important bit. Each sliver stops being a wiggly line and becomes a little picture: a picture of which pitches were loud during that instant. Deep rumbles at the bottom, hiss at the top, bright wherever there was energy. A whole second of you talking becomes a strip of these pictures laid side by side, and that strip is what every system in the rest of this page is actually looking at.
Deeper into the maze
Real slivers are about 25 milliseconds long and a new one starts every 10 milliseconds — which means they overlap. Each sliver shares three-fifths of its sound with the one before it.
That overlap is deliberate. If the slivers sat end to end, a sound that happened to land on a cut would be sliced in half and might vanish. Overlapping means nothing can hide in a gap.
Doing this to 30 seconds of audio gives you 3,000 slivers. Most systems then quietly throw away every other one, because neighbouring slivers are nearly identical — which is how 3,000 becomes the number 1,500 you meet later on.
30 s → 3,000 slivers → keep 1 in 2 → ~1,500 slivers = T
three · Delphi
Nobody in this story gives a straight answer
When you climbed the mountain to ask the Oracle at Delphi a question, she never said one clean thing. She handed you a fog of maybes and left you to sort it out. Croesus asked whether to go to war and was told a great empire would fall. It did. His own.
A computer guessing a letter behaves exactly like her. It never says that one's an A. It writes out a whole scroll: a 40% chance of A, 3% of B, 12% of C, one line for every letter it knows. Real systems know about fifty thousand letters and symbols, so it is a very long scroll — and every number on it has to add up with all the others to exactly 100%.
Click through the slivers below and watch her change her mind.
Deeper into the maze
Turning a pile of raw scores into numbers that add up to 100% is a step called softmax. It has two jobs: make every number positive, and make them sum to one. The further ahead a score was, the more of the 100% it takes.
The "nothing" option is called blank, and it works harder than it looks. Without it, a system that sees the same letter across four slivers in a row has no way to tell whether you said "caat" or "cat".
four · Theseus
Theseus and the Labyrinth
Minos built a maze nobody could escape, and Theseus went in anyway. The rule of the maze is the useful bit: he only ever goes deeper. He never doubles back.
Lay the maze out as a grid. Stepping one square right means listen to one more sliver. Stepping one square down means write one more letter. Any staircase from the top-left corner to the bottom-right is one complete theory about which sliver made which letter.
Ariadne gave him a thread so he would know precisely which turns he had taken. That thread is worth a fortune two chapters from now, so remember it.
Here is the cost. To judge a route you need to know how promising each junction is, and the only way to know that is to station a whole Oracle at every single junction. Not one Oracle for the trip. One per junction, each with her own full scroll.
36 possible routes
Deeper into the maze
Counting routes is a choosing problem: out of every step you will ever take, how many are downward? That is the number in the readout.
routes = C(T + U − 2, U − 1)
You might hope to be clever and only build the junctions you actually visit. You cannot. Scoring needs every junction on the way forward, and learning from it needs every junction again on the way back. The entire grid has to exist in memory at once, which is exactly why it is so expensive.
five · Argus
Argus Panoptes, who never blinked
Hera needed something guarded, so she posted Argus: a giant covered in a hundred eyes. He did not patrol. He did not walk a route. He stood still and took in the whole field at once — some eyes locked hard on one spot, the rest drifting half-shut.
That is the second way to solve the puzzle. Before writing each letter, do not match anything to anything. Take in the entire stretch of sound in one go, staring hard at some slivers and barely at others, blur it all into a single impression, and ask the Oracle once.
Press the letter buttons and watch his eyes open and close. Then go back up to the Labyrinth and push its sliver count to 16 — Argus still needs exactly the same number of Oracles as before. One per letter. That is all he ever needs, no matter how long you talk.
one per letter, always
Deeper into the maze
How wide each eye is open is a number, and all of those numbers add up to exactly 1 — the same trick as the Oracle's scroll. They are called attention weights.
And there is not one Argus. A real system runs about twenty of them side by side at every layer, each with its own opinion about where to look, then combines them. Some watch for the previous letter, some watch for silence, and some do things nobody has ever managed to explain.
So what does the Labyrinth cost?
9×Exactly as many times more Oracles as there are slivers of sound. Change the sliver count in the Labyrinth above and this number moves with it.
Your buttons stop at 16. Real speech gives about 1,500 slivers. At that size Argus hires around a hundred Oracles and Theseus hires a hundred and fifty thousand — the difference between a file that fits on a phone and one that needs a room full of machines humming.
Which raises an obvious question. Why would anyone choose the maze?
six · Sisyphus
Because Argus becomes Sisyphus
Theseus has Ariadne's thread. When he comes out the other side the thread shows every turn he took — so you know not only what was said but exactly when each letter happened. And because the maze only lets him go forward, he physically cannot get lost or repeat himself. He can even shout the letters back to you while he is still walking, which is how live subtitles work.
Argus has no thread and no route. Nothing stops his gaze sliding back onto the same patch of grass he already stared at, and then onto it again, and again. He turns into Sisyphus shoving the same boulder up the same hill forever. This genuinely happens: give speech software a silent recording and it will sometimes fill the page with thank you, thank you, thank you. It has no sense of place, so it cannot tell it is stuck.
So neither hero wins. One is ruinously expensive and never loses his place. One is cheap and occasionally wanders off forever.
Deeper into the maze
Here is the arithmetic behind that number, for a real system with 1,500 slivers, 100 letters and a 50,000-entry scroll:
Argus 100 × 50,000 = 5,000,000 numbers Theseus 1,500 × 100 × 50,000 = 7,500,000,000 numbers
In actual memory that is a few megabytes against roughly thirty gigabytes — for the guesses alone, before any learning, for one sentence. Which is why maze-style systems use tiny alphabets of about a thousand entries and throw away three-quarters of their slivers, while Argus-style systems can afford an alphabet of fifty thousand.
shipping
Both heroes are in products you have used
The Labyrinth is a real design — engineers call it a transducer — and it is what runs live captions on a phone, the subtitles that appear while someone is still speaking, and anything that has to say exactly when each word occurred.
Argus is the other real design, an attention encoder-decoder, and it runs the big general-purpose systems that take a finished recording and hand back a finished page of text. Cheaper, better on messy audio, and prone to exactly the Sisyphus problem.
At Phonetico we build both, for Amharic, Tigrinya, Afaan Oromo and other Ethiopian languages — and the choice between the two heroes is a genuine argument we have, repeatedly, about real systems.
the glossary
The words the engineers use
| squeezes travelling through air | a pressure wave |
| the wiggly line the microphone writes | the waveform |
| Atropos' cuts | framing, or windowing |
| a sliver, drawn as a picture | a mel-spectrogram frame (T) |
| an Oracle's scroll of maybes | a probability distribution over the vocabulary (V) |
| making the scroll add up to 100% | softmax |
| the "nothing" option | the blank symbol |
| a letter written so far | a label position (U) |
| a junction of the Labyrinth | a lattice cell |
| the whole Labyrinth | the RNN-T lattice, [B, T, U, V] |
| one of Theseus's routes | an alignment path |
| only forward, never back | monotonic alignment |
| Ariadne's thread | the recovered alignment, i.e. timestamps |
| how wide each eye is open | attention weights |
| twenty Arguses at once | multi-head attention |
| the single impression | the context vector |
| one Oracle per letter | an attention decoder, [B, U, V] |
| Sisyphus and his boulder | attention looping |
the sources
Where these ideas come from
The Labyrinth and the giant are ours; the machinery is from these papers. Each note says what the paper actually established.
- A. Graves, S. Fernández, F. Gomez & J. Schmidhuber, “Connectionist Temporal Classification”, ICML (2006). The Labyrinth itself: how to train a network to transcribe audio when nobody has marked which sliver of sound made which letter — by summing over every route through the maze at once.
- A. Graves, “Sequence Transduction with Recurrent Neural Networks” (2012). The transducer: fixes CTC's odd habit of choosing each letter as if the others didn't exist, by giving the maze-walker a memory of what it has written so far. This is the machinery inside most on-device recognisers today — including ours.
- D. Bahdanau, K. Cho & Y. Bengio, “Neural Machine Translation by Jointly Learning to Align and Translate” (2014). Where Argus opened his eyes: the invention of attention — instead of one fixed summary, the model learns to look back at the relevant part of the input for each thing it writes.
- W. Chan, N. Jaitly, Q. Le & O. Vinyals, “Listen, Attend and Spell” (2015). Attention brought to speech: a listener compresses the audio, a speller writes letters while watching it — the hundred-eyed giant's way of hearing, in its original paper.