They never sang their own song. They sang yours, and it was perfect

The witch Circe warned Odysseus about the strait before he sailed into it. Two singers wait on an island of bones, she said, and no sailor has ever heard their song and lived. Here is the strange part, and Homer wrote it down. When the ship comes close, the Sirens do not sing about themselves. They call Odysseus by name. They offer to sing him the true story of his own war, the thing he has wanted to hear for ten long years. The most dangerous voice in the oldest sea story in the world is dangerous because it is aimed at you.

For seventy years, talking machines were easy to laugh at. You could hear the robot in the first word. That time is over. A machine can now say any sentence, in any voice it has listened to for a few seconds, and get the pauses and the breath and the warmth right. This piece shows you how, because making voices is what we do at Phonetico, on purpose, with permission. And it shows you what Odysseus knew: how to get through the strait even when the song is perfect.

Written for ages 12+, and for any adult who has read about a scam call that used a family member's voice and wondered how worried to be. Every chapter has a Deeper into the maze box for anyone who wants the harder version, and the real engineering words are all at the bottom.

The strait the song is becoming someone else’s
one song, slowly turning into a different singer · the shape of the wave is the voice

the claim

The machine is not playing a recording

Start with the fact everything else in this piece stands on. When a machine says a sentence to you, nobody ever recorded that sentence. Not in a studio, not in pieces, not once. The machine builds the sound right then, out of numbers, the way the sculptor in piece three built an owl out of static. Only this statue is made of air: thirty thousand tiny pushes against your ear every second, and the machine chooses every single one.

That surprises people, because the obvious way to build a talking machine is with recordings. Hire a very patient actor. Record every word in the dictionary. When you need a sentence, find the words and glue them together. For a long time this really was the plan, and you have heard the result at every train station: “the train. now approaching. platform. FOUR.” Every word is a real person. But each word was recorded on a different day, in a different mood, and the sentence sounds like it was built out of spare parts. Because it was.

Why does gluing sound so wrong? Because a spoken sentence is not a row of words like a written one. It is one long, smooth thing. The rise and fall of your voice runs through the whole sentence like a tune, and every sound leans into the next one before it finishes. When you glue recordings together, the tune breaks at every join, and your ear catches it instantly. So the new machines stopped gluing. They learned what talking sounds like from the great pile of examples, the one from piece five, and now they make each sentence fresh and whole, the way a singer sings one.

one short song, made both ways. this instrument makes sound, so turn your volume on

Instrument 01 · The joins silent · press a button
top: glued from separate pieces · bottom: made in one go · the gold tune line gives it away
5 pieces in the glued version 0 joins in the whole one
Same sounds, same order, same singer. In the glued version, each piece keeps the pitch it was recorded at, so the tune jumps at every join. That lurch is what the old station announcements were made of. The whole version draws one tune through the whole song first, then pours the sounds along it. Your ear forgives a lot, but it never forgives a broken tune.

The sound here really is made by a machine: a tiny singing machine that lives inside this page. It only sings vowels, like its ancestors the Sirens. Everything it plays is built from the same numbers that draw the picture. Nothing on this page is a recording. That is the whole point.

Deeper into the maze

The gluing era has a name: concatenative synthesis. The serious versions were enormous — one actor recorded for tens of hours, chopped not into words but into half-sounds called diphones, cut at the steady middle of each sound because the middles match better than the edges ever could.

It still was not enough, and the reason is physical. At every join, the pitch and the exact phase of the wave jump by a little, and your ear runs a pitch-tracker all day for free. A jump of a few wiggles per second, lasting a hundredth of a second, is all the lurch it needs.

The new machines skip the library entirely and write the wave itself, number by number. The page says thirty thousand pushes per second; real systems run at 22,050, 24,000 or 48,000 — “thirty thousand” is the honest middle of the family.

old: search ~20 h of one actor for pieces · new: invent 24,000 numbers, every second

step one

First the tune, then the voice

So how do you build a sentence of sound from nothing? The way machines in this series do everything: in steps, each one easier than the whole job. The text arrives as marks on a page. First the machine rewrites it as sounds. Not letters, because letters lie all the time: the ough in “through”, “rough” and “dough” makes three completely different noises. Speech has its own small honest units, and piece one met them from the listening side. This is the same list, used backwards.

Then comes the step nobody expects, and it is the one that decides everything. Before the machine makes a single sound, it composes the music of the sentence. How long each sound will last. Where the voice will rise and where it will fall. Where the breath goes. Which word gets pushed on. Try this out loud: say “she took the boat” like a plain fact. Now say it like a question. Now say it like an accusation. Same four words. Same sounds. Three totally different sentences. Everything that changed is the music, and engineers have a word for it: prosody.

Only after the music is decided does the machine paint the sound itself. First as a picture, the spectrogram from piece one, with time running left to right and pitch stacked upward, so the machine can sketch what the sound should look like before any air moves. Then a last machine, a kind of artificial throat, turns the picture into those thirty thousand little pushes per second. Words, then sounds, then the tune, then the picture, then the voice. The Sirens' meadow is the far end of an assembly line.

same sounds, three feelings. the tune is decided before the voice exists

Instrument 02 · The music before the voice a statement
the stack of stripes is the voice · the bright gold line is the tune the machine composed
how far the tune is allowed to swingnormal
the sounds, the same in all three
Press all three buttons in a row and watch what does not change: the sounds are the same every time. Only the gold line moves. Then drag the swing slider to zero and listen. Every feeling collapses into the same flat robot, and that is exactly what the old machines were: right sounds, no music. The robot voice was never a sound problem. It was a music problem.

In the real machines, nobody writes the music rules by hand. The machine learns them from thousands of hours of people talking, and it learns the human wobbles too: the little pauses, the breaths, the almost-stumbles. It has to. A voice with no wobbles at all sounds fake, because nobody is that smooth.

Deeper into the maze

The spelling-to-sounds step is called grapheme-to-phoneme conversion. Then the music engine predicts two things for every sound: how many hundredths of a second it lasts, and the tune. The tune has an exact meaning — how many times per second your vocal folds snap shut. Around 100 for a deep adult voice, 200 and up for a kid. Engineers call that number F0, and the gold line in the instrument is literally F0 plotted over time.

The picture the machine sketches before making any air is the mel spectrogram: about 80 pitch bands stacked upward, a fresh column every hundredth of a second. “Mel” means the bands are spaced the way your ear spaces pitches — crowded where you hear finely, sparse where you don't. The last machine, the vocoder, expands each skinny column of 80 numbers into 240 pushes of air.

marks → sounds → durations + tune → picture (80 × 100 per s) → air (24,000 per s)

step two

The meadow of all voices

Now the idea that turns a talking machine into a Siren. Nothing in the steps above says whose voice comes out. That is on purpose, and it is the trick at the center of this piece: inside the machine, the voice has been separated from the words. What makes your voice yours, the deepness or brightness of it, the size of the one throat on Earth that is yours, turns out to be a short list of numbers. And a short list of numbers is a location. Imagine a huge meadow where every possible human voice has a spot to stand: deep voices gathered down one slope, bright kid voices up another, every voice that has ever existed standing somewhere in the grass.

Park the machine on one spot and it says every sentence in that voice. Walk it slowly from one spot to another and the first voice melts into the second, passing through voices that belong to nobody at all. You watched that happen in the strait at the top of this page. And it should remind you of piece two, where words like king and queen stood near each other on a map of meanings. Same trick. Except the things standing in this meadow are not words. They are people's voices, turned into map coordinates.

Read that once more, because both halves of the rest of this piece come out of it. The wonderful half: a voice is now a thing that can be kept. It can be built for a language that never had a computer voice. It can be saved for a person who is losing theirs to illness. The frightening half: a spot in a meadow does not know who it belongs to. Anyone who finds your coordinates can stand the machine on them.

every dot is a voice. drag yourself around the meadow, with sound on if you can

Instrument 03 · The meadow wandering
drag the gold marker · the wave on the right is whatever spot you are standing on
pitch of this spot brightness of this spot
The dots are strangers. The gold marker is the machine, standing wherever you put it, including in the empty grass between the dots, where it sings perfectly well in a voice that has never existed. “Walk between two voices” is the strait from the top of the page, slowed down. It is not two recordings fading into each other. It is one machine going for a stroll, and every step of the stroll is a complete voice.

The real meadow has hundreds of directions, not two, and the machine lays it out by itself during the great reading. Nobody tells it which direction should mean “deeper”. The two directions here are just the two easiest ones to hear: how fast the voice buzzes, and how big the instrument around the buzz is.

Deeper into the maze

Your spot in the meadow is a list of roughly 200 to 500 numbers called a speaker embedding. Where does it come from? A machine plays a matching game over millions of clips: same person, or two different people? To win, it must learn to keep whatever stays constant when a person talks — the size and shape of the one throat — and throw away everything that changes: the words, the mood, the microphone. The meadow is the by-product of winning that game.

Distance in the meadow means something: two spots close together are two voices a listener would confuse. And the stroll between two voices is just arithmetic — average the two lists, number by number, and every point along the way is a complete, singable voice. Engineers call the stroll interpolation.

your voice ≈ 200 numbers ≈ one spot · distance between spots = how alike two voices sound

step three

Three seconds of you

So how does a stranger find your spot? Homer answered first, as usual. The Sirens had never met Odysseus, yet the song that reached him across the water already knew his name and his story. They had heard enough about him to aim. The modern version is blunt: a machine listens to a little scrap of your speech and reads your coordinates straight off it. Not your words. Your address in the meadow. And the scrap can be tiny. A voicemail greeting. A clip from a school play that someone posted. Three seconds gets close. Thirty seconds is barely better, for the same reason a second look at a stranger's face tells you little that the first look missed. Your voice is not a long story. It is a spot on a map, and spots are quick to read.

Once your spot is found, the separation from step two does the rest. Your voice can be attached to words you never said, spoken with feeling you never felt, because the feeling is composed by the music engine: urgency, tears, a little catch of breath, all of it made to order. That is the scam call in a grandchild's voice, and it is not rare anymore. It works because the listener's ear is being told the truth about the voice while being lied to about everything else.

Now the part we owe you, because this machinery is our job. Phonetico builds voices for languages that the world's machines never bothered to learn, and every voice we build starts with a person who said yes. Recorded on purpose. Paid. Told what the voice will be used for, in writing. Understand clearly: the machine cannot tell a gift from a theft. The meadow has no fences, and a spot does not know its owner. The same machine can be a new voice for someone losing theirs, or a weapon pointed at someone's grandmother, and nothing inside the machine knows which one it is being. Only people know that. So the rules have to live with the people: consent, contracts, law, and builders who refuse to copy a voice its owner never offered.

a hidden target voice. let the copycat overhear a few seconds, and watch it find the spot

Instrument 04 · Finding the spot nothing heard yet
the ring is the real voice's spot · the amber dot is the copycat's best guess so far
0 seconds overheard of careful listeners the copy would fool
Each overheard second gives the copycat one blurry glimpse of the target's coordinates, and it simply averages its glimpses. Watch the numbers, because they are the whole story: the first three seconds do almost all the work. After about ten, the guess sits inside the ring, which is the target voice's own day-to-day wobble. From there, the two ▶ buttons are an honest test, and you will mostly fail it.

This is why “I would know that voice anywhere” stopped being protection. Notice something else too. Our contributors hand over their voice sample on purpose, with a signature. The scam's victim never handed over anything. Three seconds of them was simply lying around on the internet.

Deeper into the maze

The trick's real name is zero-shot voice cloning, and the scrap is the enrollment audio. Note what does not happen: the copier is never retrained on the victim. The scrap is pushed through the same matching-game machine from the last box, and the address falls out the other end. That is why it takes seconds, not weeks — it is reading, not learning.

And here is why three seconds gets close and thirty is barely better. Each second of tape is one blurry glimpse of the same fixed spot, and the copycat averages its glimpses. Averaging shrinks the error like √N: four times the tape buys only twice the precision. After about ten seconds the remaining error is already smaller than your voice's own day-to-day wobble — tired, excited, a different room — so more tape stops mattering at all.

The paper that made this famous trained on 60,000 hours of voices and continues a stranger's voice, feeling and room from a 3-second prompt. It is reference 3 at the bottom of this page.

error after N seconds ∝ 1 / √N · 4× the tape → only 2× closer

the trouble

Orpheus, and why your ear will lose

A second ship once passed the Sirens, and it did not use wax. When the Argonauts came close, Orpheus took out his lyre and played over them: a louder, truer music, so the Sirens' song could not land. It is a beautiful defense, and it is the one everybody reaches for today. Surely we can build a machine that hears the fake? A detector. An Orpheus. And we can! Detectors exist, and they catch yesterday's fakes well. Then next year's voice machines get trained, sometimes trained against the detector itself, the way the forger in piece three trained against the critic, and the clues the detector listened for disappear. New detector, new clues, new fakes. You have seen this before: it is the arms race from piece seven, and your unaided ear dropped out of that race years ago.

There is one more honest idea. What if the machine signed its own song? A watermark: a pattern folded into the sound, too quiet for you to hear, easy for a checking program to find, saying a machine made this. Serious builders do this, and we do it too. But only half-trust it, for a reason a twelve-year-old can spot in a second: the signature is written by the maker. Honest makers sign their fakes. A scammer's machine simply does not sign. So a watermark can prove a sound is fake. No watermark proves nothing at all.

real calls and copied calls, and an ear trying its best. then stop playing that game

Instrument 05 · The wax and the mast the ear, unaided
left hill: real calls · right hill: the copy's calls · an ear can only draw a line between them
seconds of the victim the scammer has heard2 s
best possible ear, % of calls judged rightly with the plan, no matter the slider
Drag the slider right and watch the hills merge. Past a dozen seconds, the best ear on Earth is barely better than a coin flip, because the clue it needs is no longer in the sound at all. Then press the mast. The number on the right never looks at the hills. The plan does not listen to the call, which is why the slider, and every future improvement in fake voices, cannot touch it.

The plan is anything the voice cannot carry. Hang up and call back on the number you already have. A family code word, agreed at dinner, asked out loud. Any question whose answer never touched the internet. Odysseus got every detail right centuries early. The wax: do not take the call at all. The mast: decide what you will never do on an incoming call, before it rings. No money, no codes, no secrets. And the detail everyone forgets: he told his crew that if he begged to be untied, they must tie him tighter. The begging is the attack. When a voice on the phone says right now, do not hang up, do not tell anyone, that is not a reason to bend your rule. That is the surest sign your rule is working.

Deeper into the maze

The two hills are a picture every detection engineer carries in their head. A detector's honesty is measured at the point where its two mistakes balance — calling a real call fake exactly as often as it lets a fake through. That balance point is the equal error rate, and a coin flip sits at 50%.

On fakes made by machines a detector saw during training, modern detectors are genuinely good: an equal error rate of a few percent. On fakes from a machine it has never met, the same detector can do ten times worse, because the clue it had memorized was a habit of one particular machine, not a law of fakery. That is the arms race in one sentence, and it is measured publicly every two years in a contest called ASVspoof — reference 5 below.

a watermark is strong evidence of “a machine made this” · it can never prove “a human made this”

the last thing

What Odysseus actually understood

Strip the myth down and look at what the cleverest man in Greece did not do. He did not train himself to resist the song. He assumed he would fail. He did not trust himself to judge things in the moment. He assumed his judgment would be the first thing the song took. Every piece of his plan, the wax, the rope, the orders to the crew, was made on deck, in calm water, before the first note. And the plan was built to keep working even while the smartest mind on the ship was screaming to be released. He is the only sailor who ever heard the song and lived, and he managed it with a defense that never needed him to win.

That is the lesson worth more than any detector. The song is only going to get better. One day the voice on the phone will be your mother's, perfect down to the little catch in her breath, and it will mean nothing, because a voice is a spot in a meadow and a spot can be read from three seconds of tape. It is fine to be a little sad about that. Something real was lost. Then do what Odysseus did. Move your trust out of your ear, which is losing, and into your plan, which cannot lose: the call-back, the code word, the rule you made before the phone rang.

And keep the other half of the story too, because we will not end a piece about our own work on fear alone. The same meadow holds a computer voice for languages that never had one, reading to children in Amharic. It holds the saved voice of a woman with a disease that is taking hers, still saying goodnight. It holds every audiobook and every screen reader that ever walked a blind traveler down an unfamiliar street. The Sirens and the singers stand in the same grass, built from the same numbers. The difference between them was never inside the machine. The difference is whether anyone asked the voice's owner first, and that question is guarded by people. In the end it is the only place it could be guarded.

Deeper into the maze

Security engineers have a name for Odysseus's plan: out-of-band verification. Move the check onto a channel the attacker does not hold. The scammer controls everything that arrives on the incoming call — the voice, and even the number on your screen, which can be forged. What they do not control is the phone network's routing when you dial a number you already had. Same phone, opposite direction, different channel entirely.

The family code word is what engineers call a shared secret: a key exchanged at the dinner table, never spoken online, so no amount of listening to the internet can recover it. And “tie me tighter” is the deepest principle in the whole field — the protocol is fixed in calm water and never renegotiated under pressure, because the pressure is the attack. Real security systems are built exactly this way: they assume the human in the loop will be compromised, and they are designed to hold anyway.

the voice forgeable the caller ID forgeable the call YOU place, to the number YOU had, asking the word YOU agreed — not

the glossary

The words the engineers use

making a voice out of marks on a pagespeech synthesis, or text-to-speech (TTS)
talking by gluing old recordingsconcatenative synthesis
the joins your ear catchesconcatenation artifacts, discontinuities at unit boundaries
making each sentence fresh and wholeneural (generative) TTS
the small honest units of speechphonemes
the music of a sentence: timing, pitch, breathprosody
the tune drawn through the whole sentencethe F0 (pitch) contour
the picture of the sound, sketched firstthe mel spectrogram
the artificial throat that turns picture into soundthe vocoder
thirty thousand pushes per secondthe waveform, at the sample rate
the meadow of all voicesspeaker-embedding space
one voice, one spota speaker embedding
the voice separated from the wordsdisentangled speaker identity and content
strolling between two voicesinterpolation in embedding space
reading someone's spot off a scrap of speechzero-shot voice cloning
the scrap itselfenrollment or reference audio
three seconds was enoughfew-second cloning; error falls fast at first, then flattens past a few seconds
your voice, saying words you never saidan audio deepfake
Orpheus, playing over the songsynthetic-speech detection, anti-spoofing
new clues, new listeners, foreverthe detection arms race
the machine signing its own songaudio watermarking, provenance marking
no signature proves nothinga watermark shows presence, never absence
the wax in the crew's earsscreening: do not act on unsolicited calls
the mast, and the standing ordersout-of-band verification, protocols fixed in advance
the family code worda shared secret, a challenge phrase
“tie me tighter”urgency is the attack; never renegotiate the protocol mid-call
a person who said yes, in writingconsent and licensing of voice data

One sentence to take away

A voice can now be copied from three seconds of sound, so a voice alone proves nothing, however much you love it; the real defense is not a sharper ear but a plan made in calm water, before the phone rings, and never renegotiated while the song is playing.

Odysseus heard the most persuasive sound in the world with his own ears and sailed home anyway. Not because he was strong, but because he had already decided what he would refuse before the singing began.

the sources

Where these ideas come from

The song is Homer's; the machinery that can now sing it is from these papers. Each note says what the paper actually established.

  1. A. van den Oord et al., “WaveNet: A Generative Model for Raw Audio” (2016). The moment the joins disappeared: speech generated one raw audio sample at a time, rated dramatically more natural than the glued-together voices before it.
  2. J. Shen et al., “Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions”, ICASSP (2018). Tacotron 2, the standard modern pipeline: text to a picture of sound, picture to a voice — with listeners rating it 4.53 against 4.58 for real recordings — close enough to be called comparable, though asked to choose side by side they still preferred the human, by a small but real margin.
  3. C. Wang et al., “Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers” (2023). VALL-E: the three-second copycat's paper — treat speech as tokens, train on 60,000 hours, and a stranger's voice, feeling and room can be continued from one short clip.
  4. Y. Chen et al., “F5-TTS” (2024). How accessible this became: an open-source system that clones voices quickly on ordinary hardware — the state of the meadow as of 2024, and machinery of exactly the family we build with, under consent, at Phonetico.
  5. J. Yamagishi et al., “ASVspoof 2021: accelerating progress in spoofed and deepfake speech detection” (2021). Orpheus, organised: the community benchmark where detectors compete to hear the Sirens for what they are — and the honest record of how hard that contest is.