Delayed Auditory Feedback
Play a person's own voice back to them a fifth of a second late, and something remarkable happens: fluent speech falls apart. Words stretch, syllables repeat, the voice rises and slows. The effect is called delayed auditory feedback, it was documented in 1950, and it is the single best argument against trying to listen to yourself live through any software processing chain.
The experiment that made fluent speakers stutter
Speaking feels like a one-way act — thought out, words out — but it is a closed loop. Your brain continuously monitors the sound of your own voice and uses it to regulate timing, pitch and loudness on the fly. Nobody appreciated quite how load-bearing that loop was until it became easy to tamper with. Tape recorders with separate record and playback heads made it trivial: record on one head, play back from the other a moment later, and you have a speaker hearing themselves at a precise, adjustable delay.
Bernard S. Lee ran exactly this experiment and published the results in 1950–51, coining the phrase "artificial stutter." Under delays of a few hundred milliseconds, his fluent subjects began repeating syllables, prolonging vowels, slowing dramatically and involuntarily raising their voices. Some could push through with effort; none could simply ignore it. The delayed signal wasn't a distraction in the ordinary sense — it was corrupted feedback injected into a control loop, and the loop obediently tried to correct for an error that didn't exist.
What made Lee's finding so striking is that the subjects knew exactly what was happening and still couldn't override it. Speech monitoring runs below the level of deliberate attention: you don't decide to check your own voice any more than you decide to balance while walking. Give that automatic system stale data and it corrects against the past — holding a vowel because the confirmation hasn't arrived, re-launching a syllable it believes failed. The result looks like stuttering from the outside, but it's really a well-functioning controller being fed a lie.
A paradoxical clinical tool
The strangest part of the DAF literature is its asymmetry. The disruption is worst for fluent speakers at delays around 175–200 ms — roughly a syllable's length, which is presumably why it lands so squarely in the machinery of speech timing. Yet for some people who stutter, short delays of around 50–75 ms have the opposite effect: fluency improves. That paradox turned DAF from a curiosity into a clinical tool, and altered-feedback devices and therapy techniques have drawn on it for decades. The same knob, turned to different positions, either builds fluency or demolishes it.
The demolition setting has its own folklore. In 2012, Japanese researchers Kazutaka Kurihara and Koji Tsukada built the SpeechJammer — a handheld device combining a directional microphone and a directional speaker that bounces a talker's own words back at them a fraction of a second late, reliably jamming their speech from across a room. It earned an Ig Nobel Prize and a permanent place in every article about DAF, this one included, because it makes the underlying point vividly: delayed self-hearing isn't merely annoying. It reaches into speech production itself.
Why this dooms live self-monitoring on calls
Now consider what happens when you toggle on "listen to my processed mic" in any audio software. The signal must be captured in buffered blocks, pass through the processing chain — which, if it includes look-ahead stages, must add delay by design — then queue through a playback route and an output buffer before it reaches your ears. Each step is individually reasonable; the sum lands somewhere between tens and hundreds of milliseconds. In other words: a software monitor path delivers your own voice at almost exactly the delays Lee found most disruptive.
This is not a bug any vendor can fix with cleverer code. Hardware direct monitoring solves it for musicians by routing the mic to the headphones before the computer, at sub-millisecond analog speed — but then you're hearing the raw signal, which defeats the purpose of checking your processing. The moment the monitored signal must pass through the processing you want to evaluate, DAF physics wins. You can hold a brief phrase together while self-monitoring; you cannot conduct a natural conversation that way.
Bone conduction and the stranger in the recording
There's a second, quieter reason live monitoring misleads: even at zero delay, you have never heard your voice the way others hear it. When you speak, sound reaches your inner ear along two paths — through the air, and through bone conduction, vibration carried directly through the bones and tissue of your skull. The bone path favours low frequencies, so the voice you know from the inside is warmer and deeper than the one your microphone captures. This is why recordings of yourself sound thin and foreign — "do I really sound like that?" — and the answer is yes, to everyone but you. A monitoring signal fights not just your timing loop but a lifetime of miscalibrated expectation about your own tone.
The record-then-replay alternative
Both problems share one solution: stop judging your voice while producing it. Record a short passage, then listen back when you're no longer speaking. The feedback loop is out of the circuit, so DAF cannot touch you; and because you're now purely a listener, the bone-conduction bias becomes a known constant rather than a live distraction. This is, not coincidentally, how audio professionals have always worked — nobody mixes a vocal while singing it. Tracking and evaluating are separate activities, done with separate ears.
The detail that makes replay genuinely useful is the comparison. Listening to a processed recording in isolation tells you whether you sound good; it can't tell you what the processing did, because memory of your raw voice is unreliable at exactly the frequencies bone conduction distorts. Add an instant switch between the raw and processed versions of the same recording — flipped mid-sentence, at the same playback position — and the difference stops being a matter of memory. Identical performance, identical room, only the processing different: the one comparison live monitoring can never give you.
In DeskBroadcast
DeskBroadcast ships both approaches and is honest about which is for what. The Monitor toggle plays the processed signal live — and with processing plus its playback route it can add anywhere from roughly 60 to 250 ms, squarely in DAF territory. It exists for headphone ambience checks, not for talking. The Mic Check exists precisely because of DAF: it records eight seconds of the raw microphone and the processed chain simultaneously, then replays them with a Raw/Processed switch that flips mid-playback at the same position. You judge your own sound — the noise removal, the compression, all of it — without ever speaking over yourself, safely on speakers, with no feedback loop to fight.
Listen without the loop
The only fair way to hear your processed voice is after the fact. DeskBroadcast's Mic Check records eight seconds and lets you flip between raw and processed mid-playback — the comparison DAF makes impossible to do live.
Download DeskBroadcast