Dynamic Range Compression
The most-used, least-understood processor in audio: a circuit that turns a voice down the instant it gets loud, so that everything can be turned up afterwards. Here's where it came from, why every broadcast voice you've ever heard runs through one, and how to set one for speech.
What a compressor actually does
Human speech is wildly uneven. In a single sentence, the difference between a stressed syllable and the tail of a trailing word is routinely 20 dB — a hundred-fold difference in power. Your ears handle this effortlessly. Playback systems and listening environments don't: on a laptop speaker in a kitchen, the loud syllables are fine and the quiet ones vanish.
A compressor is an automatic volume control that closes that gap. It watches the level of the signal, and whenever the level crosses a threshold, it reduces the gain by an amount determined by the ratio. A 3:1 ratio means that for every 3 dB the input rises above the threshold, the output rises only 1 dB. Loud moments are pulled down; quiet moments pass untouched. Then makeup gain raises the whole, now-more-even signal back up. The result is not "quieter" — it's denser. The average level rises while the peaks stay controlled, which is why compressed speech sounds closer, more confident, and more intelligible on bad speakers.
A short history: from saving transmitters to defining a sound
Compression wasn't invented to sound good. It was invented to protect equipment. In the 1930s, radio broadcasters had a hard technical constraint: over-modulating an AM transmitter distorted the signal and could damage the plant, while under-modulating wasted precious coverage. Engineers needed the signal as hot as possible but never over. Human "gain riding" — an operator's hand on a fader — didn't scale, so the industry built automatic gain-reduction amplifiers. Western Electric's 110A limiter and, most famously, the WE 1126 (1937) were installed at transmitter sites to catch what the operator missed.
What began as protection became an aesthetic. Recording studios noticed that gain-reduced vocals sat better against a band. The 1950s and '60s produced the units engineers still argue about today: the Fairchild 660/670 variable-mu limiter (the sound of Beatles-era Abbey Road), Bill Putnam's LA-2A (1965) with its famously gentle optical gain cell, and the Urei 1176 (1967), the first FET compressor, fast enough to grab individual consonants. In the 1970s David Blackmer's dbx VCA designs made precise, repeatable ratio-based compression cheap enough for every studio and, eventually, every broadcast chain. By the digital era, compression had become fully mathematical — a gain computer and an envelope detector in code — which is exactly what runs inside your Mac today.
Why it exists: the dynamic range mismatch
Every audio medium has a window. Analog tape gave you maybe 60 usable decibels between hiss and saturation. FM broadcast, vinyl, telephone lines, laptop speakers, earbuds on a train — each has its own, mostly narrower, window. A raw voice recorded in a quiet room can span more range than the medium (or the listening room) can present. Something has to fold the signal into the window, and there are only two candidates: a human riding a fader, or a compressor.
For conversational media — radio, podcasts, video calls — there's a second reason: consistency between talkers. When one participant whispers and another booms, listeners reach for the volume control every time the speaker changes. Broadcast chains solved this decades ago with compression at multiple stages. Video call platforms apply some of their own, but by then a bad signal is already bad; compressing at the source, before transmission and before the platform's codec, always works better.
The controls, translated
- Threshold — the level where gain reduction starts. Set it so normal speech just tickles it and emphatic speech drives it. For a voice peaking around −12 dBFS, thresholds near −18 dBFS are typical.
- Ratio — how firmly loud material is held. 2:1–4:1 is transparent "glue" for speech; 10:1 and beyond is limiting, a safety net rather than a tone.
- Attack — how fast gain reduction engages. Too fast (<1 ms) and it clamps the crisp onset of consonants, dulling speech; too slow and loud syllables escape. A few milliseconds is the sweet spot for voice.
- Release — how fast gain returns afterwards. Too fast causes audible "pumping"; too slow and a loud sentence squashes the quiet one after it. 100–300 ms suits speech cadence.
- Makeup gain — the whole point. After the peaks are controlled, raise the signal so the average lands where you want it.
A useful mental model: the threshold and ratio decide how much control you apply; the attack and release decide whether anyone can hear you applying it. Well-set speech compression is only noticeable when you turn it off.
Compressing the spoken voice
Speech differs from music in ways that make it friendlier to compress. It's a single source, its syllabic rhythm is predictable (around 4 syllables per second), and nobody expects it to breathe dynamically the way a ballad does. For calls, podcasts and streams, the goal is 3–6 dB of gain reduction on emphatic syllables and none between phrases. That's enough to lift intelligibility on poor playback chains without the flattened, fatiguing quality of over-compression.
The classic mistake is treating a compressor as a loudness machine: driving 10–15 dB of reduction, then compensating with heavy makeup gain. That raises the room tone and breath noise between words (which is why compression is best applied after noise reduction), and it removes the natural emphasis speakers rely on to communicate. If everything is loud, nothing is.
In DeskBroadcast
DeskBroadcast's mic chain places its compressor exactly where broadcast tradition says it belongs: after noise reduction and the voice EQ, so it glues the finished tone rather than amplifying noise. The settings are speech-tuned and deliberately unadjustable: threshold −18 dBFS, 3 ms attack, 200 ms release, and +3 dB of makeup gain — a "studio glue" preset that evens out loud and quiet passages without audible pumping. It's a single toggle, on by default in the Studio Booth and Podcast voice presets, and the RMS and peak meters above it show you exactly what it's doing to your level. Downstream, auto-level handles the slower job of keeping your overall speech peaks near the −12 dBFS studio delivery target.
Hear it on your own voice
DeskBroadcast's Mic Check records eight seconds of your mic and replays it raw vs. processed — the fastest way to hear what compression, EQ and noise removal actually do.
Download DeskBroadcast