Sample Rates & 48 kHz
Digital audio is a series of snapshots — tens of thousands of measurements of a waveform every second. How many snapshots you take per second is the sample rate, and behind that one number sits a beautiful theorem, a quirk of 1970s video recorders, and the reason your video call chain almost certainly runs at 48,000 of them.
What a sample rate is
A microphone produces a continuously varying voltage. To get that into a computer, an analog-to-digital converter measures the voltage at strictly regular intervals and stores each measurement as a number. The sample rate is how many of those measurements happen per second: at 48 kHz, the converter takes 48,000 snapshots every second, one every 20.8 microseconds. Playback reverses the process — a digital-to-analog converter turns the stream of numbers back into a smooth, continuous voltage for a speaker or headphone amplifier. The question that defined early digital audio was: how many snapshots are enough?
Nyquist and Shannon, in plain words
The answer came from telecommunications theory decades before digital audio existed. Harry Nyquist's 1928 work on telegraph transmission, and Claude Shannon's later formal proof of the sampling theorem, established a result that still surprises people: if a signal contains no frequencies above some limit, then sampling it at more than twice that limit captures it completely. Not approximately — completely. The samples aren't a jagged, stair-step caricature of the wave; they are sufficient information to reconstruct the one and only band-limited waveform that passes through them.
The theorem cuts both ways. Sample at 48 kHz and you can perfectly represent everything up to 24 kHz — the Nyquist frequency. But let any energy above that limit reach the converter and it doesn't just disappear; it folds back down into the audible range as aliasing, a ghost tone at the wrong frequency. This is why every real converter has an anti-aliasing filter in front of it, removing content above Nyquist before sampling. Human hearing tops out around 20 kHz in the young and healthy, so a rate in the mid-40s of kilohertz covers hearing with a little room left for the filter to do its work. That arithmetic is why the standard rates land where they do.
44.1 versus 48: an accident and a decision
The odd-looking 44.1 kHz of the compact disc is a fossil of production practicality. In the late 1970s, before dedicated digital audio recorders were affordable, the way to store digital audio was a PCM adaptor: a box that encoded audio samples as pseudo-video so they could be recorded on a video tape machine. The sample rate therefore had to mesh with television line structure, and 44.1 kHz was a rate that could be carried neatly within both major TV standards' geometry. When Sony and Philips finalized the CD format, the rate used by the studio equipment mastering the discs became the rate of the medium, and consumer audio inherited 44.1 kHz for decades.
48 kHz, by contrast, was chosen on purpose. Professional film and video sound standardized on it partly because it divides cleanly against common frame rates — at 24 frames per second, each frame spans exactly 2,000 samples; at 25 fps, 1,920 — which makes locking sound to picture arithmetic instead of approximation. It also buys slightly more margin above the audible band for anti-alias filtering. Digital video formats, broadcast infrastructure, and eventually the entire film and video production world settled on 48 kHz, and AES recommendations endorse it as the default professional rate. Real-time communication stacks — the plumbing under Zoom, Meet and Teams — grew up in the same ecosystem and speak 48 kHz natively.
Why more kilohertz isn't more quality
If 48 kHz covers hearing, what do 96 or 192 kHz buy? For delivery — the file or stream a listener receives — essentially nothing except bandwidth. The sampling theorem says everything below Nyquist is already captured perfectly; doubling the rate extends the representable band into ultrasonic territory no listener perceives, while doubling the data. Higher rates have legitimate uses in production — some processing is more convenient with the extra spectral room, and relaxed filter requirements simplified early converter design — but "192 kHz sounds twice as detailed as 96" is not how the math works. A voice call carried at 48 kHz is not a compromised version of some higher-rate truth. For speech the case is even more lopsided: nearly everything that makes a voice intelligible and pleasant lives below 12 kHz, so 48 kHz already carries the entire signal with generous margin.
It's also worth separating sample rate from its frequent companion, bit depth. Sample rate decides how often you measure — which sets the frequency range. Bit depth decides how finely you measure each snapshot — which sets the dynamic range between the quietest representable detail and the loudest, the territory covered by dBFS metering and headroom. Sixteen bits give roughly 96 dB of range; 24 bits, far more than any acoustic chain delivers; 32-bit floating point adds enormous computational headroom, so intermediate processing stages effectively can't clip internally. "48 kHz, 32-bit float" answers two independent questions: how wide, and how deep.
One rate, once resampled
The practical sample-rate sin isn't picking the wrong number — it's changing numbers mid-chain. Every sample rate conversion is a real DSP operation with a real cost in filtering and latency, and a chain that hops between 44.1 and 48 at every stage pays that cost repeatedly. Well-built audio software does what broadcast facilities do: pick one house rate, convert incoming signals once at the boundary, and keep everything inside the building at that rate. For anything that touches calls or video, the house rate to pick is 48 kHz, because that's what the platforms at the end of the chain speak anyway.
In DeskBroadcast
DeskBroadcast runs its entire audio chain natively at 48 kHz in 32-bit float. Whatever rate your hardware happens to capture at is resampled exactly once, at the input boundary — from there, everything downstream operates at 48 kHz: noise reduction, EQ, the compressor, the AI model, and the virtual mic that hands the result to your call. The rate isn't just a convention here; it's structural. The DeepFilterNet 3 denoiser consumes frames of exactly 480 samples, which at 48 kHz is precisely 10 milliseconds of audio per frame — the model's timing and the chain's clock are the same arithmetic. One rate, one resample, no drift between what the app processes and what Zoom, Meet or Teams expects to receive.
Hear the chain, not the specs
Sample rates are invisible when they're right — what you can hear is everything running on top of them. DeskBroadcast's Mic Check records eight seconds of your voice and replays it raw versus processed, so you can judge the 48 kHz chain with your own ears.
Download DeskBroadcast