Listen to any radio show: the music bed sits comfortably loud… until the host speaks, and it politely steps back — then swells again in the pauses. Nobody is riding a fader by hand. That move is called ducking, and it's the single production trick that separates "someone talking over a song" from "a produced show."
Voice and music share the same critical frequencies — roughly 1-4 kHz, where consonants live and intelligibility is decided. Set the music loud enough to feel present, and it masks the words; set it low enough to never interfere, and it disappears entirely. A static volume cannot win this fight, because the right level changes every second depending on whether someone is speaking.
Ducking solves it dynamically: the voice acts as a trigger, and the music's volume follows — down when you talk, up when you breathe.
Every ducker, hardware or browser, turns the same three knobs:
The sweet spot for speech beds is -12 to -18 dB. Less and the words swim; more and the music vanishes instead of stepping back.
Fast enough that your first word isn't drowned (~50-150 ms). Too fast sounds like a gate slamming.
The musicality lives here: 0.5-1.5 s lets the music breathe back in during pauses without pumping on every comma.
AudioKit's Podcast tool does the whole radio move without a DAW:
Half of good ducking is the bed itself: instrumentals only (two voices at once is chaos, even when one is sung), steady arrangements without huge dynamic jumps, and mid-heavy songs duck deeper than airy ambient pads — if your bed lives in the same octaves as your voice, increase the amount by 3-4 dB.
As a rule of thumb: if you notice the music while listening to the words, it's too loud; if you forget it exists during pauses, it's too quiet. Numerically: bed peaks around -18 dB under speech works for most voices.
A static lower level works for background ambience. The moment your music has a role — an intro that swells, transitions, energy between segments — ducking is what makes it possible to have both presence and clarity.
-16 LUFS (stereo) is the accepted target across podcast platforms. Louder gets turned down; quieter makes listeners reach for the volume in traffic.