AudioKit 5 min read Podcast & voice

Listen to any radio show: the music bed sits comfortably loud… until the host speaks, and it politely steps back — then swells again in the pauses. Nobody is riding a fader by hand. That move is called ducking, and it's the single production trick that separates "someone talking over a song" from "a produced show."

Why your voiceover fights your music

Voice and music share the same critical frequencies — roughly 1-4 kHz, where consonants live and intelligibility is decided. Set the music loud enough to feel present, and it masks the words; set it low enough to never interfere, and it disappears entirely. A static volume cannot win this fight, because the right level changes every second depending on whether someone is speaking.

Ducking solves it dynamically: the voice acts as a trigger, and the music's volume follows — down when you talk, up when you breathe.

The three settings that matter

Every ducker, hardware or browser, turns the same three knobs:

Amount (how far it dips)

The sweet spot for speech beds is -12 to -18 dB. Less and the words swim; more and the music vanishes instead of stepping back.

Attack (how fast it dips)

Fast enough that your first word isn't drowned (~50-150 ms). Too fast sounds like a gate slamming.

Release (how slowly it comes back)

The musicality lives here: 0.5-1.5 s lets the music breathe back in during pauses without pumping on every comma.

Do it automatically, in your browser

AudioKit's Podcast tool does the whole radio move without a DAW:

  1. Drop your voice track and your music (MP3, WAV…).
  2. Set the ducking amount — the tool detects speech and rides the music for you.
  3. Preview, adjust, export — one mixed file, ready to publish.
Two companion moves make it shine: run your voice through Enhance Voice first if it was recorded on a laptop mic (the ducker follows a clean voice more accurately), and finish with the LUFS Normalizer at -16 LUFS — the standard loudness for podcast platforms.
Mix your voice over music, automatically Speech detection rides the music for you — free, in your browser.
Open Podcast tool →

Choosing music that ducks well

Half of good ducking is the bed itself: instrumentals only (two voices at once is chaos, even when one is sung), steady arrangements without huge dynamic jumps, and mid-heavy songs duck deeper than airy ambient pads — if your bed lives in the same octaves as your voice, increase the amount by 3-4 dB.

FAQ

How loud should the music be when I'm speaking?

As a rule of thumb: if you notice the music while listening to the words, it's too loud; if you forget it exists during pauses, it's too quiet. Numerically: bed peaks around -18 dB under speech works for most voices.

Ducking or just lowering the music once?

A static lower level works for background ambience. The moment your music has a role — an intro that swells, transitions, energy between segments — ducking is what makes it possible to have both presence and clarity.

What loudness should the final podcast be?

-16 LUFS (stereo) is the accepted target across podcast platforms. Louder gets turned down; quieter makes listeners reach for the volume in traffic.

Related tools

Written by AudioKit — free audio tools that run in your browser.