Skip to main content
Skrrol logo
GUIDE

Audio Ducking Explained — Keep Music Under Your Voice, Automatically

Skrrol AI Editorial11 min read

Table of Contents

  1. 1.What audio ducking is (and what it is not)
  2. 2.Why mixes without ducking sound amateur
  3. 3.Manual keyframed volume vs automatic ducking
  4. 4.How Skrrol's mixer ducks music automatically
  5. 5.Dialing in the duck: amount, base levels, and pumping
  6. 6.Pair ducking with EQ so voice and music share the spectrum
  7. 7.A practical walkthrough: podcast episode, then a vlog
  8. 8.Common ducking mistakes (and the fixes)

Close your eyes during any well-produced video and listen to what the music does. The moment someone speaks, it steps back. The moment they stop, it swells to fill the space. That movement is called ducking, and it is one of the clearest dividing lines between a finished mix and two audio files playing at the same time. This guide explains what ducking is, why mixes without it read as amateur within seconds, how automatic ducking in Skrrol's mixer replaces an afternoon of volume keyframes, and how to tune the amount so the music breathes with the voice instead of lurching underneath it.

What audio ducking is (and what it is not)

Ducking is a volume relationship between two signals: a priority source, usually a voice, and a background source, usually music. Whenever the priority source is active, the background drops by a set amount. Whenever the priority source goes quiet, the background rises back to its normal level. Radio has run on this principle for decades: the presenter talks over the intro of a song, the song sits politely underneath, and it blooms back the instant the talking stops. Nobody is riding a fader by hand; a processor listens to the microphone and pushes the music down automatically.

Be precise about what ducking is not. It is not the same as simply setting the music quieter. A static low level means the music is just as timid when nobody is talking, which wastes the energy you added the track for. Ducking is dynamic: loud when it can be, quiet when it must be. It is also not a tone tool. EQ and noise reduction change what a sound is made of; ducking only changes how loud one sound is allowed to be while another is happening. The two jobs are complementary, which is why the second half of this guide pairs them.

Why mixes without ducking sound amateur

Unducked mixes fail in one of two directions. In the first, the music is set loud enough to carry real energy, and every spoken line has to fight through it. It might survive on good headphones; on the phone speaker where most short-form video is actually watched, the voice loses — consonants smear, quiet word endings vanish, and the listener has to work to follow the sentence. Nobody consciously diagnoses the problem. They just feel that the video is tiring and scroll on.

In the second direction, a creator burned by the first mistake drags the music fader down until the voice always wins. Now the dialog is clear but the track underneath is wallpaper. Intros, pauses, montage sections, the beat that should land after a punchline — all of it sits at the same apologetic volume. A ducked mix gets both halves: music at full presence in every gap, voice unchallenged the moment it starts. That rise and fall is most of what people are hearing when they describe a mix as produced.

Manual keyframed volume vs automatic ducking

The traditional way to build that movement is by hand. You add volume keyframes to the music clip — one to start the dip before a line begins, one at the ducked level, one to hold, one to recover after the line ends. Four keyframes per spoken phrase. A tight three-minute vlog might need forty of them. A forty-minute podcast episode has hundreds of speech gaps, which means a thousand or more keyframes, each placed by scrubbing to the phrase, guessing the ramp, and playing back to check.

The deeper problem is not the initial work, it is the maintenance. Volume keyframes are anchored to time, not to the dialog. Trim one line, ripple one cut, swap the order of two segments, and every keyframe downstream now dips in the wrong place. The mix you spent an afternoon riding is broken by a five-second edit, and the only fix is to ride it again.

Automatic ducking replaces the keyframes with a standing rule. In Skrrol's mixer, the music track is told to listen to the dialog track; whenever speech is detected, the music bed drops by the amount you chose, and when speech stops it lifts back cleanly. The duck is computed in real time from the dialog itself, so there are no volume keyframes to place and nothing to break when you re-edit — move a line and the dip moves with it. Manual keyframes still earn their place for one-off moments, like a swell you want to land on a specific beat, but ducking handles the systematic work.

How Skrrol's mixer ducks music automatically

The mixer will look familiar to anyone who has stood in front of a hardware console. Every audio track on your timeline appears in the Mixer panel as a channel strip with a fader, a pan knob, solo, mute, and a level meter, and the master bus carries a loudness meter that reads true peak and integrated LUFS. Setting up ducking takes three moves: right-click the music channel, choose Sidechain to Dialog, and set the duck amount — typically 9 to 12 dB. From that point on, the mixer listens to the dialog channel and pulls the music down by that amount every time someone speaks, then lifts it back the moment they stop.

The sidechain source is flexible. A solo creator ducks the music off a single voice track. A two-host podcast can sidechain the music off either dialog channel, or — cleaner — route both hosts into a dialog bus and sidechain the music to the bus, so the bed dips under whichever person is speaking. There is no fixed ceiling on track count; a modern laptop comfortably runs twenty or more stereo tracks with effects, enough for anything from a talking-head video to a fully layered short-film mix. And the whole mix is processed locally in your browser through the Web Audio API — your audio is never uploaded anywhere to be ducked.

Dialing in the duck: amount, base levels, and pumping

Two numbers control how ducking feels: the music base level and the duck amount. Get the base level right first, because the duck subtracts from it. A reliable starting point is dialog peaking between -12 dB and -6 dB, the music bed sitting at -18 to -24 dB, and sound effects balanced against the dialog by ear. From there, a 9 to 12 dB duck puts the music clearly underneath the voice during speech while leaving it obviously present in the gaps. If you duck from a base level that is too hot, you need a huge duck amount to rescue the voice, and huge duck amounts are what make a mix pump.

Pumping is the audible artifact where the music lurches down and rebounds hard enough that the listener notices the movement itself — the bed sounds like it is gasping between phrases. It has three usual causes. The duck amount is too large for the material. The dialog track is littered with non-speech noise — breaths, desk bumps, chair squeaks — so the duck triggers on and off in rapid succession. Or the base level is so hot that even a correct duck amount produces a dramatic swing. The fixes map one to one: reduce the amount, clean the dead air out of the dialog track so speech detection has honest material to follow, or pull the music fader down so a smaller duck does the job.

Pair ducking with EQ so voice and music share the spectrum

Ducking manages when the music is loud. EQ manages where in the frequency spectrum the music and the voice overlap, and the two together outperform either alone. Skrrol's parametric EQ gives any clip or track a multi-band equalizer with a visual frequency response curve, draggable band points, and per-band frequency, gain, and Q controls. Four bands are active by default and an instance can grow to twelve for surgical work. You can apply it to a single clip or across an entire channel, which is what you want for a dialog chain.

On the voice, the moves are well established. Cut everything below 80 Hz with a steep high-pass to remove room rumble, air handling, and mic handling noise. Cut 2 to 4 dB around 250 Hz if the voice sounds boxy or muddy. Boost 2 to 3 dB around 3 kHz to bring the voice forward, and lift a high shelf 1 to 3 dB from 10 kHz for air. The dialog clarity and podcast voice presets start you at a known-good curve, and the bypass toggle lets you A/B every change to confirm it actually helped.

On the music, make the mirror-image move. The presence region that makes a voice intelligible sits roughly between 2 and 5 kHz, so a modest, wide cut on the music track in that range clears the lane the voice needs most. The practical payoff is that you can duck less. When the bed no longer competes head-on with the voice in the same frequencies, a shallower duck keeps the dialog clear, the music stays fuller through spoken sections, and the whole mix moves less. Both the EQ and the duck render into the audio when you export, so what you hear in preview is what ships.

A practical walkthrough: podcast episode, then a vlog

Here is the full recipe applied to the most common case: a two-host podcast episode with an intro music bed and a few stings.

  1. Drop each recording on its own track — one per host, one for the music bed, one for stings. Every clip you add creates its own track.
  2. Open the Mixer tab. Each track appears as a channel strip with fader, pan, meter, mute, and solo.
  3. Set base levels before touching the duck: host mics peaking between -12 and -6 dB, the music bed at -18 to -24 dB, stings balanced against the dialog.
  4. Route both host channels into a dialog bus and sidechain the music to the bus, so the bed ducks under whichever host is speaking.
  5. Set the duck amount to 9 to 12 dB, then play the fastest back-and-forth stretch of conversation and listen for pumping. Reduce the amount if the bed audibly bounces.
  6. EQ each voice: high-pass at 80 Hz, a small cut around 250 Hz if the room made the voice boxy, a presence boost near 3 kHz — or start from the podcast voice preset and tune.
  7. Keep dialog panned center and place stings or ambience slightly off-center for width.
  8. Play the loudest stretch and watch the master meter: target -14 LUFS integrated for YouTube and streaming delivery, with the true-peak limiter at its -1 dBTP default ceiling. Then export — the mix renders into the finished video.

A vlog is the same recipe with fewer tracks. One voice track — often the camera audio — and one music bed. Sidechain the bed to the voice and duck around 9 dB. The structure of a vlog is what makes ducking shine: talk-to-camera segments alternate with b-roll montages, and the duck gives each mode what it needs — full music over the montage, clear voice the moment you cut back. Because vlog audio is usually recorded in uncontrolled locations, the 80 Hz high-pass earns its keep twice over, cutting the traffic rumble and air-conditioning hum that location microphones soak up.

Common ducking mistakes (and the fixes)

Ducking is one of those techniques that costs almost nothing once you know it exists and pays out on every project afterward. Set the base levels, sidechain the bed to the dialog, choose an amount the ear cannot catch, carve the presence region, and check the meter on the way out. Do it once and you will hear the absence of it everywhere — in other people's videos, and never again in yours.

Related Features

Skrrol capabilities that connect to this article

Related Use Cases

Workflows and industries that use this technique

Frequently Asked Questions

What is audio ducking in video editing?

Ducking automatically lowers a background track — usually music — whenever a priority signal, usually a voice, is active, then restores it when the voice stops. It keeps dialog clear without flattening the music in the gaps. In Skrrol's mixer it is a sidechain: the music channel listens to the dialog track and dips by a configurable amount in real time.

How much should music duck under a voice?

Start at 9 to 12 dB, with the music bed at a base level of -18 to -24 dB and dialog peaking between -12 and -6 dB. Tune from there by ear: if the voice fights, duck more or lower the bed; if you can hear the music audibly diving, duck less.

Can ducking follow more than one speaker?

Yes. Sidechain the music off any dialog channel, or route all voice tracks into a dialog bus and sidechain the music to the bus — the bed then ducks under whichever speaker is talking.

Do I still need volume keyframes if I use automatic ducking?

Not for the routine dips. The duck is computed in real time from the dialog, so it needs no keyframes and survives re-edits. Keep manual keyframes for deliberate one-off moments, like a musical swell you want to land on a specific beat.

Why does my ducked music sound like it is pumping?

Either the duck amount is too large, the music base level is too hot, or the dialog track is full of non-speech noise that triggers the duck erratically. Reduce the amount, pull the bed fader down, and clean the dead air out of the dialog track.

Is my audio uploaded to a server for ducking?

No. The entire mix, ducking included, runs locally in your browser through the Web Audio API — your audio never leaves your device.

Related Articles

Start Creating Today

Try Skrrol AI's free studio now and bring your creative vision to life