Skip to content
2xKit

How Silence Detection and Removal Works in Podcast Editing

The threshold and duration settings behind automatic silence removal, and why getting them wrong clips your breaths.

Quick answer

Silence detection works by scanning an audio track for stretches where volume stays below a set threshold, such as -40dB, for longer than a set minimum duration, such as 500 milliseconds, then flagging or automatically cutting those stretches. The two settings that matter most are the threshold (how quiet counts as silence) and the minimum duration (how long a quiet gap must last before it's removed), since setting either too aggressively clips natural pauses and breaths rather than genuine dead air. The Silence Detector and Silence Remover handle both steps.

Manually scrubbing through an hour-long podcast recording to cut out dead air is tedious enough that automatic silence removal has become a standard editing step, but the tools behind it are simpler than they seem, and understanding the two core settings they rely on is the difference between a natural-sounding edit and one that clips breaths and awkwardly chops sentences.

Threshold: how quiet counts as "silence"

Every silence detection tool needs a volume threshold, a decibel level below which audio is treated as silence rather than content. Set it too high (closer to 0dB, meaning it only counts truly dead air) and background hiss or a quiet breath won't register as silence, so nothing gets removed. Set it too low (very negative, like -60dB) and it becomes overly permissive, potentially treating quiet spoken words or trailing sentence endings as silence and cutting them out. A typical starting point around -35dB to -45dB catches genuine gaps between sentences while leaving quiet speech intact, though the right value depends heavily on the recording's own noise floor, checked easily with the Audio Spectrum Visualizer before choosing a threshold.

Minimum duration: avoiding choppy, unnatural cuts

The second setting, minimum duration, determines how long a quiet stretch has to last before it counts as removable silence rather than a normal conversational pause. Setting this too short (say, 100ms) removes brief natural pauses between words and phrases, making speech sound rushed and unnatural; a more typical setting around 400-800ms only targets genuinely long gaps, like a pause between distinct thoughts or an edited-out mistake. The Silence Detector identifies these stretches first so you can review them, and the Silence Remover then cuts them automatically, which is usually a safer workflow than blind automatic removal for anything you plan to publish. For manual review and precise trims around specific moments the automatic pass missed, Trim Audio handles those individually, and exporting chapter markers at natural topic breaks with Podcast Chapter Export is a common next step once the silence-trimmed episode is ready.

Frequently asked questions