
EQ for speech intelligibility works best when you remove what hides consonants (rumble, mud, boxiness, harsh peaks) before you add “clarity.” In practice, that usually means a high-pass filter plus one or two small cuts, and only a gentle presence lift if the voice still feels veiled.
Start with the smallest set of moves that can possibly work
If you only have time for one principle: cut before you boost, and change less than you think. Speech intelligibility is mostly about keeping the midrange clean enough that consonants stay audible at normal listening levels. Big boosts often sound impressive for a moment, but they also raise noise, sibilance, and listener fatigue.
A minimal-intervention approach is a short loop:
- identify the single biggest problem you hear,
- fix it with one EQ move,
- re-check intelligibility at a realistic playback level,
- stop as soon as words are easy to understand.
Step 1: High-pass filter to remove rumble (nearly always the first move)
Most spoken-word recordings carry low-frequency energy that adds nothing to intelligibility: mic handling, desk vibration, HVAC, traffic, footfalls, proximity effect. That energy steals headroom and can mask the low mids.
- Typical starting point (adult voice): high-pass around 70–100 Hz.
- If the voice is very deep or you want more warmth: start lower (e.g., 60–80 Hz).
- If it’s a thin headset mic or already bright: don’t force a high cutoff; keep it conservative.
Use a gentle slope if your EQ offers it. The goal is not “thin,” it’s “clean.”
Step 2: One cut that makes the words pop (mud and boxiness zones)
After rumble, the most common intelligibility killer is low-mid build-up that makes speech feel “covered,” “boxy,” or “muffled.” Two ranges matter most:
A) “Mud” and bloom: roughly 120–300 Hz
Too much here makes the voice sound thick and indistinct, especially on phones and laptop speakers.
- Try a small cut: 1–3 dB with a moderate Q (not razor-thin).
- Move the center frequency until the voice stops feeling “cloudy.”
B) “Boxy” / “roomy” tone: roughly 300–600 Hz
This is where small rooms, reflective desks, and cheap mic placement show up. Cutting a little here can make consonants feel more separated without adding brightness.
- Again: 1–3 dB is often enough.
- If the recording sounds “hollow” after your cut, you went too far or too wide.
Minimal intervention means you choose one of these (or the smallest change that fixes both), not an aggressive smile-curve.
Step 3: Don’t “EQ for clarity” until you’ve removed masking
Many people jump straight to boosting “clarity” (upper mids) and end up with a sharp, spitty voice that is still hard to understand—just louder in the wrong places. If the voice is muddy, presence boosts won’t fix the underlying masking; they mainly highlight mouth noise and sibilants.
A good check: after your low/low-mid cleanup, listen again. If the words are now understandable, stop. If they’re understandable but still feel slightly veiled, then consider a tiny presence lift.
Step 4: Gentle presence, not hype (where intelligibility lives)
A lot of speech information that helps recognition sits in the midrange, particularly around 1–2 kHz, with “presence” and edge often perceived higher than that. Boosting too high too fast makes sibilance and harshness, not intelligibility. Shure’s guidance is a useful reality check here: intelligibility is not simply “more highs.” (service.shure.com)
Practical approach:
- If the voice feels dull after cleanup, try +1 to +2 dB somewhere in 2–4 kHz (wide, gentle).
- If it becomes edgy or fatiguing immediately, undo it and look for a narrow harsh spot to cut instead (next section).
Think of presence as seasoning. You should barely notice the EQ as an effect—you should just notice that words land more easily.
Step 5: Harshness and “nasal” resonances (cutting is safer than boosting)
Speech often has one or two narrow resonances that jump out on certain syllables. Removing a small peak can improve clarity more than boosting.
“Nasal,” honky, or megaphone tone: roughly 800 Hz–1.2 kHz
- Use a narrower cut (a bit higher Q than your mud cut).
- Reduce 1–4 dB, just enough that the voice sounds natural.
“Bite,” glare, or painful consonants: roughly 2.5–5 kHz
This range overlaps perceived clarity, which is why boosting here is risky. If you already boosted presence and it hurts, reverse the boost and instead find the specific sharp frequency:
- Sweep a narrow band gently (don’t crank it) to locate the painful spot.
- Cut that spot slightly.
Minimal intervention often means you replace one broad “clarity boost” with one small corrective cut.
Step 6: Sibilance is not intelligibility (5–8 kHz)
“S” and “sh” live here. Too much energy in this band can make speech tiring and, paradoxically, less intelligible because the sibilants dominate.
With EQ alone:
- If “S” is aggressive, try a tiny cut around 6–8 kHz (narrow-ish), 1–2 dB.
- If the whole voice becomes dull, you cut too wide.
If sibilance is severe, EQ can help a bit, but it’s easy to overdo. Minimal intervention means you aim for “not distracting,” not “gone.”
Step 7: Match loudness when you A/B, or you’ll fool yourself
EQ changes often increase perceived loudness. Louder usually seems “clearer,” even if it’s not. When you compare before/after:
- Lower the output level of the EQ (or your channel fader) so the after is roughly the same loudness as the before.
- Then decide if intelligibility truly improved.
This single habit prevents most over-EQ.
Step 8: Use the right listening test: small speaker, low volume, real words
Intelligibility problems hide on big speakers at a comfortable level. To test whether your minimal changes worked:
- Listen quietly (just above a whisper level).
- Listen on a phone/laptop speaker or a single small monitor.
- Pay attention to consonants at the ends of words and to similar-sounding words (“fifty” vs “sixty,” “can” vs “can’t”).
If words stay understandable under those conditions, you’re done. If they don’t, don’t add more boosts—look for another masking cut.
A minimal “default chain” you can try (and then stop)
This is not a preset; it’s a starting structure that stays minimal:
- High-pass: 70–100 Hz (adjust by voice).
- One corrective cut: either 150–300 Hz (mud) or 300–600 Hz (boxy), 1–3 dB.
- Optional gentle presence: +1–2 dB at 2–4 kHz only if needed.
- Optional narrow cut for harshness (2.5–5 kHz) or sibilance edge (6–8 kHz) if it’s distracting.
If you find yourself adding step 5, 6, and 7, you’ve left “minimal intervention” territory and should reassess the recording or mic placement instead of stacking more EQ.
Why does this matter
Better intelligibility with minimal EQ means speech stays natural, fatigue stays low, and your audio translates across phones, cars, and earbuds without sounding “processed.”
Sources
- Shure — Equalization and Speech Intelligibility (service.shure.com)
- Audacity Manual — Filter Curve EQ (manual.audacityteam.org)
- RØDE — A Guide to Audio Processing and FX For Podcasting (rode.com)