Workflow · the bar itself
Course video audio standards: what platforms actually expect from your narration
Audio forums will tell you a course needs studio-grade sound. Course platforms tell you something humbler and far more useful: it needs to be unnoticeable. Here's the published bar, what reviewers actually reject for, and how each requirement translates into a recording-side move you make before the file exists.
Platform requirements below were checked against Udemy's published instructor documentation on 20 September 2026 (paraphrased — republishing their checklists verbatim isn't our place). Most other marketplaces keep detailed review criteria behind instructor logins, so we generalize only where we could read the source.
The published bar, in plain terms
Udemy — the marketplace with the most public paper trail — reviews every course before it goes live, and its audio expectations reduce to five checks. Read them slowly, because what's absent matters as much as what's there:
- Both channels
- In sync
- No distracting noise
- No distortion
- Consistent volume
- Audio comes out of both channels. Not "in stereo," not "spatially immersive" — just present in both ears. Single-channel narration (a classic recorder misconfiguration) is an explicit review problem.
- Audio is in sync with the video. A capture-pipeline concern: drift creeps in through mismatched sample rates and exotic device chains, and it's miserable to repair afterwards.
- No distracting background noise. Note the adjective. The standard is not silence; it's that a learner's attention never leaves the lesson. Steady faint room sound passes; the intermittent lawnmower doesn't.
- No distortion. Udemy's instructor guidance singles out gain-too-high static as one of the most damaging and common audio faults — the exact failure our levels ritual exists to prevent, and the one no software un-bakes.
- Reasonable, consistent volume. Learners set their volume in lesson one and expect lesson nine to respect it.
What's absent: any demand for expensive capture, any named noise floor number, any specific tooling. The platforms are measuring the learner's experience, not your equipment list. That should relax you — and then focus you, because every one of the five checks is decided at recording time.
Translated into recording-side moves
| Review check | Where it's decided | The recording-side move |
|---|---|---|
| Both channels carry audio | Recorder/export config | Record one test lesson start-to-finish and play it in headphones, both ears, before recording lesson two. Left-only narration is a settings bug you catch in minute one or inherit for a whole course. |
| Sync holds | Capture pipeline | Keep the chain boring: one recorder, standard sample rate, no daisy-chained virtual devices you don't need. Run a long test — sync problems hide in minute twenty, not minute one. |
| No distracting noise | The room + the live chain | Stop noise before the file: OBS's live filter or a suppression layer under Loom/Camtasia/ScreenFlow, and a session scheduled around your building's loud hours. |
| No distortion | Input gain, before REC | Speech peaks at −12…−6 dB, set at true narration volume, re-checked each session. Thirty seconds; non-negotiable. |
| Consistent volume across lessons | Session discipline | Same mic distance, same gain, same chain, every session — write the numbers down. (This is also the argument for recording a module in scheduled sessions rather than heroic scattered bursts.) |
The desk's numbers, for people who want numbers
Since platforms decline to print figures, here are the ones this desk works to — conventions, clearly labeled as ours, chosen to clear every published check with margin:
- Peaks: −12 to −6 dB on the recorder's meter, ceiling untouched. Headroom is what makes "no distortion" a guarantee instead of a hope.
- Loudness feel: match your narration level against a couple of major-platform tutorial videos at the same system volume — if yours is dramatically quieter or louder, fix gain now, at the source. (Formal loudness targets are an export-time craft; recording with consistent levels is what makes any target reachable later without damage.)
- Noise test: headphones at normal volume, eyes closed, ten seconds of your take's pauses. If you can name what you hear ("fan," "fridge," "street"), a learner eventually will too — treat it at the chain, not in post.
- Continuity kit: thirty seconds of room tone per session, so every future edit preserves the consistency the checks reward.
Accessibility is part of the audio standard
Platforms increasingly expect captions, and clean narration is quietly the input that makes them cheap: auto-captioning services transcribe a clear, noise-free voice dramatically better than a muddy one, which means less correction time per lesson. Clean audio isn't just the listening experience — it's the metadata pipeline too.
Self-hosted courses: whose standard applies now?
Selling from your own site with no reviewer between you and the learner? The checklist above still applies — it just arrives as refund requests instead of review notes. If anything, hold the bar higher: marketplace learners blame the platform for rough audio; your own customers blame you. The five checks plus the desk numbers make a fine self-review: run them against your first exported lesson before you record twenty more.
When to stop improving
The standards above are a finish line, not a starting block. If your test lesson passes all five checks in honest headphones, your audio is done — further polish is procrastination wearing headphones. The learners came for the lesson; the audio's whole job is to never remind them it exists. Get there the cheap, permanent way: the recording-side signal chain, one-pass narration, and the ninety-second ritual before every red light.