How I generate first-draft captions with Whisper and clean them up
← Blog
🎬 Captioning & Subtitles8 min read

How I generate first-draft captions with Whisper and clean them up

💡 To get a first-draft SRT from any audio or video file with Whisper, run: whisper audio.mp4 --model turbo --language en --output_format srt --word_timestamps True --max_line_width 42 --max_line_count 2 --output_dir ./captions/. Whisper produces a usable .srt in under a minute on most machines. You still need a cleanup pass in Subtitle Edit: check overlapping timecodes, CPS violations, and proper-noun spelling before delivering.

Key takeaways

  • Use --model turbo for English: it runs ~8x faster than large with comparable accuracy for most content. For multilingual or low-resource language work, use --model large-v3.
  • Always set --word_timestamps True --max_line_width 42 --max_line_count 2 so Whisper breaks lines at word boundaries with a sensible line length instead of dumping a whole segment on one line.
  • Five cleanup checks after every Whisper run: overlapping timecodes, CPS violations, proper-noun errors, hallucinated lines on silence, and missing sentence-end punctuation.
  • Vietnamese content requires more than ASR cleanup: Whisper misses tone marks frequently, so I use its output as a timing scaffold and re-type the dialogue in Subtitle Edit.
  • Whisper output is a first draft. For broadcast, streaming, or ADA/CVAA compliance, a human review pass is still required.

Which Whisper model should I start with?

Whisper ships with six model sizes. Here is how I choose between them:

ModelParametersSpeedMy recommendation
tiny39 M~10xQuick local tests only
base74 M~7xVery short clips, clear audio
small244 M~4xCreator drafts where speed matters most
medium769 M~2xBalanced choice for most content
large-v31550 M1x baselineBest accuracy; multilingual or non-English
turbo809 M~8xMy default for English content

For the three smallest sizes, Whisper also provides English-only variants (tiny.en, base.en, small.en) that perform better on English-only audio. I skip these and go straight to turbo or medium for any content I am delivering to a client.

The turbo model is not trained for translation tasks. If you need to transcribe non-English audio and produce an English translation in one step, use medium or large-v3 with --task translate.

The command I use for a first-draft SRT

This is my standard Whisper command for English creator content:

whisper audio.mp4 \
  --model turbo \
  --language en \
  --output_format srt \
  --word_timestamps True \
  --max_line_width 42 \
  --max_line_count 2 \
  --output_dir ./captions/

What each flag does:

  • --model turbo: ~8x faster than large with comparable English accuracy. Not suitable for translation tasks.
  • --language en: Skips the 30-second language detection step. Always specify the language if you know it. Without this flag, Whisper reads the first 30 seconds and guesses, which fails on content that starts with music or silence.
  • --output_format srt: Produces a standard .srt file. Other options: vtt (WebVTT), txt (plain text), tsv (tab-separated with timestamps), json (full data including word-level timing), all (all five formats at once).
  • --word_timestamps True: Attaches timing to each word internally. Required for --max_line_width and --max_line_count to work correctly.
  • --max_line_width 42: Maximum characters per line. Without this flag, Whisper produces long single-line cues. I use 42 for standard players.
  • --max_line_count 2: Maximum lines per subtitle event. 2 is the standard for captions and most subtitle work.
  • --output_dir ./captions/: Where to write the output. Defaults to the current directory if omitted.

What does --word_timestamps True actually add?

Without --word_timestamps True, Whisper writes one subtitle event per audio segment: roughly one sentence or breath. Those events often run 80 to 100 characters on a single line, far above the 42-character-per-line standard.

With --word_timestamps True, Whisper attaches individual timing to each word. The --max_line_width and --max_line_count flags then use those word-level timestamps to split long segments into shorter, properly timed events at word boundaries.

There is also --highlight_words True, which wraps the currently-spoken word in italics in the SRT output. It is a useful starting point for karaoke-style word-by-word highlight captions on social content:

whisper audio.mp4 \
  --model turbo \
  --language en \
  --output_format srt \
  --word_timestamps True \
  --highlight_words True \
  --max_line_width 42 \
  --max_line_count 2 \
  --output_dir ./captions/

The highlight_words output needs cleanup, but it saves time compared to adding word-level timing manually in a caption editor.

Where does Whisper go wrong - and how do I catch it?

Whisper is good at transcription but it is not error-free. These are the five categories I check on every job:

Error typeWhy it happensHow I catch it
Proper nouns and product namesDomain-specific terms outside the training distributionCtrl+F for names from the brief; use --initial_prompt with expected terms
Hallucinated text on silenceModel generates plausible text during near-silence or musicListen through the output; add --hallucination_silence_threshold 2.0
Missing sentence-end punctuationInconsistent punctuation across segment boundariesFix common errors in Subtitle Edit (Tools > Fix common errors)
CPS violationsShort-duration events from word-level splitsTools > List errors (Ctrl+F8) in Subtitle Edit
Vietnamese diacritics stripped or wrongVietnamese is a low-resource language with a complex tonal systemUse Whisper timing only; re-type dialogue in Subtitle Edit

For content with long pauses or ambient-music sections, I add --hallucination_silence_threshold 2.0. This tells Whisper to skip periods of near-silence longer than 2 seconds instead of inventing text to fill them:

whisper audio.mp4 \
  --model turbo \
  --language en \
  --output_format srt \
  --word_timestamps True \
  --max_line_width 42 \
  --max_line_count 2 \
  --hallucination_silence_threshold 2.0 \
  --output_dir ./captions/

My Subtitle Edit cleanup pass

Once Whisper produces the .srt, I open it in Subtitle Edit and work through the same five steps every time:

  1. Fix common errors (Tools > Fix common errors): I enable "overlapping lines" (adds a 1-frame gap between adjacent cues), "remove empty lines", and "fix double spaces". This catches most Whisper timing artifacts in under 10 seconds.
  2. CPS check (Tools > List errors, Ctrl+F8): Subtitle Edit highlights any event above my target CPS limit. I use 20 CPS for most creator content, 17 for children's programming. Word-level splits with short durations show up here.
  3. Name pass: I Ctrl+F for names and brand terms from the project brief. Whisper transcribes proper nouns phonetically, which produces subtle errors that spellcheck misses.
  4. Silence scan: I play through at 1.5x speed and listen for lines that start on silence or near-silence. Those are almost always hallucinated lines that need to be deleted.
  5. Punctuation pass: Whisper sometimes ends a cue mid-sentence with no punctuation, making the next cue read awkwardly. I read through the text view and add missing periods and commas.

For clean English interview content, this cleanup takes roughly one minute per ten minutes of source audio. For content with heavy accents or technical terminology, allow three to four minutes per ten minutes. For Vietnamese, I essentially re-caption from scratch using Whisper output only as a timing reference.

When does professional captioning make more sense?

Whisper plus Subtitle Edit cleanup is a legitimate workflow for creator content: a YouTube video, a podcast, a social reel with one clear speaker. The bar is higher for broadcast, streaming platforms, and accessibility compliance.

The point where I recommend bringing in a human captioner:

  • Broadcast and streaming deliverables requiring SCC or SMPTE-TT: Whisper cannot produce these formats natively, and the conversion requires a QC pass that a human should run anyway.
  • Content subject to ADA, CVAA, or Section 508: FCC caption quality rules require captions that are accurate, synchronous, complete, and properly placed. Whisper output, even after cleanup, is not reliably accurate enough for legal compliance without a human review pass.
  • Multiple speakers with overlapping or accented speech: ASR accuracy drops significantly here; manual captioning is often faster than correcting bad ASR output.
  • Vietnamese or other diacritic-heavy languages: I offer Vietnamese subtitle work as a human service, not an ASR one. If you need accurate Vietnamese subtitles, take a look at my Vietnamese subtitle translation service.

If you need a clean English SRT for a YouTube video or a soft-subtitle track for social, Whisper plus cleanup is the right starting point. If you need broadcast compliance or legal accuracy, start with a human captioner.

FAQ

Can Whisper transcribe Vietnamese correctly?

In my testing, the large-v3 model follows the general meaning of Vietnamese speech but misses tone marks frequently and produces words that are phonetically close but semantically wrong. Vietnamese is tone-sensitive: one wrong diacritic changes the meaning entirely. For Vietnamese content where accuracy matters, I use Whisper only for initial timing and re-type the dialogue manually in Subtitle Edit.

What is the difference between --max_line_width and --max_words_per_line?

--max_line_width limits characters per line (I use 42). --max_words_per_line limits words per line instead. I prefer the character limit because different languages have very different word lengths. A 10-word limit that works for English is far too long for German compound words or too short for certain Vietnamese structures.

Should I set --condition_on_previous_text to False?

The default is True: Whisper feeds the previous segment as context for the next window, which helps with consistent terminology and punctuation. The downside is that hallucinated words in one segment can carry into the next. I leave it True for clips under 10 minutes. For long recordings with frequent silent sections, I sometimes set it to False to reset the context at each window.

What does --initial_prompt do and when should I use it?

--initial_prompt pre-seeds the decoder with context text. I use it when the content includes a specific name or brand that Whisper keeps misspelling. For example: --initial_prompt "Transcript of a DaVinci Resolve Fairlight tutorial." The model then tends to spell those terms correctly throughout the run. Keep the prompt under 200 words.

Why does the same subtitle line appear twice in the output?

A known issue in original Whisper (not Faster Whisper or WhisperX variants): consecutive segments can overlap when the attention window shifts, producing duplicate events. Running Fix common errors in Subtitle Edit catches this under the "same text" category. Setting --condition_on_previous_text False also reduces repetition, with the accuracy trade-off noted above.

Official Sources

Written by Dao Huy (Lucas), Vietnamese translator & localization specialist (EN · ZH · FR → Vietnamese). See translation services →

Get QuoteWhatsApp