How to Extract Embedded CEA-608/708 Captions from MP4, TS and MXF
← Blog
🎬 Captioning & Subtitles7 min read

How to Extract Embedded CEA-608/708 Captions from MP4, TS and MXF

💡 ccextractor input.mp4 -o out.srt extracts embedded CEA-608 captions from an MP4, MPEG-TS or MXF file as an SRT. Run ffprobe first to see what streams a file carries. For MXF, pass --input mxf. To split Field 1 and Field 2 or target a CEA-708 service, use --output-field and --service.

Key takeaways

  • CCExtractor 0.96.5 auto-detects MP4, MPEG-TS and MXF; force input with --input mxf when auto-detection fails on older file wrappers.
  • Probe streams first: ffprobe -v error -show_entries stream=index,codec_type,codec_name,codec_tag_string -of default input.mp4 reveals separate data tracks (codec c608 or eia_608); TS and MXF embed captions in the video SEI so no separate track appears.
  • CEA-608 Field 1 carries primary English captions (default). Use --output-field 2 for Field 2, which often holds a second-language caption track.
  • CEA-708 DTVCC services: --service 1 for the primary language service, --service 1,2 for both primary and secondary.
  • Default output is SRT. Use --out scc to keep the native CEA-608 encoding when you need a broadcast-ready file for re-mastering or QC.

Checking for embedded captions before you extract

Before I run CCExtractor on any file, I check what the container actually holds. The method differs by container type. For MP4 and MOV, ffprobe lists all streams including any separate data tracks. Run this on any MP4:

ffprobe -v error   -show_entries stream=index,codec_type,codec_name,codec_tag_string   -of default   input.mp4

What each flag does:

  • -v error: suppress informational output so only stream data prints.
  • -show_entries stream=...: list only these fields per stream.
  • codec_tag_string: shows the FourCC. A data stream tagged c608 or eia_608 is a separate CEA-608 track embedded at the MP4 container level.
  • -of default: plain key=value output, easy to read without extra tools.

For MPEG-TS and MXF, CEA-608 and CEA-708 data rides inside the video stream as H.264 SEI NAL units. ffprobe will not show a separate caption stream, and that is expected. If ffprobe shows only video and audio streams in a TS or MXF, that does not mean captions are absent: run CCExtractor and check its log output for detected caption data.

The basic CCExtractor command for MP4, TS and MXF

The simplest extraction converts CEA-608 captions to SRT. CCExtractor auto-detects the input format for MP4 and TS. For MXF I always force the input type because older MXF wrappers can trip the auto-detection:

# MP4 or MOV: auto-detect, output SRT
ccextractor input.mp4 -o out.srt

# MPEG-TS: auto-detect works reliably
ccextractor input.ts -o out.srt

# MXF: force input format to avoid auto-detection failures
ccextractor --input mxf input.mxf -o out.srt

# Output SCC instead of SRT (preserves native CEA-608 encoding)
ccextractor --out scc --input mxf input.mxf -o out.scc

Flag notes:

  • -o out.srt: output file path. Without -o, CCExtractor writes a file named after the input with an automatic suffix.
  • --input mxf: forces the MXF reader. Accepted values: ts, mp4, m2ts, mkv, mxf, scc.
  • --out scc: output format. Other options: ass, webvtt, webvtt-full, sami, ccd, bin.

CCExtractor logs a summary at the end of every run. Watch for "Extracted X frames" and lines mentioning EIA-608 or CEA-708 data. If both caption counts are zero, the file carries no detectable caption data in the streams CCExtractor reads.

How do I target a specific field, channel or CEA-708 service?

CEA-608 carries caption data in two fields and two channels per field. CEA-708 carries numbered DTVCC services. By default, CCExtractor extracts both 608 and 708 data when present, producing two output files. The flags below let you be precise:

TargetFlagWhen to use
608 Field 1 only--output-field 1Primary English captions (the default when no field flag is given)
608 Field 2 only--output-field 2Secondary captions, often a second-language track
Both 608 fields--output-field bothWrites two files with _1 and _2 suffixes
Channel 2 within a field--cc2Channel 2 of the active field
708 Service 1--service 1Primary language DTVCC service
708 Services 1 and 2--service 1,2Primary and secondary DTVCC services
All 708 services--service allExtract every DTVCC service present

Examples:

# Extract 608 Field 2 as SRT
ccextractor --output-field 2 input.ts -o field2.srt

# Extract 608 Field 1, channel 2
ccextractor --output-field 1 --cc2 input.mp4 -o channel2.srt

# Extract CEA-708 Service 1 only
ccextractor --service 1 input.mp4 -o service1.srt

# Extract both 608 fields and both 708 services
ccextractor --output-field both --service 1,2 input.mxf -o captions

Most US broadcast material uses 608 Field 1, Channel 1 for English captions. Field 2 sometimes carries a Spanish caption track. CEA-708 Service 1 is the primary language service in HDTV digital captions. When I am not sure what a file carries, I extract with --output-field both --service all first and check what output files are created.

What output formats can I get from CCExtractor?

The default is SRT, which covers most editing and upload needs. The --out flag switches formats. My usual choices and when I reach for each:

Format flagOutput fileUse case
--out srt (default)SubRip (.srt)Edit in Subtitle Edit, upload to YouTube, translation workflow
--out sccScenarist Closed Caption (.scc)Preserve 608 encoding; re-master or QC-decode with caption-inspector
--out webvttWebVTT (.vtt)Web video player, HTML <track> element
--out webvtt-fullWebVTT with stylingPreserves color and position metadata where present
--out assSubStation Alpha (.ass)Styled subtitles for playback or burn-in with ffmpeg
--out ccdCCD disassembly (.ccd)Low-level 608 byte inspection

After extracting to SCC, I QC-decode with caption-inspector to verify the text and positioning before delivering:

docker run --rm -v "$(pwd)":/data nebulabroadcast/caption-inspector   -o /data /data/out.scc

This writes a -C1.608 text decode file. Reading it confirms the extracted text matches what I expect, catching encoding problems that happen during extraction before they reach the client.

When should I hand caption extraction to a specialist?

CCExtractor is free and handles the common cases well. I use it as part of every QC workflow. The extraction command is only the starting point in these situations:

  • Broadcast deliverable with QC sign-off required: extracting an SRT from a client MXF is quick, but if the extracted captions need correction, re-timing and delivery back as an SCC or SMPTE-TT file for broadcast ingest, that is a full caption QC job.
  • CEA-708 DTVCC service content: CCExtractor's 708 decoder is documented as a work in progress. For a delivery that depends on the 708 service layer, I verify against caption-inspector and may need manual correction on window styling.
  • MXF with ancillary VBI data: some broadcast MXF files embed caption data in ancillary data packets rather than video SEI. CCExtractor may not find these. A broadcast caption tool with dedicated MXF ancillary reading is needed.
  • Extraction plus translation: extracting the timing is step one. Producing a correct translated subtitle for a localised delivery is a separate task with its own CPS checks and QC pass.

If you need English captions extracted, corrected and delivered as a broadcast or streaming file, or Vietnamese subtitles produced from the extracted timing, my caption QC and delivery service covers both steps end to end.

FAQ

CCExtractor says it found no captions. What do I check first?

First confirm the video codec is H.264 or MPEG-2, the most common carriers for 608/708 SEI data. Then try with --output-field both --service all to ensure you are not filtering out the active field or service. If CCExtractor still finds nothing, the file may have no embedded closed captions, or captions may be in MXF ancillary data outside the video stream, which CCExtractor does not decode.

Can CCExtractor extract subtitles from an MKV file?

CCExtractor reads MKV containers and looks for CEA-608/708 data in the video SEI. If the MKV has a separate SubRip or SubStation Alpha soft subtitle track instead of embedded 608/708 data, CCExtractor will not find those. For soft subtitle tracks, use ffmpeg directly: ffmpeg -i input.mkv -map 0:s:0 out.srt extracts the first subtitle track without any re-encode.

How do I know which field carries the English captions?

Run CCExtractor with --output-field both --service all first and check which output files contain readable English text. US broadcast material almost always uses CEA-608 Field 1, Channel 1 for primary English captions. CEA-708 Service 1 is the primary HDTV digital service. Field 2 sometimes carries a second-language track such as Spanish.

The extracted SRT has garbled characters. What went wrong?

CEA-608 uses a custom Latin character set. CCExtractor converts it to UTF-8 by default, but a non-standard or incorrectly flagged character set in the source can produce garbled output. Try adding --unicode explicitly. For byte-level inspection, run with --out ccd to get the raw hex byte pairs, which makes it easier to identify the encoding problem before fixing it.

Can I convert the extracted file to SMPTE-TT directly?

Not from CCExtractor directly: its output formats are SRT, SCC, ASS, WebVTT, SAMI and binary. Extract to SRT first, then convert with SeConv: seconv out.srt SMPTE-TT2052 --output-folder . --overwrite --fps 29.97. Deliver SMPTE-TT (ST 2052-1) XML for broadcast and streaming targets that require a 708-level file rather than a hand-built MCC.

Official Sources

Written by Dao Huy (Lucas), Vietnamese translator & localization specialist (EN · ZH · FR → Vietnamese). See translation services →

Get QuoteWhatsApp