How to Extract Embedded CEA-608/708 Captions from MP4, TS and MXF
💡
ccextractor input.mp4 -o out.srtextracts embedded CEA-608 captions from an MP4, MPEG-TS or MXF file as an SRT. Run ffprobe first to see what streams a file carries. For MXF, pass--input mxf. To split Field 1 and Field 2 or target a CEA-708 service, use--output-fieldand--service.
Key takeaways
- CCExtractor 0.96.5 auto-detects MP4, MPEG-TS and MXF; force input with
--input mxfwhen auto-detection fails on older file wrappers. - Probe streams first:
ffprobe -v error -show_entries stream=index,codec_type,codec_name,codec_tag_string -of default input.mp4reveals separate data tracks (codecc608oreia_608); TS and MXF embed captions in the video SEI so no separate track appears. - CEA-608 Field 1 carries primary English captions (default). Use
--output-field 2for Field 2, which often holds a second-language caption track. - CEA-708 DTVCC services:
--service 1for the primary language service,--service 1,2for both primary and secondary. - Default output is SRT. Use
--out sccto keep the native CEA-608 encoding when you need a broadcast-ready file for re-mastering or QC.
Checking for embedded captions before you extract
Before I run CCExtractor on any file, I check what the container actually holds. The method differs by container type. For MP4 and MOV, ffprobe lists all streams including any separate data tracks. Run this on any MP4:
ffprobe -v error -show_entries stream=index,codec_type,codec_name,codec_tag_string -of default input.mp4
What each flag does:
-v error: suppress informational output so only stream data prints.-show_entries stream=...: list only these fields per stream.codec_tag_string: shows the FourCC. A data stream taggedc608oreia_608is a separate CEA-608 track embedded at the MP4 container level.-of default: plain key=value output, easy to read without extra tools.
For MPEG-TS and MXF, CEA-608 and CEA-708 data rides inside the video stream as H.264 SEI NAL units. ffprobe will not show a separate caption stream, and that is expected. If ffprobe shows only video and audio streams in a TS or MXF, that does not mean captions are absent: run CCExtractor and check its log output for detected caption data.
The basic CCExtractor command for MP4, TS and MXF
The simplest extraction converts CEA-608 captions to SRT. CCExtractor auto-detects the input format for MP4 and TS. For MXF I always force the input type because older MXF wrappers can trip the auto-detection:
# MP4 or MOV: auto-detect, output SRT
ccextractor input.mp4 -o out.srt
# MPEG-TS: auto-detect works reliably
ccextractor input.ts -o out.srt
# MXF: force input format to avoid auto-detection failures
ccextractor --input mxf input.mxf -o out.srt
# Output SCC instead of SRT (preserves native CEA-608 encoding)
ccextractor --out scc --input mxf input.mxf -o out.scc
Flag notes:
-o out.srt: output file path. Without-o, CCExtractor writes a file named after the input with an automatic suffix.--input mxf: forces the MXF reader. Accepted values:ts,mp4,m2ts,mkv,mxf,scc.--out scc: output format. Other options:ass,webvtt,webvtt-full,sami,ccd,bin.
CCExtractor logs a summary at the end of every run. Watch for "Extracted X frames" and lines mentioning EIA-608 or CEA-708 data. If both caption counts are zero, the file carries no detectable caption data in the streams CCExtractor reads.
How do I target a specific field, channel or CEA-708 service?
CEA-608 carries caption data in two fields and two channels per field. CEA-708 carries numbered DTVCC services. By default, CCExtractor extracts both 608 and 708 data when present, producing two output files. The flags below let you be precise:
| Target | Flag | When to use |
|---|---|---|
| 608 Field 1 only | --output-field 1 | Primary English captions (the default when no field flag is given) |
| 608 Field 2 only | --output-field 2 | Secondary captions, often a second-language track |
| Both 608 fields | --output-field both | Writes two files with _1 and _2 suffixes |
| Channel 2 within a field | --cc2 | Channel 2 of the active field |
| 708 Service 1 | --service 1 | Primary language DTVCC service |
| 708 Services 1 and 2 | --service 1,2 | Primary and secondary DTVCC services |
| All 708 services | --service all | Extract every DTVCC service present |
Examples:
# Extract 608 Field 2 as SRT
ccextractor --output-field 2 input.ts -o field2.srt
# Extract 608 Field 1, channel 2
ccextractor --output-field 1 --cc2 input.mp4 -o channel2.srt
# Extract CEA-708 Service 1 only
ccextractor --service 1 input.mp4 -o service1.srt
# Extract both 608 fields and both 708 services
ccextractor --output-field both --service 1,2 input.mxf -o captions
Most US broadcast material uses 608 Field 1, Channel 1 for English captions. Field 2 sometimes carries a Spanish caption track. CEA-708 Service 1 is the primary language service in HDTV digital captions. When I am not sure what a file carries, I extract with --output-field both --service all first and check what output files are created.
What output formats can I get from CCExtractor?
The default is SRT, which covers most editing and upload needs. The --out flag switches formats. My usual choices and when I reach for each:
| Format flag | Output file | Use case |
|---|---|---|
--out srt (default) | SubRip (.srt) | Edit in Subtitle Edit, upload to YouTube, translation workflow |
--out scc | Scenarist Closed Caption (.scc) | Preserve 608 encoding; re-master or QC-decode with caption-inspector |
--out webvtt | WebVTT (.vtt) | Web video player, HTML <track> element |
--out webvtt-full | WebVTT with styling | Preserves color and position metadata where present |
--out ass | SubStation Alpha (.ass) | Styled subtitles for playback or burn-in with ffmpeg |
--out ccd | CCD disassembly (.ccd) | Low-level 608 byte inspection |
After extracting to SCC, I QC-decode with caption-inspector to verify the text and positioning before delivering:
docker run --rm -v "$(pwd)":/data nebulabroadcast/caption-inspector -o /data /data/out.scc
This writes a -C1.608 text decode file. Reading it confirms the extracted text matches what I expect, catching encoding problems that happen during extraction before they reach the client.
When should I hand caption extraction to a specialist?
CCExtractor is free and handles the common cases well. I use it as part of every QC workflow. The extraction command is only the starting point in these situations:
- Broadcast deliverable with QC sign-off required: extracting an SRT from a client MXF is quick, but if the extracted captions need correction, re-timing and delivery back as an SCC or SMPTE-TT file for broadcast ingest, that is a full caption QC job.
- CEA-708 DTVCC service content: CCExtractor's 708 decoder is documented as a work in progress. For a delivery that depends on the 708 service layer, I verify against caption-inspector and may need manual correction on window styling.
- MXF with ancillary VBI data: some broadcast MXF files embed caption data in ancillary data packets rather than video SEI. CCExtractor may not find these. A broadcast caption tool with dedicated MXF ancillary reading is needed.
- Extraction plus translation: extracting the timing is step one. Producing a correct translated subtitle for a localised delivery is a separate task with its own CPS checks and QC pass.
If you need English captions extracted, corrected and delivered as a broadcast or streaming file, or Vietnamese subtitles produced from the extracted timing, my caption QC and delivery service covers both steps end to end.
FAQ
CCExtractor says it found no captions. What do I check first?
First confirm the video codec is H.264 or MPEG-2, the most common carriers for 608/708 SEI data. Then try with --output-field both --service all to ensure you are not filtering out the active field or service. If CCExtractor still finds nothing, the file may have no embedded closed captions, or captions may be in MXF ancillary data outside the video stream, which CCExtractor does not decode.
Can CCExtractor extract subtitles from an MKV file?
CCExtractor reads MKV containers and looks for CEA-608/708 data in the video SEI. If the MKV has a separate SubRip or SubStation Alpha soft subtitle track instead of embedded 608/708 data, CCExtractor will not find those. For soft subtitle tracks, use ffmpeg directly: ffmpeg -i input.mkv -map 0:s:0 out.srt extracts the first subtitle track without any re-encode.
How do I know which field carries the English captions?
Run CCExtractor with --output-field both --service all first and check which output files contain readable English text. US broadcast material almost always uses CEA-608 Field 1, Channel 1 for primary English captions. CEA-708 Service 1 is the primary HDTV digital service. Field 2 sometimes carries a second-language track such as Spanish.
The extracted SRT has garbled characters. What went wrong?
CEA-608 uses a custom Latin character set. CCExtractor converts it to UTF-8 by default, but a non-standard or incorrectly flagged character set in the source can produce garbled output. Try adding --unicode explicitly. For byte-level inspection, run with --out ccd to get the raw hex byte pairs, which makes it easier to identify the encoding problem before fixing it.
Can I convert the extracted file to SMPTE-TT directly?
Not from CCExtractor directly: its output formats are SRT, SCC, ASS, WebVTT, SAMI and binary. Extract to SRT first, then convert with SeConv: seconv out.srt SMPTE-TT2052 --output-folder . --overwrite --fps 29.97. Deliver SMPTE-TT (ST 2052-1) XML for broadcast and streaming targets that require a 708-level file rather than a hand-built MCC.
Official Sources
- CCExtractor GitHub repository - README and
--helpoutput confirming flags:--input,--out,--output-field,--service,--cc2. Verified October 2026. - ffprobe documentation (FFmpeg.org) -
-show_entries,-select_streamsand stream output fields. Verified October 2026. - Subtitle Edit (Nikse.dk) - SeConv command-line converter, SMPTE-TT2052 format name,
--fpsflag. Verified October 2026.
Written by Dao Huy (Lucas), Vietnamese translator & localization specialist (EN · ZH · FR → Vietnamese). See translation services →
