CEA-608 vs CEA-708: technical differences and which to deliver
💡 CEA-608 uses a fixed 32-column × 15-row grid with a Latin character set, three caption modes, and a 2-bytes-per-frame transmission rate: its file format is SCC. CEA-708 adds window-based rendering, up to 63 caption services, Unicode characters, and flexible fonts and colours: its recommended open-source deliverable is SMPTE-TT (ST 2052-1), not MCC.
Key takeaways
- CEA-608: 32 columns × 15 rows grid, pop-on / roll-up / paint-on modes, Latin character set, SCC file format, 29.97 drop-frame timecode
- CEA-708: window-based display, up to 63 services, 8 windows per service, Unicode characters (font-dependent), SMPTE-TT (ST 2052-1) for open-source delivery
- CEA-708 embeds 608 compatibility bytes so legacy decoders fall back gracefully: a clean SMPTE-TT file carries both layers
- CEA-608 is a Latin-only standard. Vietnamese diacritics, Chinese characters, and other non-Latin scripts belong on a separate subtitle track, not inside the broadcast caption stream
- To detect which standard a file contains, run
ccextractor input.mp4 --out report: a 608-only file produces one output stream; a 708 file produces two
What does CEA-608 specify technically?
CEA-608 (published by CTA as CTA-608-E, still widely referenced as CEA-608 or EIA-608) is the Line 21 closed caption system used in North American broadcast from the 1970s through the ATSC digital transition. When I open an SCC file in a text editor, each line is a timecode followed by space-separated 4-character hex words: each word is 2 bytes of caption data, transmitted at one word per video frame at 29.97 fps drop-frame.
The display grid is fixed: 32 columns across 15 rows. The encoder places text in one of three modes:
- Pop-on: text is pre-loaded into a non-visible buffer, then appears all at once on a cue (most common for prerecorded content)
- Roll-up: 2, 3, or 4 lines scroll upward as new lines arrive (common for live broadcast)
- Paint-on: characters appear individually left to right as they are transmitted (real-time, character by character)
| Feature | CEA-608 value |
|---|---|
| Grid size | 32 columns × 15 rows |
| Caption channels | CC1, CC2 (Field 1); CC3, CC4 (Field 2) |
| Text channels | T1–T4 (Field 1); T5–T8 (Field 2) |
| Caption modes | Pop-on, roll-up (2, 3, or 4 rows), paint-on |
| Character set | Latin (English, French, Spanish, Portuguese, German, plus Northern European extended pack) |
| Non-Latin scripts | Not supported (see note below) |
| Timecode | 29.97 drop-frame (broadcast standard) |
| Transmission rate | 2 bytes per video frame |
| File format | SCC (Scenarist Closed Captions) |
| Signal path (analog) | VBI Line 21 |
| Signal path (digital) | VANC / H.264 SEI NAL unit |
The character set constraint is the most important fact I state upfront in every broadcast proposal. CEA-608 covers English, Spanish, French, and other Western European languages through basic and extended character packs. Vietnamese diacritics and Chinese characters are outside this set entirely. English is the correct deliverable for a CEA-608 stream; other languages go on a separate subtitle track in a Unicode-capable format like SRT or TTML.
What does CEA-708 add, and how does it carry 608 compatibility bytes?
CEA-708 (now CTA-708-E) is the digital television caption standard introduced for ATSC DTV and HDTV. Instead of a fixed grid, 708 uses independently positioned windows: each caption service can define up to 8 windows, each with its own position, size, font style, font size, foreground and background colour, edge type, and scroll direction.
The additions over 608 that matter for delivery:
- Up to 63 caption services (Service 1 = primary language, Service 2 = secondary language service)
- 8 font styles: Default, Mono Serif, Proportional Serif, Mono Sans-Serif, Proportional Sans-Serif, Casual, Cursive, Small Caps
- 4 font sizes: Small, Standard, Large, Extra Large
- 8 foreground and background colours (black, white, red, green, blue, yellow, magenta, cyan) with 4 opacity levels each
- 5 edge types: None, Raised, Depressed, Uniform (outline), Drop Shadow
- Unicode character support, though practical rendering depends on the display device's font set
The compatibility design is central to delivery quality. CEA-708 embeds CEA-608 data inside its packet stream: the 608 CC1-CC4 channels ride as compatibility bytes within the 708 data. A viewer with a legacy 608 decoder sees the compatibility text from CC1 or CC2; a 708-capable display renders the full 708 window layout. A broadcast deliverable that properly implements both layers must pass QC on both decoders. I use caption-inspector (the Comcast open-source tool, run via Docker) to verify that a SMPTE-TT file decodes cleanly at both the 608 and 708 layers before submitting a file.
How do I detect which standard my file contains?
The most direct method is CCExtractor's report mode. This reads the file, prints what caption data it finds, and writes nothing to disk:
ccextractor input.mp4 --out report
If the file contains only CEA-608 data, CCExtractor reports a single output stream. If it also contains CEA-708 data, it reports two streams. Once I know which standard is present, I use these extraction commands:
# Extract CEA-608 Field 1 (CC1) to SRT
ccextractor input.mp4 -1 -o output_608.srt
# Extract CEA-708 Service 1 to SMPTE-TT (TTML)
ccextractor input.mp4 --service 1 --out smptett -o output_708.ttml
# Extract both CEA-608 and CEA-708 simultaneously
ccextractor input.mp4 -1 --service 1
-1: process Field 1 (the default); use-2for Field 2 or-12for both fields--service 1,2: extract 708 Services 1 and 2 (primary and secondary language services)--out smptett: output as SMPTE Timed Text (TTML), the recommended format for 708 deliverables in an open-source workflow--out report: print caption content and statistics to stdout without writing any file
To check whether an MP4 has an embedded closed caption stream at the ffprobe level, I probe the first video stream for the closed_captions field by analyzing a short read interval:
ffprobe -v error -select_streams v:0 -show_entries stream=closed_captions -analyze_frames -read_intervals "%+#30" -of default=noprint_wrappers=1 input.mp4
-select_streams v:0: inspect only the first video stream-analyze_frames -read_intervals "%+#30": decode 30 frames to surface SEI-embedded caption data that only appears at the frame level-show_entries stream=closed_captions: returnsclosed_captions=1if caption data is present in the video track
Which delivery format goes with which standard?
For US broadcast and regulatory submissions I use SCC for a CEA-608 deliverable and SMPTE-TT (ST 2052-1 XML) for a CEA-708 deliverable. MCC is the historical option for 708, but every open-source MCC export I have tested fails the 708 service layer in caption-inspector: the tool decodes only the 608 compatibility bytes, not the 708 service content. SMPTE-TT is the reliable open-source path for 708 deliverables.
| Delivery target | Standard | File format | Timecode |
|---|---|---|---|
| US analog broadcast (legacy) | CEA-608 | SCC | 29.97 drop-frame |
| US digital broadcast / cable (ATSC) | CEA-608 + CEA-708 | SMPTE-TT (ST 2052-1) | 29.97 drop-frame |
| US streaming (Netflix, web SDH) | SDH / closed captions | TTML/IMSC or SRT | Media-relative (HH:MM:SS,mmm) |
| DTV compliance submission | CEA-708 | SMPTE-TT (ST 2052-1) | 29.97 drop-frame |
The SeConv command to produce an SCC (608) file from an SRT source:
seconv input.srt ScenaristClosedCaptions --fps 29.97 --output-folder out --overwrite
And for a SMPTE-TT (708-compatible) file:
seconv input.srt SMPTE-TT2052 --output-folder out --overwrite
What can I handle myself, and where does specialist delivery change the outcome?
Reading an SCC file in a text editor and spotting the hex data takes seconds. Running CCExtractor in report mode, then exporting an SRT to SCC via SeConv, is a straightforward workflow using free tools. If you are producing captions for a web video, WCAG 1.2.2 is satisfied by an SRT or VTT file with the right content: sound effects, speaker IDs, and dialogue. No broadcast-encoded format is required for that use case.
The point where format knowledge changes the result: US broadcast compliance and regulated submission. 47 CFR Part 79 (the FCC closed captioning rule) requires the caption data to be encoded in the correct format for the distribution channel: SCC/608 for legacy broadcast, SMPTE-TT/708 for ATSC digital. A correctly timed SRT file does not satisfy that requirement. The column limit (32 characters per row in 608) and the 29.97 drop-frame timecode requirement both need validation before delivery. A visual check of the SRT cannot catch a hex encoding that silently truncates a line at the 32-character boundary.
For content going to a US streaming or broadcast platform, I run the finished SCC or SMPTE-TT through caption-inspector to confirm that both the 608 and 708 layers decode to the correct text and positioning. That QC step catches errors a text-level review cannot find.
For English closed-caption and SDH delivery on US streaming and broadcast titles, I cover both the content and the format through to QC. If you want an independent decode check on a file you have already produced, the caption QC service runs the full caption-inspector decode and reports exactly what a broadcast decoder sees.
FAQ
Is CEA-608 still required in 2026, or has 708 replaced it?
Both are still required for US digital broadcast. ATSC rules require that a digital broadcast carry the CEA-608 compatibility layer inside the 708 stream so legacy decoders continue to work. Producing a 708-only deliverable without the 608 compatibility bytes would fail QC for an ATSC broadcast submission. In practice, a correctly produced SMPTE-TT file carries both layers and satisfies both requirements.
My file is .mcc: does it contain 608 or 708 data?
An MCC file carries both CEA-708 and the 608 compatibility bytes. MCC is the binary format produced by commercial tools like Telestream MacCaption for ATSC delivery. If you received an MCC from a client and need to extract the 608 layer, CCExtractor handles it: ccextractor input.mcc -1 -o output.srt. For producing a new 708 deliverable in an open-source workflow, use SMPTE-TT: the open-source MCC export from Subtitle Edit produces a 708 service layer that caption-inspector cannot decode in my testing.
Can CEA-708 carry Vietnamese or Chinese caption text?
CEA-708 declares Unicode support, but practical rendering depends on the display device's font set. More importantly, for US broadcast delivery the entire pipeline is calibrated for English. The correct path for Vietnamese or Chinese language tracks is a separate SRT, TTML, or VTT file delivered alongside the English caption file. CEA-608 carries Latin characters only and does not support non-Latin scripts at all.
How do I verify that a SMPTE-TT file decodes correctly?
I use caption-inspector, the Comcast open-source QC tool. Run it via Docker: docker run --rm -v /path/to/dir:/data nebulabroadcast/caption-inspector -o /data /data/output.ttml. It produces a -C1.608 file showing the 608 compatibility text and a -S1.708 file showing the 708 service content. If the 708 file is empty or shows only errors, the 708 service layer is broken. If you only see 608 output, the SMPTE-TT file may be carrying only the compatibility bytes with no real 708 service content.
What if my master is 23.976 fps: can I still deliver CEA-608?
CEA-608 was designed for 29.97 fps broadcast. For a 23.976 fps master, confirm with the broadcaster whether they accept a 23.976-timed SCC or require a frame-rate conversion to 29.97 first. Most US broadcast delivery specifications are 29.97 drop-frame, and a 23.976 master is typically converted before caption encoding. For streaming delivery at 23.976 fps, TTML/IMSC with media-relative timing is cleaner and avoids the frame-rate mismatch entirely.
Official Sources
- CCExtractor GitHub repository - command line flags (-1, --service, --out smptett, --out report) verified against the installed 0.96.5 help output. Verified Oct 2026.
- W3C TTML2 Recommendation (W3C) - TTML namespace and document structure used in SMPTE-TT deliverables. Verified Oct 2026.
- WCAG 2.1 Understanding 1.2.2: Captions (Prerecorded) - caption content requirements (dialogue, sound effects, speaker IDs) for WCAG Level A compliance. Verified Oct 2026.
- ffprobe documentation (ffmpeg.org) - -analyze_frames, -read_intervals, and -show_entries stream=closed_captions options for detecting embedded caption data. Verified Oct 2026.
Written by Dao Huy (Lucas), Vietnamese translator & localization specialist (EN · ZH · FR → Vietnamese). See translation services →
