SDH sound effects, music and speaker IDs: the notation I use
← 博客
🎬 字幕与隐藏字幕7 min read

SDH sound effects, music and speaker IDs: the notation I use

💡 In SDH, sound effects go in [SQUARE BRACKETS] on their own line, off-screen speakers are flagged in italics (SPEAKER: dialogue), and music wraps in ♪ symbols. These three patterns cover the vast majority of non-dialogue events in any SDH project, per the W3C WAI transcribing guidelines and WCAG 1.2.2.

Key takeaways

  • Sound effects: [DESCRIPTION IN BRACKETS], uppercase, on their own line before or during the sound
  • Off-screen speakers: italicize the full cue and add SPEAKER NAME: before the dialogue
  • Music: ♪ mood or title ♪ for instrumentals in italics; ♪ actual lyrics ♪ when the song is sung on screen
  • Delivery context: (whispers), (laughing), (sighs) placed inline before the dialogue line
  • Shouting or yelling: ALL CAPS dialogue text, no bracket tag needed

What actually makes an SDH track different from a plain subtitle file?

SDH (Subtitles for the Deaf and Hard of Hearing) carries the full audio picture: all dialogue, meaningful sound effects, music cues, and speaker identifications when the camera does not show who is speaking. A plain subtitle file carries only the dialogue and maybe a burned-in song lyric.

WCAG 1.2.2 states the requirement directly: captions must convey "not only dialogue, but identify who is speaking and include non-speech information conveyed through sound, including meaningful sound effects." SDH is not a stricter optional layer - it is the accessibility baseline for pre-recorded video published on the web.

In practice, the extra events show up in roughly one cue in four. That is where the craft of SDH lives: a door slamming off-camera that signals someone has entered, a whispered exchange that shifts the power dynamic in a scene, a music cue that sets tone before the first word is spoken.

The SDH notation conventions table

I use this as my reference sheet at the start of every SDH project. The square-bracket convention follows the format YouTube documents ("add text like [applause] or [thunder]") and is dominant in US streaming deliverables. The W3C WAI guidelines use parentheses instead, but square brackets survive plaintext conversion without looking like editorial notes, which is why I prefer them in SRT deliverables.

Audio eventNotationExample in SRT text
On-screen sound effect[ALL CAPS DESCRIPTION][GLASS SHATTERS]
Off-screen sound effect[DESCRIPTION][THUNDER IN DISTANCE]
Off-screen named speakerSPEAKER: dialogue (full cue italic)ANNA: Get away from there.
Off-screen unidentified voiceMAN: / WOMAN: (full cue italic)MAN: Who is out there?
Non-diegetic instrumental music♪ mood or title ♪ in italics♪ tense orchestral score ♪
On-screen song with lyrics♪ lyrics ♪♪ I can not stop the rain ♪
Whispered speech(whispers) before dialogue(whispers) We need to leave now.
Shouting or yellingALL CAPS dialogue textI TOLD YOU TO LEAVE.
Emotional delivery context(descriptor) before dialogue(laughing) I can not believe it.

A note on italics in SRT: mark off-screen cues with <i>...</i> tags around the full event. Some NLE timelines strip these on ingest; for broadcast CEA-708 delivery the styling lives in the 708 service layer, not in SRT tags, but the notation convention stays the same.

How do I format speaker IDs for on-screen and off-screen voices?

The rule I follow: if the camera shows the speaker's face at the moment of speech, no ID tag is needed. Viewers can see who is talking. The tag is for when they cannot.

For named characters heard off-screen (voice-over, phone calls, through a door), I italicize the full cue and put the character's name in ALL CAPS before a colon: DETECTIVE: You should not have come back here. I use the full name the first time a character speaks off-screen in a scene; after that the name alone is enough.

For unidentified voices (a distinguishable line in crowd noise, a caller not yet revealed by name), I use a descriptive label: DISPATCHER: or CHILD:. I avoid vague labels like VOICE: unless nothing more specific is identifiable. For documentary narration I establish the narrator once with NARRATOR: on the first cue; subsequent narration in the same run drops the tag because context carries the ID.

Here is how these rules look in a real four-event SRT sequence from a thriller scene:

1
00:00:04,200 --> 00:00:06,100
[DOOR SLAMS]

2
00:00:06,200 --> 00:00:09,400
<i>ANNA: (breathing heavily)
Are you in there?</i>

3
00:00:09,500 --> 00:00:11,800
(whispers) Get away from the window.

4
00:00:11,900 --> 00:00:14,200
<i>♪ low, dissonant strings ♪</i>
  • Cue 1: [DOOR SLAMS] is a standalone event. It tells the viewer something happened before the camera shows what.
  • Cue 2: Anna is off-screen (heard, not yet seen). The <i> tag wraps the whole cue; ANNA: names the speaker; (breathing heavily) is an inline delivery note.
  • Cue 3: on-screen dialogue, no speaker tag needed. (whispers) before the text flags the delivery mode.
  • Cue 4: non-diegetic music. The ♪ wrapper and italics mark it as music; I describe the mood rather than just "music playing" because the mood is what matters to the viewer.

When should I include a sound effect vs leave it out?

The W3C WAI guideline offers a practical test: caption a sound "if it is necessary for the understanding and/or enjoyment of the media." I apply it by asking: would a Deaf viewer miss something story-relevant if this event were silent? Yes means it goes in. No means skip it.

Events I always include: a gunshot, a crash or collision, a phone ringing, a door opening to reveal someone, an alarm that triggers a character's reaction, a notification the plot hinges on. Events I typically skip: ambient traffic in the background of a city scene when nothing story-relevant is happening, crowd murmur that is pure atmosphere, a generic music bed that only fills silence.

The trap I see most often in amateur SDH: over-tagging ambient background ([AMBIENT CITY SOUNDS] before every urban scene) and under-tagging reactions to off-screen events. A character who flinches at something the viewer cannot hear - that is the event that needs the tag, not the city ambience three seconds before.

What I handle myself vs when broadcast SDH needs a specialist

Anyone with Subtitle Edit and the notation table above can build a solid SDH track for streaming or web video. The "Fix Common Errors" pass in Subtitle Edit catches mismatched italic tags. The SDH filter under the Tools menu shows events carrying sound-effect tags so you can audit coverage against the script before delivery.

The threshold where specialist handling earns its cost: US broadcast delivery. Broadcast SDH in the US is delivered as CEA-608 encoded in SCC or SMPTE-TT format, where the sound-effect and speaker-ID text feeds into a structured data stream. The notation I described still drives the text content, but the delivery format and QC requirements are different from a plain SRT file. Getting the notation right is a prerequisite; producing a correctly structured broadcast deliverable is the step where a specialist in CEA-608/708 workflow saves significant re-delivery risk.

For English SDH and closed-caption deliverables on US streaming titles, my English SDH captioning service covers both the notation and the delivery format. If you are pairing that English SDH track with a Vietnamese subtitle, the subtitle translation workflow keeps timing and speaker IDs synchronized.

FAQ

Should sound-effect tags go on their own SRT line or inline with dialogue?

Own line, as a separate event, when the sound does not overlap with speech. Inline (a bracketed descriptor before the dialogue) when the sound and speech are simultaneous and splitting the event would break timing. Most story-relevant sound effects occur in audio gaps between dialogue lines, so a standalone event is the natural fit in the majority of cases.

Do italics mark all off-screen audio, or just off-screen speakers?

Italics mark off-screen audio sources in general: voices over a phone, speakers in the next room, voice-over narration, and non-diegetic music. On-screen music - a band visible in frame, a singer performing - is not italicized. The W3C WAI transcribing guidance specifically recommends italics for off-screen speech.

How do I caption background music I cannot identify by title?

Describe the mood and instrumentation: ♪ gentle acoustic guitar, slow tempo ♪. When the piece is identifiable, include the title: ♪ Beethoven's Moonlight Sonata ♪. For documentary contexts where the piece is the content, the WCAG 1.2.2 Understanding document gives a worked example: [Orchestral Suite No. 3.2 in D major, BWV 1068, Air] ♪ Calm melody with a slow tempo ♪.

My distributor asks for both SDH and closed captions. Are those the same file?

Often no. A typical US streaming delivery includes an English closed-caption file (SCC or SMPTE-TT format, CEA-608/708 encoded) and a separate SDH subtitle file (SRT or TTML). The closed-caption file carries the 608/708 data required under US law. The SDH file is the text-subtitle version available alongside a dubbed audio track. Confirm with your distributor which deliverables they need and in what format before you produce them.

Does every SDH event need a speaker ID?

No. A speaker ID is only needed when the viewer cannot tell who is speaking from the visual context. If the character's face is on screen at the moment of speech, no tag is needed. Tags are for off-screen voices, overlapping dialogue where more than one person speaks at once, and scenes where the speaking character is not shown for multiple seconds.

Official Sources

Written by Dao Huy (Lucas), Vietnamese translator & localization specialist (EN · ZH · FR → Vietnamese). See translation services →

报价WhatsApp