← All articles

5 min read

What a Caption Needs Beyond the Words

Automatic captions transcribe speech and nothing else. The speaker, the sounds that matter and the tone are a manual pass, with settled conventions worth following.

Automatic captions transcribe speech. A caption that serves a deaf or hard-of-hearing viewer has to carry more than speech, and nothing in a transcription tool will add the rest for you. That gap is the difference between a file that is technically captioned and one that is actually usable.

What speech-only captions leave out

Watch a clip with the sound off and read only the words. Three things go missing, and each of them changes what the scene means.

Who is talking. Fine when the speaker is on screen and their mouth is moving. Not fine in an interview cut to a listening shot, in a voiceover, or when someone speaks from off frame. The viewer reads a line and cannot attribute it.

Sounds that carry meaning. A knock at the door, a phone ringing, a car pulling away. If a hearing viewer would react to it, its absence changes the scene rather than merely thinning it.

Tone the words do not carry.Sarcasm, a whisper, shouting, a voice breaking. Written flat, a sarcastic line reads as sincere and means the opposite of what was said.

The conventions, and they are worth following

These are long-settled in subtitling, and following them matters because viewers already know how to read them. Improvising your own notation makes the reader work out your system before they can follow the video.

Non-speech information goes in square brackets, lower case inside: [door slams], [phone ringing], [upbeat music]. Speaker identification goes in front of the line, as a name and colon or a dash when two people share a cue, which is covered in more depth in the guide to captioning two speakers. Tone goes in brackets too, and only when it is not obvious: [sarcastic], [whispering].

Music is the case people overthink. Name it when it sets the mood, [tense music], and give the lyrics when the lyrics matter to the video. A three minute clip does not need every bar annotated.

The mistake is annotating everything

Having learned the brackets, the temptation is to describe every sound in the room. Do not. A caption track that reports the air conditioning and every footstep buries the dialogue in noise, and reading is slower than hearing to begin with.

The test is whether a hearing viewer would notice. Ambient hum: no. A door closing that makes someone look up: yes. If the sound has no consequence on screen, leaving it out loses nothing.

The same restraint applies to speaker labels. Name the speaker when the speaker changes and cannot be identified from the picture, not on every cue. Repeating a name through a long answer spends characters the words need, and on vertical video the line limit is already tight.

Where this sits in the workflow

Generate the transcript first, since the speech is the bulk of the work and a machine does it faster than you will. Then make one pass for the things a transcript cannot know: who is speaking, what else is audible, and where tone changes the meaning.

Do that pass before checking line lengths, because every bracket and every speaker label adds characters. A name and a colon can be ten characters before a word of dialogue. Check lengths afterwards with the subtitle validator and re-break with the line breaker, which leaves timing alone.

Keep the result as a subtitle file rather than burning it into the video. Burnt-in captions cannot be turned off by a viewer who does not need them, cannot be translated, and cannot be read by assistive software, which is the trade laid out in hardcoded versus soft subtitles. To build the track in the first place, start in the caption generator and export SRT or VTT when the wording is settled.

The short version

A transcript is not a caption track. Add the speaker when the picture does not say it, the sounds a hearing viewer would react to, and the tone that the words alone reverse. Then stop, because everything past that point is noise that slows the reader down.

Quick answers

What is the difference between subtitles and SDH?

Subtitles assume you can hear and carry the dialogue only, usually for someone who does not follow the language. SDH stands for subtitles for the deaf and hard of hearing, and adds what a viewer would otherwise get from the audio: who is speaking, sounds that matter to the scene, and tone that the words alone do not convey.

How should non-speech sounds be written?

Bracketed, in lower case, at the point the sound occurs. The notation is decades old and that is precisely the argument for it: your reader has met it before and does not have to learn anything. A private system, however tidy, costs them a few seconds of decoding on every cue.

Should I caption every sound?

Almost never, and over-annotating is what people do immediately after learning the brackets. Ask whether the sound has a consequence on screen. If nobody reacts to it and nothing changes, recording it adds reading time and returns nothing, and reading is the slower channel to begin with.

When do I need to name the speaker?

At the moment attribution becomes ambiguous, which is narrower than it sounds: a voice from off frame, a voiceover, or a cut away to whoever is listening. On camera and mid-answer, the picture is already doing the work, and repeating the name only eats into a line budget that vertical video makes tight.

Does this affect my line lengths?

Substantially, and in a direction people forget. Every bracket occupies a line and every attribution costs roughly ten characters before the first word of dialogue. Sequence the work so measurement comes last, or you will be measuring a file that no longer exists.

Should accessible captions be burnt into the video?

A separate file is the better default, because it leaves the viewer in control and stays machine-readable. Pixels cannot be switched off, translated, or handed to assistive software. Feeds that autoplay silently are the exception worth hedging on, and there the answer is to ship both rather than to abandon the file.

Caption EditingAccessibilityCaption Readability

Related caption tools

More caption guides

Caption your next video with AutoCaption.

Upload a video in the browser, generate editable captions, then export subtitle files or a styled captioned video.

Open caption generator