Very short, punchy delivery
If the line is three words long, there is nothing to read ahead to and the wipe has nowhere to travel. A word-pop style carries short hooks better.
Caption style
The most-used caption style of the moment, and the one most often faked. A real karaoke caption tracks the speech word for word. A fake one divides the line evenly and drifts out of sync by the second syllable.
Karaoke — The gold wipe crosses each word exactly as it is said.
Closest alternatives — point at one to play it
Hormozi
Hormozi Pop
Podcast Duo
Word spans come from the transcript's own timings wherever the audio provides them.
This is the one structural difference between karaoke captions and the Hormozi family, and it decides which one you should use. Word-at-a-time styles withhold the sentence: you cannot read ahead, so you stay locked to the audio. Karaoke shows the sentence and marks your place in it, so you can read ahead — and for anything explanatory, reading ahead is how comprehension works.
That makes it the right choice for the content word-pop styles handle badly. Dense scripts. Lists of figures. Anything where the viewer needs to hold a whole clause in mind. It is also why it is the standard for lyric videos and educational content, where the text is the point rather than the punctuation.
The translucent plate is there for a specific failure. A wipe needs both states legible at once — the said words and the unsaid ones — over whatever is behind them. A stroke alone handles one word in isolation; across a full line over moving footage, the unsaid half of the line is what breaks up first. A 32% plate holds the whole line steady without blacking out the frame.
If the line is three words long, there is nothing to read ahead to and the wipe has nowhere to travel. A word-pop style carries short hooks better.
A full line plus a plate is a wide, solid block. Over footage where the speaker sits low in frame it covers them; a behind-the-subject style solves exactly that.
The wipe is only honest if the word spans are. Heavy music, overlapping speakers or a very noisy recording degrade word timings, and the wipe falls back to an even division across the cue.
Like any highlight, the gold wipe marks position by hue. It is a style, not an accessibility feature — export a real subtitle file for that.
Cleaner audio gives tighter word timings, so use the best version you have.
In the Creators family. The wipe runs on your own audio in the preview.
Watch one line land before exporting video or a subtitle file.
Really synced, wherever the audio allows it. Transcription is asked for word-level timestamps and the renderer uses each word's own start and end, so a word held for 0.78 seconds is lit for 0.78 seconds. The timings are sanitised first — a word cannot start before its cue or before the previous word ends. If a clip returns no usable word spans, the line falls back to an even split, which keeps the motion honest-looking but not exact.
People use the terms for two different things. Karaoke keeps the whole line on screen and moves a highlight through it. Word-by-word — the Hormozi family — shows one or two words at a time and nothing else. Karaoke lets the viewer read ahead; word-by-word deliberately does not.
It is the style lyric videos use, and it works as long as the vocal is separable enough for the transcript to time. Over a dense mix the word spans get loose, and a lyric video is the one format where viewers notice immediately.
Yes — the highlight, the resting colour, the plate opacity and the outline are all editable, and the preview updates on your own footage.
Karaoke, usually. Podcast clips are explanatory and the speech is conversational rather than punchy, which is exactly the case where letting the viewer read ahead helps. The Speakers family has presets built specifically for two-person audio.
Upload a clip and watch one line land before you commit to the style.