Crowds and multiple subjects
The matte is built around a subject. With several people at similar depth the boundary becomes ambiguous and the caption pops in front of someone it should be behind.
Caption style
The one caption effect that cannot be faked with a font choice. The word has to be occluded by the person in front of it, frame by frame, as they move — which means the tool has to actually know where they are.
Butter Lock — A butter-yellow hero word springs in behind you, white subtitles in front.
Closest alternatives — point at one to play it
Neon Lock
Crimson Lock
Deep Focus
Segmentation runs on your footage in the preview, so you can see it hold before exporting.
Every other style on this site is a typographic choice: pick a face, a colour, a motion, and it renders the same way over any footage. This one is not a typographic choice at all. It is a compositing operation that depends on what is actually in the frame, and it either holds or it does not.
That is why it is worth doing properly. The effect is widely copied by hand — editors cut the subject out on a few key frames and hope the gap is short enough that nobody notices the edges swim. For a static shot that works. For anything where the speaker moves, gestures, or steps toward camera, hand-masking either takes a long time or looks wrong, and both of those are why most captioning tools do not offer this at all.
Because the caption sits behind the person, it also solves the problem every other style has to work around: covering the subject. A caption block placed over a speaker's chest is a compromise between legibility and hiding them. Behind them, there is nothing to trade off — the text can sit where it composes best.
The matte is built around a subject. With several people at similar depth the boundary becomes ambiguous and the caption pops in front of someone it should be behind.
Dark clothing against a dark room gives segmentation very little to work with, and the edge will flicker. Check the preview before committing.
A blurred limb has no crisp boundary, so the occlusion edge softens exactly where the eye is looking.
There is no subject to go behind. Use a normal style — the Faceless preset is built for that footage.
A single speaker, reasonably separated from the background, works best.
Butter Lock is the neutral one; the family has thirty more.
Check the boundary holds while the speaker moves, then export.
The frame is segmented into subject and background, and that subject matte is composited over the caption. So the caption is genuinely drawn first and the person drawn on top, per frame, rather than a static shape masking a fixed region. It follows the speaker as they move because the matte is recomputed rather than keyframed.
No. That is the point of it. The usual way to get this effect is to cut the subject out on key frames in an editor, which is slow and tends to swim at the edges between keys.
It depends on the shot, and the honest answer is that you should look rather than take a promise. A single speaker with reasonable separation from the background holds well. Crowds, very low contrast and heavy motion blur are the three cases that break it. The preview runs on your own clip, so you can see which you have before exporting anything.
That is the trade. If the speaker is centred and large, they will occlude a lot of the word, which looks good and reads worse. The presets place type where it usually survives a centred subject, but on a tight close-up you may want a normal style instead.
The style for footage with no subject to go behind.
Where a caption can sit without the interface covering it.
The same workflow, framed for a Reels edit.
A depth effect only exists burned in. What that costs you.
Upload a clip and watch the edge while the speaker moves.