← all posts

How to Add Captions Behind a Person in Your Video

Add the text behind person effect to your video captions automatically. No masking, no layers, no keyframes. Drop a video, pick a template, export.

The text behind person effect puts your text behind the subject in the frame, so captions look like part of the scene instead of a sticker on top of it. Every tutorial you find teaches it with static text that you type, position, and time by hand. None of them do it with real captions that follow your speech.

Tscaps does. You drop a video, get word-level captions, pick a template with the effect, and export. The AI finds where a person is in the frame and composites the cutout above your captions, frame by frame.

What the text behind person effect is

The idea is three layers stacked on top of each other: the original background at the bottom, your text in the middle, and a cutout of the person on top. Because the person’s pixels sit above the text, the text looks like it is behind them.

On short-form video this creates a depth that flat captions cannot match. It works well for intros, hooks, and any moment where a person talks straight into the camera and the caption needs to grab attention without covering their face.

The effect became popular on TikTok and Reels around 2025. Tools like CapCut, Rotato, and FlexClip added support for it, but all of them work with a static text box you place by hand.

You type the words, you drag them into position, and you resize until they look right. If the video has ten lines of dialogue, you are doing that ten times.

The problem with doing it by hand

In After Effects or Premiere Pro, the text behind person effect means rotoscoping or keying a mask for every frame where the person appears. On a 60-second clip at 30 fps, that is 1,800 frames. Professional editors automate parts of it, but the setup still takes time and the tools cost money.

Lighter tools like CapCut and Rotato skip the manual masking. They run an AI model to separate the person from the background, and you get the three-layer composite without drawing a single mask. That part is solved.

What none of them solve is the text. You still type the words, choose the font, position the box, and decide when it appears and disappears. For a single title card that is fine. For captions on a talking-head video with 40 lines of dialogue, it is the same work repeated for every line.

And if you change the wording or fix a typo, you reposition again.

How tscaps adds text behind person to your captions

Tscaps is an online caption editor. Nothing to download. You drop a video, it transcribes the audio into word-level captions, and you pick a visual style from the template gallery.

One of those templates activates the text behind person effect. When you pick it, the editor runs a person segmentation model on the video, finds the scenes where a person is visible, steady, and sharp, and marks those as valid windows for the effect.

Every caption segment that falls inside a valid window gets the effect automatically. The captions shift up behind the person, and the cutout paints on top. You see it in the preview as you edit, and it burns into the exported video the same way.

Three things that matter here:

  • You do not position anything. The template controls where the captions sit and how they move when the effect activates. The timings come from the transcription. You edit the words if you need to, the rest is handled.
  • Per-segment control. If the effect activated on a segment and you want it off, or it did not activate and you want it on, there is a toggle in the segment settings. Force-on runs the segmentation model on the frames that segment covers, even if they were not in a valid window.
  • It works on every template. The auto-activation is tied to templates that opt into it, but the per-segment toggle is available on all of them. You can force the effect on for a single line, on any style.

What about export?

The effect composites into the final MP4 the same way it shows in the preview. The segmentation masks are captured once and reused at export, so the burn does not re-run the AI model. What you see in the editor is what you get in the file.

Tscaps local is free, needs no account, and exports without a watermark. The cloud lane adds AI transcription and AI styling on top; its free tier carries a watermark on exports, paid plans remove it. The text behind person effect is available on both surfaces, across every tier.

How it works under the hood

The editor sends each video frame to a MediaPipe person segmentation model running inside a web worker. The model returns a mask: which pixels belong to a person and which belong to the background.

The editor scans the video once and identifies scenes where a person is present, steady, and in focus. Those scene windows are cached together with downsampled masks for every frame inside them. At render time, three layers are painted in order: the background frame, the caption text, and the actor cutout on top using the cached mask.

All of this runs in the browser. The video does not upload to a server for segmentation. On devices with WebGL2 support the model runs on the GPU; otherwise it falls back to CPU.

When to use it and when to skip it

The effect works best when:

  • A person is centered in the frame and visible from the waist or shoulders up.
  • The camera is steady or moves slowly. Fast pans blur the mask edges.
  • The background has enough contrast with the person to separate them cleanly.

It is less useful when:

  • The video has fast cuts every second. The model needs a stable scene to build a valid window.
  • Multiple people overlap in the frame. The mask covers all of them, so the effect applies to the group, not to one person.
  • Nobody is on screen. No person, no cutout, no effect. The captions show normally.

The editor handles all of these cases: if a segment does not qualify, the effect does not activate, and the captions render in their default position. Nothing breaks.

Try it

Drop a video at tscaps.io and pick a template with the text behind person effect. Free, online, no account required.