9:16 Video Cleanup

Remove Text from a Vertical Video

Published July 29, 2026 ยท Reviewed against Wipe AI product facts

Direct answer: Use a tight, time-based selection to remove text from an authorized 9:16 clip without changing its aspect ratio. Preview center captions carefully because they often cover faces and products.

Vertical videos pack captions, usernames, stickers, titles, and calls to action into a narrow frame. Cropping one edge can cut off the subject or break the 9:16 composition, so region-based reconstruction is often the more practical option.

The tradeoff is that center-screen captions may hide important detail. A clean export or editable source is still better than generating replacement pixels.

Open Text RemoverBrowse All Guides
Authorized use only: Process video you own, license, or have permission to edit. Removing a visible element does not change copyright, attribution, disclosure, privacy, or platform obligations.

Protect the 9:16 composition

Keep the full frame when the subject, product, hands, or interface already uses the available width. Select only the visible letters and their shadow or plate instead of cropping the entire caption band.

If the text jumps between top, center, and bottom positions, create separate time ranges. A single tall mask would erase clean visual information throughout the clip.

Test dynamic social captions

Word-by-word captions change size and position quickly. Include the largest footprint within a stable layout, but do not cover every possible caption position if the design changes between sections.

Preview hair, faces, fingers, product edges, and patterned clothing. Small reconstruction errors are more noticeable when the text sits directly over the subject.

Prepare the cleaned clip for reuse

Wipe AI exports H.264 MP4 at up to 1080p. Confirm the output dimensions and safe zones before adding new captions, platform UI, or branded text.

Keep new captions editable or in a separate subtitle track when the publishing platform allows it. This makes future localization and corrections easier.

Recommended workflow

  1. Keep the 9:16 master and identify every distinct text layout.
  2. Select the smallest complete text, shadow, and background-plate footprint.
  3. Split top, center, and bottom layouts into separate time ranges.
  4. Preview frames where text overlaps faces, products, or fine detail.
  5. Export and add replacement captions within the destination safe zone.
Source-first alternatives: Return to the social editing project, export without captions, use a clean camera roll file, reframe only when the subject remains safe, or cover old text with an accurate new design.

Frequently asked questions

Can I keep the 9:16 aspect ratio?

Yes. Region-based cleanup does not require cropping the vertical frame, though the output may be re-encoded.

Can Wipe AI remove animated captions?

Visible caption regions can be processed, but changing layouts should be split by time and tested where they overlap the subject.

Will face details be restored exactly?

No. If text hides a face, inpainting estimates the missing area. Use a clean source whenever accuracy matters.

What output format is used?

Wipe AI exports an H.264 MP4 at up to 1080p.

Related video cleanup guides