Subject-Aware Auto Reframing

Editors converting landscape footage to vertical video get an editable camera path that follows the real speaker or action, while uncertain sections are flagged for human review.

When a short-form video editor converts a landscape interview, game, or livestream into vertical video, the hardest part is not cropping—it is keeping the frame on the wrong person. The user drops the video into a timeline, selects the main speaker, player, or current action, and the product generates an editable virtual-camera path. The path stays steady within a safe zone so it does not crop off heads or captions; when the system predicts that another person is about to speak or receive the ball, it also begins a smooth move in advance.

Editors can drag key points on the timeline and adjust tracking strength and the safe zone. Every automated move remains editable camera-movement data rather than being baked into an irreversible crop. For conversations with multiple people, the editor can choose “prioritize the current speaker” or “keep everyone visible.” Action footage can follow a ball, car, hand, or another selected subject.

When the subject is occluded, people overlap, the subject exits the frame quickly, or confidence drops sharply, the product stops making the decision automatically, marks the section for review, and shows the candidate subjects it detected. The first version focuses on common people and sports objects and outputs keyframe tracks that remain editable in Premiere, CapCut, or Final Cut. It does not handle captions, color grading, or the final edit.

Why now

A September 11, 2026 post on r/EntrepreneursGrind complained that after trying three landscape-to-vertical tools, the result was still just a centered crop; existing options had not reliably followed the speaker or action. S1 As of September 12, the post had 2 points and 0 comments, and the problem surfaced at the delivery point where creators were converting landscape YouTube content into Reels. S1

Target user

The core users are editors who regularly reformat podcasts, interviews, courses, and sports footage for vertical video. They often receive Reels or Shorts deliverables only after the landscape master has been locked. At that point, captions, cuts, and pacing usually cannot be rebuilt, so they need to produce a reliable composition quickly. The longer the footage and the more often subjects change, the harder it is to finish shot-by-shot keyframing on deadline.

Minimal entry point

The technical core has three layers: subject tracking, shot planning, and export. After the user selects the subject in the first frame, SAM 2 propagates its mask forward. It supports point, box, and mask prompts and lets users correct the result in later frames. S3 Interview footage then goes through offline speaker diarization, matched against face tracks. Shot planning generates only position, scale, and Bézier keyframes, with constraints for headroom, the caption area, and movement speed. The first version supports one subject and two-person interviews, but not catch prediction for arbitrary ball sports. Export starts with Final Cut Pro’s FCPXML format, which can describe a project timeline and lets applications exchange project data with Final Cut Pro. S4

Punching above its weight

Find the first users among small editing teams that specialize in podcast clips, course distribution, and sports highlights. These teams handle repetitive footage and can directly compare the time required to reframe landscape video manually. Use the same multi-person interview to demonstrate centered cropping, automatic tracking, and the result after manual corrections. Then offer a small amount of free processing in exchange for failed sections and final keyframes, gradually covering the most common occlusion cases.

Competitors & gaps

Adobe Premiere Pro Auto ReframeGoogle
Premiere Pro can duplicate a sequence, convert it to a target aspect ratio, and apply Auto Reframe. S2 It offers three motion presets—Slow, Default, and Fast. Fast mode follows movement and generates more keyframes. S2 When footage contains multiple points of interest or fast motion, editors still need to adjust keyframes manually. S2 The current workflow mainly asks users to choose the frame shape and motion speed. Adobe’s documentation does not offer controls for selecting a lead subject, prioritizing the current speaker, or keeping everyone visible. It also does not isolate uncertain decisions or show the candidate subjects the system considered. Turn-taking interviews may still require a shot-by-shot review, while passes and fast exits in sports footage lack editor-facing ambiguity handling. The opportunity is to turn the result into a reviewable camera-movement track, so editors only handle genuinely ambiguous sections.

How it makes money

Charge a monthly per-seat subscription, with usage tiers based on the amount of video processed. The basic plan includes subject tracking and FCPXML export; higher tiers add team review, batch processing, and custom safe zones.

The case against

Speaker diarization and face matching can fail together, causing the camera to follow the wrong person continuously. If the system moves early toward the wrong next speaker, the finished video will contain an unexplained pan. Sports footage adds occlusion, cuts, and small fast-moving objects that are difficult for a single tracking model to cover. To reduce false positives, the system must retain confidence scores, candidate subjects, and edit history, increasing storage and interface complexity. Coordinates, scaling, and interpolation after FCPXML import also need validation across versions. If editors still have to watch all the footage, the time savings from automation fall sharply. Private interviews and unreleased sports footage may also restrict cloud processing, forcing the product to absorb the performance cost of local inference.

Evidence and sources

4 checkable sources cited
Trend observation· Reddit
Subject-aware video reframing
Sources
Telegram channel