How to make a short-form video with AI: script, storyboard, generate in segments, edit

Short-form made with AI fails when you generate thirty seconds at once and crop landscape to vertical. One 30-second explainer, ratio first, one shot per generation, with a checklist.

Tutorial
Published
Updated

Just open CapCut, drop your clips into a template, done. You can skip the rest of this article; that is enough to post. About a thousand other shorts are going out through the same template while you read this, so it clearly works. Simple. No need to read on.

Still here? Then let us make one that is not among the thousand.

The usual way to make a short-form video with AI goes like this: think of an idea, write a prompt, hit generate, wait for thirty seconds of video, crop it from landscape to vertical, add captions, post.

Two things break on that road. First, at thirty seconds in one generation, the model has forgotten what you asked for by second ten. Second, cropping landscape to vertical throws away the picture you paid to generate: the subject ends up at the edge and there is nowhere for the captions to go.

A vertical short is a different film from the aspect ratio onward. This is one thirty-second explainer, made from the start.

Can you just crop landscape to vertical? No, vertical is a different film

16:9 and 9:16 differ by far more than a rotation. A landscape frame reads left to right; the subject can stand at one third and leave a wide field of environment. A vertical frame reads top to bottom; the eye lands just above centre, there is almost no room to the sides, and the top and bottom belong to captions and the platform's buttons.

So a shot composed well in 16:9 (a person at the left third, a whole street to the right) becomes, in 9:16, half a body and a lamp post. What you lose is the composition, and no amount of resolution brings it back.

The reverse holds too: a shot generated vertical, subject centred with headroom, becomes a narrow figure between two black bars when forced into landscape.

So where you will publish decides which ratio you generate in, and the ratio is decided before the script is written. If you need both landscape and vertical, that is two storyboards and two rounds of generation.

How do you write a thirty-second script? One point, and the first three seconds keep them

A short-form script differs from any other script in two ways: it makes exactly one point, and the first three seconds have to stop the thumb.

Our example:

Why is convenience-store ice clear, and the ice from your freezer white?

One question, one answer, thirty seconds. Write it by the second:

Seconds Section Content
0–3 Hook Two ice cubes side by side, one clear, one cloudy. Voice-over: "Same water. Why the difference?"
3–10 Setup A home freezer freezes from all sides toward the middle; the air and impurities in the water get pushed to the centre and trapped there, a white core
10–22 Answer An ice machine freezes the water slowly from one direction only; the air has somewhere to go and is pushed out along the way, leaving clear ice
22–28 Turn You can do it at home: fill an insulated cup, leave it open at the top, freeze it; it freezes from the top down and the upper part comes out clear
28–32 Close The two cubes again. Voice-over: "The difference is where the freezing starts."

Thirty-two seconds, five sections. The hook puts the contradiction on screen and skips the "here is a fun fact" opening. The first three seconds have one job, which is to keep the viewer from scrolling.

How do you storyboard a vertical frame? Five rules

With a timed script, break it into shots. The method is the same as for any storyboard, with five extra rules for a vertical frame:

  1. Subject in the centre, filling the middle 60 percent. The top 20 percent is platform interface and your title text; the bottom 20 percent is captions and the button column.
  2. One thing per shot. A vertical frame cannot hold two subjects side by side; to compare two things, alternate shots or stack them vertically.
  3. Six to ten shots, three to five seconds each. Beyond ten shots a thirty-second video turns into a slideshow; under six it drags.
  4. Consistent direction of motion. If one shot pushes upward, do not have the next fall downward; vertical frames make up-and-down very visible.
  5. Start and end states for every shot. As with a longer film, this is the only thing that lets separately generated shots join.

The script, broken down:

Shot Seconds Picture Size Start state End state
1 0–3 Two cubes stacked vertically on black, clear above, cloudy below Close-up Both cubes still Same
2 3–6 A home ice tray goes into the freezer, door closes Medium Door open Door closed
3 6–10 Cross-section of the tray: ice grows from the walls inward, a white core forms in the centre Close-up (diagram) Water clear Centre turns white
4 10–14 Inside an ice machine: the water touches the cold plate on one face only Medium (diagram) Surface still Freezing begins at the top
5 14–22 Cross-section: the ice layer thickens from the top down, bubbles pushed downward and out of the bottom Close-up (diagram) Thin ice layer Thick clear layer
6 22–25 An insulated cup filled with water goes into the freezer, open side up Medium Cup on a table Cup in the freezer
7 25–28 The ice column tipped out: clear on top, white at the bottom Close-up Column standing Same
8 28–32 Back to the two cubes from shot 1 Close-up As shot 1 Same

Eight shots. Shots 3, 4 and 5 are marked "diagram" as a reminder that these are explanatory pictures; legibility matters more than realism for them.

How do you write the prompt, and how many seconds per generation? One shot, five to eight seconds

Only now do you generate. Three rules:

One shot per generation. Five to eight seconds per segment, one row of the storyboard each. The model can hold a short segment together, and you can judge a short segment quickly.

The ratio goes in the first line of the prompt, every time. 9:16, vertical, subject centred, headroom and footroom. Every segment is an independent generation, so every prompt says it again.

Write the start and end states in. Shot 2 ends on "door closed", so shot 3 starts with "the tray in a dark freezer". Separately generated shots join because of those two lines; a simple storyboard to connect every shot covers this in full.

The prompt for shot 5 looks something like this (an illustration; match your model's own format):

Vertical 9:16. Close-up, diagram style. Cross-section of a transparent container: an ice layer thickens slowly from the top down, tiny bubbles pushed downward by the ice and out through the bottom of the container. Clean frame, subject centred, empty space above and below. Start: a thin layer of ice at the water surface. End: clear ice filling the upper half of the container. Static camera, 8 seconds.

When a generation is wrong, run the same prompt again first; change one line only after two takes fail the same way. The diagram shots (3, 4, 5) are usually the hardest to land, because models handle realism far better than explanation. Budget more takes for them, or simplify the picture (film only the ice surface and let the captions do the explaining).

Edit, captions, voice-over: most people are watching muted

Eight segments in a folder are not yet a short-form video.

  • Assemble, then trim to the script's seconds. Generated segments never match the script's timing exactly; cut the extra, and fill a missing second or two with a slow-motion or a freeze; regenerating for that is a waste.
  • Captions are part of the picture. Most viewers watch muted. One sentence at a time in the bottom 20 percent, never a paragraph, large enough to read on a phone without squinting.
  • Something changes every three to five seconds. A new shot, a new caption line or a new sound, at least one of the three. A vertical frame is small and attention drops faster than in landscape.
  • Sound. Record or generate the voice-over from the script; lay ambience under the whole video so eight separately generated shots sound like one world. Sound starts with the picture in the first three seconds; no silent opening.

How long can a Reel, a Short or a TikTok be? The limit is not the target

Vertical short-form is 9:16 at 1080 by 1920. Keep the top and bottom 15 to 20 percent free of anything important, and the right edge free for the button column. Length limits differ by platform and they change:

Platform Limit (checked 2026-08-25) Note
YouTube Shorts 3 minutes Since 15 October 2024, vertical or square uploads up to three minutes are classified as Shorts; music use has separate per-track limits
Instagram Reels Help Center says 3 minutes; the marketing site says 20 minutes Two official sources disagree; go by the limit the app enforces when you post
TikTok Changed several times; not fixed here Check the platform's help page before you post

The limit is a ceiling. Our example runs thirty-two seconds because one point takes thirty seconds to make; when the point is made, stop. Whether anyone watches it afterwards, or whether it earns anything, is a matter of content and the platform's algorithm, and no workflow can promise it. This one only promises that the video you post is complete.

Will a thirty-second short made with AI cost more than $10?

It is the most common question around short-form production, and there are two questions inside it. A quote from a studio or an agency prices people and a service; generating one yourself is a separate number, and it can be worked out.

Using this article's example, thirty-two seconds in eight shots, each generated as a 6-second take, one pass over the whole video is 48 seconds of generation. At the list prices of two video models that publish a per-second rate (official price pages checked 2026-08-25; list prices, limited-time discounts excluded):

Takes per shot Seconds generated MiniMax H3 768P ($0.08/s) Seedance 2.0 720p ($0.15/s)
1 48 about $4 about $7
5 240 about $19 about $36
20 960 about $77 about $144

So, will a thirty-second short cost more than $10? It stays under only when every shot lands first time; at a normal five takes per shot it is $19 to $36.

Generation for one short runs from a few dollars to a few tens of dollars, and the real variable is how many takes each shot needs; the diagram shots (3, 4 and 5) usually eat the most. Voice-over, music and editing tools are extra.

As for "which AI can generate video" and "which app to use for shorts": this workflow is not tied to a tool. Any video model that lets you set 9:16, generate five to ten seconds at a time, and ideally take a reference image or a starting frame will do; for editing and captions, whatever editor you already have is enough.

The whole short-form workflow in one table

Step Do Done when
1 Ratio Choose the platform, choose 9:16 Everything after this uses that ratio
2 Script One point, written by the second, hook in the first three 20 to 45 seconds; you can say what the video is about in one sentence
3 Storyboard Six to ten shots, subject centred, start and end states per shot Table filled, diagram shots marked
4 Generate One shot per generation, 5 to 8 seconds, ratio and states in the prompt Every shot has one accepted take
5 Edit Assemble, trim to time, one caption line at a time, ambience under everything Makes sense muted, works with sound
6 Publish Check the platform's specs and limits Nothing important sits under the interface

Short-form looks like the lowest bar there is: thirty seconds, a phone, an idea. With AI the bar drops further, low enough that many people skip every step in the middle and jump from idea straight to generate.

The skipped steps are the third of the frame that got cropped away, the subject that drifts after second ten, the captions that do not fit in the safe zone. Those are workflow problems, and AI has nothing to do with them. Do the workflow and thirty seconds is just thirty seconds.

RECOMMENDED
TutorialExtended an AI video and the join stutters? No frame repeats — the motion breaks for half a secondTutorialSeedance 2.5 Prompt Skill: Every Reference Needs a Job