How to make a short-form video with AI: script, storyboard, generate in segments, edit
Short-form made with AI fails when you generate thirty seconds at once and crop landscape to vertical. One 30-second explainer, ratio first, one shot per generation, with a checklist.
Just open CapCut, drop your clips into a template, done. You can skip the rest of this article; that is enough to post. About a thousand other shorts are going out through the same template while you read this, so it clearly works. Simple. No need to read on.
Still here? Then let us make one that is not among the thousand.
The usual way to make a short-form video with AI goes like this: think of an idea, write a prompt, hit generate, wait for thirty seconds of video, crop it from landscape to vertical, add captions, post.
Two things break on that road. First, at thirty seconds in one generation, the model has forgotten what you asked for by second ten. Second, cropping landscape to vertical throws away the picture you paid to generate: the subject ends up at the edge and there is nowhere for the captions to go.
A vertical short is a different film from the aspect ratio onward. This is one thirty-second explainer, made from the start.
Can you just crop landscape to vertical? No, vertical is a different film
16:9 and 9:16 differ by far more than a rotation. A landscape frame reads left to right; the subject can stand at one third and leave a wide field of environment. A vertical frame reads top to bottom; the eye lands just above centre, there is almost no room to the sides, and the top and bottom belong to captions and the platform's buttons.
So a shot composed well in 16:9 (a person at the left third, a whole street to the right) becomes, in 9:16, half a body and a lamp post. What you lose is the composition, and no amount of resolution brings it back.
The reverse holds too: a shot generated vertical, subject centred with headroom, becomes a narrow figure between two black bars when forced into landscape.
So where you will publish decides which ratio you generate in, and the ratio is decided before the script is written. If you need both landscape and vertical, that is two storyboards and two rounds of generation.
How do you write a thirty-second script? One point, and the first three seconds keep them
A short-form script differs from any other script in two ways: it makes exactly one point, and the first three seconds have to stop the thumb.
Our example:
Why is convenience-store ice clear, and the ice from your freezer white?
One question, one answer, thirty seconds. Write it by the second:
| Seconds | Section | Content |
|---|---|---|
| 0–3 | Hook | Two ice cubes side by side, one clear, one cloudy. Voice-over: "Same water. Why the difference?" |
| 3–10 | Setup | A home freezer freezes from all sides toward the middle; the air and impurities in the water get pushed to the centre and trapped there, a white core |
| 10–22 | Answer | An ice machine freezes the water slowly from one direction only; the air has somewhere to go and is pushed out along the way, leaving clear ice |
| 22–28 | Turn | You can do it at home: fill an insulated cup, leave it open at the top, freeze it; it freezes from the top down and the upper part comes out clear |
| 28–32 | Close | The two cubes again. Voice-over: "The difference is where the freezing starts." |
Thirty-two seconds, five sections. The hook puts the contradiction on screen and skips the "here is a fun fact" opening. The first three seconds have one job, which is to keep the viewer from scrolling.
How do you storyboard a vertical frame? Five rules
With a timed script, break it into shots. The method is the same as for any storyboard, with five extra rules for a vertical frame:
- Subject in the centre, filling the middle 60 percent. The top 20 percent is platform interface and your title text; the bottom 20 percent is captions and the button column.
- One thing per shot. A vertical frame cannot hold two subjects side by side; to compare two things, alternate shots or stack them vertically.
- Six to ten shots, three to five seconds each. Beyond ten shots a thirty-second video turns into a slideshow; under six it drags.
- Consistent direction of motion. If one shot pushes upward, do not have the next fall downward; vertical frames make up-and-down very visible.
- Start and end states for every shot. As with a longer film, this is the only thing that lets separately generated shots join.
The script, broken down:
| Shot | Seconds | Picture | Size | Start state | End state |
|---|---|---|---|---|---|
| 1 | 0–3 | Two cubes stacked vertically on black, clear above, cloudy below | Close-up | Both cubes still | Same |
| 2 | 3–6 | A home ice tray goes into the freezer, door closes | Medium | Door open | Door closed |
| 3 | 6–10 | Cross-section of the tray: ice grows from the walls inward, a white core forms in the centre | Close-up (diagram) | Water clear | Centre turns white |
| 4 | 10–14 | Inside an ice machine: the water touches the cold plate on one face only | Medium (diagram) | Surface still | Freezing begins at the top |
| 5 | 14–22 | Cross-section: the ice layer thickens from the top down, bubbles pushed downward and out of the bottom | Close-up (diagram) | Thin ice layer | Thick clear layer |
| 6 | 22–25 | An insulated cup filled with water goes into the freezer, open side up | Medium | Cup on a table | Cup in the freezer |
| 7 | 25–28 | The ice column tipped out: clear on top, white at the bottom | Close-up | Column standing | Same |
| 8 | 28–32 | Back to the two cubes from shot 1 | Close-up | As shot 1 | Same |
Eight shots. Shots 3, 4 and 5 are marked "diagram" as a reminder that these are explanatory pictures; legibility matters more than realism for them.
How do you write the prompt, and how many seconds per generation? One shot, five to eight seconds
Only now do you generate. Three rules:
One shot per generation. Five to eight seconds per segment, one row of the storyboard each. The model can hold a short segment together, and you can judge a short segment quickly.
The ratio goes in the first line of the prompt, every time. 9:16, vertical, subject centred, headroom and footroom. Every segment is an independent generation, so every prompt says it again.
Write the start and end states in. Shot 2 ends on "door closed", so shot 3 starts with "the tray in a dark freezer". Separately generated shots join because of those two lines; a simple storyboard to connect every shot covers this in full.
The prompt for shot 5 looks something like this (an illustration; match your model's own format):
Vertical 9:16. Close-up, diagram style. Cross-section of a transparent container: an ice layer thickens slowly from the top down, tiny bubbles pushed downward by the ice and out through the bottom of the container. Clean frame, subject centred, empty space above and below. Start: a thin layer of ice at the water surface. End: clear ice filling the upper half of the container. Static camera, 8 seconds.
When a generation is wrong, run the same prompt again first; change one line only after two takes fail the same way. The diagram shots (3, 4, 5) are usually the hardest to land, because models handle realism far better than explanation. Budget more takes for them, or simplify the picture (film only the ice surface and let the captions do the explaining).
Edit, captions, voice-over: most people are watching muted
Eight segments in a folder are not yet a short-form video.
- Assemble, then trim to the script's seconds. Generated segments never match the script's timing exactly; cut the extra, and fill a missing second or two with a slow-motion or a freeze; regenerating for that is a waste.
- Captions are part of the picture. Most viewers watch muted. One sentence at a time in the bottom 20 percent, never a paragraph, large enough to read on a phone without squinting.
- Something changes every three to five seconds. A new shot, a new caption line or a new sound, at least one of the three. A vertical frame is small and attention drops faster than in landscape.
- Sound. Record or generate the voice-over from the script; lay ambience under the whole video so eight separately generated shots sound like one world. Sound starts with the picture in the first three seconds; no silent opening.
How long can a Reel, a Short or a TikTok be? The limit is not the target
Vertical short-form is 9:16 at 1080 by 1920. Keep the top and bottom 15 to 20 percent free of anything important, and the right edge free for the button column. Length limits differ by platform and they change:
| Platform | Limit (checked 2026-08-25) | Note |
|---|---|---|
| YouTube Shorts | 3 minutes | Since 15 October 2024, vertical or square uploads up to three minutes are classified as Shorts; music use has separate per-track limits |
| Instagram Reels | Help Center says 3 minutes; the marketing site says 20 minutes | Two official sources disagree; go by the limit the app enforces when you post |
| TikTok | Changed several times; not fixed here | Check the platform's help page before you post |
The limit is a ceiling. Our example runs thirty-two seconds because one point takes thirty seconds to make; when the point is made, stop. Whether anyone watches it afterwards, or whether it earns anything, is a matter of content and the platform's algorithm, and no workflow can promise it. This one only promises that the video you post is complete.
Will a thirty-second short made with AI cost more than $10?
It is the most common question around short-form production, and there are two questions inside it. A quote from a studio or an agency prices people and a service; generating one yourself is a separate number, and it can be worked out.
Using this article's example, thirty-two seconds in eight shots, each generated as a 6-second take, one pass over the whole video is 48 seconds of generation. At the list prices of two video models that publish a per-second rate (official price pages checked 2026-08-25; list prices, limited-time discounts excluded):
| Takes per shot | Seconds generated | MiniMax H3 768P ($0.08/s) | Seedance 2.0 720p ($0.15/s) |
|---|---|---|---|
| 1 | 48 | about $4 | about $7 |
| 5 | 240 | about $19 | about $36 |
| 20 | 960 | about $77 | about $144 |
So, will a thirty-second short cost more than $10? It stays under only when every shot lands first time; at a normal five takes per shot it is $19 to $36.
Generation for one short runs from a few dollars to a few tens of dollars, and the real variable is how many takes each shot needs; the diagram shots (3, 4 and 5) usually eat the most. Voice-over, music and editing tools are extra.
As for "which AI can generate video" and "which app to use for shorts": this workflow is not tied to a tool. Any video model that lets you set 9:16, generate five to ten seconds at a time, and ideally take a reference image or a starting frame will do; for editing and captions, whatever editor you already have is enough.
The whole short-form workflow in one table
| Step | Do | Done when |
|---|---|---|
| 1 Ratio | Choose the platform, choose 9:16 | Everything after this uses that ratio |
| 2 Script | One point, written by the second, hook in the first three | 20 to 45 seconds; you can say what the video is about in one sentence |
| 3 Storyboard | Six to ten shots, subject centred, start and end states per shot | Table filled, diagram shots marked |
| 4 Generate | One shot per generation, 5 to 8 seconds, ratio and states in the prompt | Every shot has one accepted take |
| 5 Edit | Assemble, trim to time, one caption line at a time, ambience under everything | Makes sense muted, works with sound |
| 6 Publish | Check the platform's specs and limits | Nothing important sits under the interface |
Short-form looks like the lowest bar there is: thirty seconds, a phone, an idea. With AI the bar drops further, low enough that many people skip every step in the middle and jump from idea straight to generate.
The skipped steps are the third of the frame that got cropped away, the subject that drifts after second ten, the captions that do not fit in the safe zone. Those are workflow problems, and AI has nothing to do with them. Do the workflow and thirty seconds is just thirty seconds.
