
Can Planning the Action in Blender Save AI Video Tokens?
We already subscribe to Claude Code and Codex. Could we let an agent block out the action in Blender before paying for video generation? We tried a fight scene and a drawn camera route.
Can planning the action in Blender save some of the tokens we spend generating AI video? That's what we wanted to try.
We subscribe to Claude Code. We subscribe to Codex. We're paying for both subscriptions. Might as well put them to work. If an agent can operate Blender locally, could we use it to arrange the characters and camera, fix what looks wrong, and only then send the result off for video generation?
Every generation costs credits. Submit a scene before you've worked out the action, and you're paying for each round of trial and error. We'd like to do more of that locally: watch the blockout move, adjust the camera, try again. The agent still uses our subscription allowance, but changing a move doesn't immediately mean paying for another generated video.
Having tried it, we think this workflow is worth using. We can see the blockout's actions and the drawn route in the finished clips, though some details changed along the way. We haven't calculated the savings yet. We have, however, seen where the references help.
Last time, a blocky cat. This time, a fight.
In our earlier Blender blockout experiment, we checked the positions of a cat, a paper boat, and a storefront. That article covered using the blockout to inspect the scene; it didn't include a comparison with a video generated directly from that blockout.
This time, we sent it in.
In Blender, an orange character on the left fights unarmed. A teal character on the right carries a sword. They exchange attacks for six seconds. The models don't need much detail yet; we just need to show the action. Then we hand the clip to Seedance 2.5: replace the left character with the character in red, the right one with the woman in purple, and the background with our chosen scene. Follow this fight.
Here's an English translation of the prompt we submitted in Chinese. We've replaced the asset UIDs with readable names.
[Generation goal] Generate a fight video using [motion reference video] as the reference. The fighting should feel powerful, with a sense of impact.
[Motion reference video] is the action guide. [Purple character reference image] is the character on the right in the action guide. [Red character reference image] is the character on the left in the action guide, without a sword, using only punches and kicks. [Scene reference image] is the actual setting.
Replace the setting in the action guide with [scene reference image], the right character with [purple character reference image], and the left character with [red character reference image].
Generate both characters' movements according to the choreography in [motion reference video].
[Restriction] [Red character reference image] must not hold any blades or swords. Use only punches and kicks.
Most of that prompt says what to replace with what. As for the choreography, we handed it over with one instruction: follow the motion reference.
With text alone, we'd have to describe who raises the sword first, which way it swings, how the other person blocks it, and how far apart they stand. Then there's the camera. Blender has already shown all of that, so we left out much of the description. We still have to choreograph the fight, but we can watch it as we make changes, instead of imagining the action while wondering whether a sentence actually explains it.
It follows the action. Then the character in red changes a move.
We used FFmpeg to extract frames in our earlier article on why extended AI video clips can stutter at the join. Here, we used it again to put the blockout and output side by side at the same timestamps.
The generated fight runs at 24 fps. Frame numbers below start at zero, and times were checked against the original file timestamps. We haven't stretched the blockout or interpolated frames.

| Time / frame in both clips | Blender blockout | Generated output |
|---|---|---|
| 0.500 s / 12 | Unarmed guard on the left, sword on the right | Same roles, with a closer camera |
| 1.000 s / 24 | Sword arm rises and draws behind the body | The woman in purple reaches the same stage of the motion |
| 2.000 s / 48 | Body lowers after the swing | The woman in purple also lowers her body; her legs are outside the frame |
| 2.500 s / 60 | Rising and approaching the opponent | The characters close the distance again |
| 3.750 s / 90 | The unarmed character punches forward | The character in red kicks |
| 4.000 s / 96 | Arm still extended in a punch | Leg still raised in an attack |
The sword lift, downward swing, lowered stance, and recovery line up. Then the character in red counters with a kick where the blockout has a punch.

A kick doesn't violate our instruction to use “only punches and kicks.” But we did animate a punch. If that particular punch matters to the scene, we'd still have work to do here. The earlier exchange lines up, though. We can see the blockout's influence.
Something else caught my attention: the orange and teal models don't appear in the comparison frames. The simple ground plane has become stone paving, a reclining Buddha, and distant mountains. The simplified models' motion carries over while their appearance changes. From this result, we'd guess ByteDance has become fairly good at this kind of replacement. At least we don't see an orange mannequin suddenly turning up in the sampled frames.
The camera has moved closer, too. The blockout shows full bodies; some leg movements fall outside the generated frame. We didn't write a long camera prompt, and the exchange still came through. If keeping both characters fully in shot were essential, we'd need to add that requirement.
As for impact, we only wrote one brief request: make the fighting feel powerful. There's room to work on that. We wouldn't jump to saying the model can't do it when we haven't spent much time on the prompt ourselves.
I left the extra four seconds on purpose
The blockout lasts six seconds. I requested ten seconds of output. I wanted to see what would happen after the reference ran out.
They kept fighting. The woman in purple swings her sword; the character in red dodges, blocks, and counters. They return to a standoff near the end.

It doesn't freeze on the blockout's last pose. It adds another exchange. We can't put an exact boundary at six seconds and say that's where it starts improvising—it already changed a punch into a kick earlier. The extra four seconds were there to see what that continuation would look like.
If we can act out a fight, can we draw a camera move?
A while ago, people were trying another idea: draw a line on an image and see whether a video model follows it with the camera. We tried it with a harbor scene. Start beside the little blue boat, skim the water around the white boats, enter the waterfront buildings, then climb toward the church. Here's the route.

Without the line, there's quite a bit to explain. There are two white boats—which side do we go around? Which opening on shore should we face afterward? The church is up above, but do we fly straight toward it or pass through the lower buildings first? A camera move that feels obvious in your head becomes a series of turns you have to describe.
That's the convenience of drawing it: “Start here, go through here, end there.” You can point at the image. We wanted to see whether this could save some of the time spent writing prompts and explaining the route again.
Higgsfield's 3D Jutsu takes a related approach: an agent builds an editable 3D scene, then you adjust objects, cameras, and animation. Arranging the scene before generating the video is close to what we want to do with Blender locally.
The line doesn't remove the need for a prompt. We still have to say how low to fly, where to rise, how to pass through windows—and that the purple line must not appear in the video. We even wrote, “Do not turn the markings into light trails, ropes, or scene objects.” Use the route, please. Leave our annotations out of the shot.
We split the twelve seconds into four stages: 0–3 seconds from the blue boat to the white boats, 3–6 circling the boats, 6–9 entering the buildings, and 9–12 rising to the church. The boats stay moored. The camera moves.
The window entry is worth pausing on. At about seven seconds, the camera passes through the mint-green building's window frame. The bright harbor falls behind, and a dark passage opens ahead. We didn't draw the interior. It filled in that stretch itself.

| Requested movement | What appears in the output |
|---|---|
| 0–3 s: skim from the blue boat toward the white boats | At 0.000 s / frame 0, the blue boat is close on the left; by 2.500 s / frame 60, the camera approaches a white boat's bow |
| 3–6 s: make one full circuit around both white boats | At 3.000 s / frame 72, the camera passes another boat side; by 4.000 s / frame 96, it heads ashore, leaving early. These frames don't establish a complete circuit |
| 6–9 s: enter the buildings | Still outside at 6.500 s / frame 156; through the window frame at 7.000 s / frame 168; inside the passage by 7.500 s / frame 180 |
| 9–12 s: rise toward the church | Looking out from an archway at 10.000 s / frame 240; outside the bell tower at 12.000 s / frame 288 |
The window passage and climb to the church are there. The sampled frames don't show purple lines or numbers. But the boat circuit ends early, so the timing differs from our instructions. Watching once, you might think it looks close enough. Pausing at matched timestamps makes the differences easier to see.
Taking this into Blender would let us preview the camera traveling along the route. We could fix an obstruction or an awkward turn locally. Before going further, though, we wanted to try one more thing.
Take the line away. It still gets there.
What if we remove the route guide and generate another clip using only the clean harbor image and text?
It still gets there. The blue boat, white boats, building passages, and church all remain in the sequence. The text already explains quite a lot. The model can work without the guide image.
We have not verified the full submitted prompt and model settings for the version without the guide. Treat these as two example outputs, not a controlled single-variable experiment.



Left: with the route guide. Right: without it. Both clips are shown at the same timestamps, with zero-based frame numbers and no speed adjustment.
The difference at four seconds is easy to spot. The guided version is heading ashore, while the other is still beside the boats. At seven seconds, both reach an opening in the mint-green building. At ten, they use different passage exits. Their final compositions differ, too: the guided version is closer to the bell tower; the other includes the dome and the sea behind it.
The version without the guide has a spacious ending and a route of its own. These two clips aren't enough to say a drawn line is always better, or to attribute every difference to it.
I'd still draw the line, though. Especially when I've already decided which boat to start beside and which window to enter. An image makes that easier to explain and check later. Being able to generate without the image doesn't tell us how long it would take to describe the route from scratch.
The same goes for a fight. With that many movements between two characters, it's more direct to act it out and point to the move you want to change. When you know exactly what you want, a reference is worth preparing.
We're trying to reduce the moments where “I meant this” turns into “you went somewhere else.” We saw the references provide direction in this attempt. Measuring how much error or how many retries they save will take more runs.
Why does it still change things when it has a reference?
Look again at 3.750 seconds in the fight: the blockout punches, but the character in red kicks. In the harbor clip, we're already heading ashore at four seconds, when we'd planned to still be circling the boats. The broad direction survives; details drift. We've shown it what to do. Why does this still happen?
We looked through the Seedance technical reports. The 1.5 Pro report explicitly describes a Diffusion Transformer, with MMDiT as the basis for joint audio-video generation. The 2.0 model card explains that text, images, audio, and video can be supplied together, but doesn't give enough detail about how references enter the generation process for us to trace it further. The public material helps us understand the series' technical background. It doesn't reconstruct the full workings of 2.5. Seedance 1.5 Pro technical report, Seedance 2.0 model card
Start with publicly described diffusion methods. A model begins with noise and progressively generates an image, guided by text or other conditions along the way. Latent Diffusion uses cross-attention to bring conditions such as text and bounding boxes into generation. The model learns to produce suitable images given those conditions. Supplying them doesn't, by itself, make every step execute like an animation instruction in Blender. Latent Diffusion paper
What if we strengthen the guidance? Classifier-Free Guidance combines conditional and unconditional model predictions to steer generation; its paper also discusses the trade-off between quality and diversity. Stronger guidance still doesn't add a check that says “frame 90 must contain a punch.” Locking a pose or path requires an appropriate constraint. Classifier-Free Guidance paper
If Seedance 2.5 uses references through this kind of learned conditioning, that would help explain how a reference can be useful while some details still drift. The “if” matters: we don't know which guidance implementation 2.5 uses, and these papers don't tell us why this particular punch became a kick.
We'd use this workflow again. Act out a complicated exchange in Blender. Draw the intended camera route. That gives us something concrete to compare against. Then check the finished clip for the punch or turn that must stay as planned. The source clips, prompt, and matched frames are above, so you can see what carried over and what changed.
Our conclusion is simple: arrange the action and camera locally, get them working, then spend the credits to generate. We still need to calculate the token savings. But we don't have to pay for a new video every time we're still figuring out a move.
