Article cover for AI Video Previsualization: We Built a Box Cat in Blender

AI Video Previsualization: We Built a Box Cat in Blender

Before generating the paper boat and stray cat, we used Claude Code and Blender MCP to build a 30-second, eight-camera previz for blocking, framing, occlusion and the final exit.

Tutorial
Published
Updated

The finished paper-boat film has rain, wet fur, puddles and neon. Our first screening featured a cat made of boxes.

It had no fur. The boat did not glow. Yet this 30-second Blender blockout answered several expensive questions early: where the cat stood, how the boat approached, whether the camera could see the action and which way both subjects would leave.

We had Claude Code operate Blender through Blender MCP, building simplified versions of the cat, paper boat, cardboard box, store, glass door, puddle and street. Eight cameras covered P01 through P08. The result was our previsualization, or previz: one hard look at space and camera logic before the final video clips existed.

30-second Blender clay previz; the original 1920×1080, 24fps frame is preserved inside the Article player's 16:10 presentation frame.

An eight-panel storyboard can still lose the room

The formal storyboard already carried the whole sequence: the cat curled beside the box, the boat approaching, a paw testing the water, the nose moving closer, the boat lifting and the pair leaving. Landscape and portrait each had a dedicated board with its own composition.

Landscape eight-panel storyboard showing the cat beside the ruined box, the boat approaching and glowing, the cat testing the water, the boat lifting and both subjects leaving for the rainy street.

Formal 16:9 storyboard submitted for the Desktop generation; read P01–P08 left to right and top to bottom.

Portrait eight-panel storyboard showing the cat beside the ruined box, the boat approaching and glowing, the cat testing the water, the boat lifting and both subjects leaving for the rainy street.

Formal 9:16 storyboard submitted for the Mobile generation; the same P01–P08 sequence is recomposed vertically.

For the basic first-frame, last-frame and cross-shot continuity workflow, see Create Video with AI: Use a Simple Storyboard to Connect Every Shot. This project added Blender because all eight shots shared one store, one puddle and one exit route. Distances left vague on a flat board become impossible to ignore in 3D.

Those eight panels tell us what each shot should contain. Move them into Blender and three-dimensional space starts asking follow-up questions:

  • How far apart are the box, cat and storefront glass?
  • Will the cat or box hide the boat as it crosses the puddle?
  • When the P06 nose close-up opens into the P07 medium shot, where should the store door appear?
  • Can the P08 route down the street continue in the same direction as the previous shot?

The blockout could ignore rain, wet fur and neon. Once the geometry held together, the later references and generated clips had a baseline. Without that baseline, eight beautiful panels can turn into eight parallel universes at the edit.

Claude Code called the moves; Blender kept the evidence

We chose the story beats and wrote down what each camera needed to answer. Claude Code took the operating instructions and used Blender MCP to create, position and arrange the objects and cameras. Blender kept the 3D scene and rendered the videos and board.

The surviving .blend file has these specifications:

Item Recorded value
Timeline Frames 1–720 at 24fps, 30 seconds
Frame 1920×1080
Cameras Eight, named CAM_01 through CAM_08
Scene objects 101 objects
Main outputs Clay video, segmentation-ID video and board_2x4.png

We replaced much of the object-by-object clicking with instructions that arranged the scene. Then we watched the renders. MCP can execute “move the camera back.” It cannot decide whether the shot deserves that move. The tool works hard; the director's chair remains occupied.

For the difference between pans, tracking shots, dollies and orbit moves, see A Complete Guide to AI Video Camera Moves. The blockout made the camera run those ideas through a real scene before we committed them to the prompt.

Eight cameras rehearsed 30 seconds

The blockout placed all eight beats on one timeline:

Camera Time Action under review
CAM_01 0–2 seconds The cat curls beside the box
CAM_02 2–6 seconds The cat looks up at the boat
CAM_03 6–10 seconds The cat reaches out a paw
CAM_04 10–14 seconds Close-up of the paper boat
CAM_05 14–18 seconds The cat steps toward the water
CAM_06 18–21 seconds The cat's nose approaches the boat
CAM_07 21–25 seconds The boat lifts and the door enters the shot
CAM_08 25–30 seconds The cat and boat leave the storefront

These timings belong to the previz. They do not form an edit decision list for the generated film. The finished landscape version runs 45 seconds and the portrait version 46.5 seconds. Generation stretched some actions, and the two ratios developed different pacing. At this stage, the blockout only had to prove that all eight actions could run in sequence.

Clay catches occlusion; segmentation ID catches visual soup

The clay render gives every object the same grey treatment. Fur, neon and rain leave the room, so shot size, silhouette, occlusion and ground direction become easier to read. If the cat hides the boat, or a close-up cannot hold both subjects, grey geometry reports the crime immediately.

The segmentation-ID render separates object groups with flat colours. Beauty is outside its job description. Its job is to stop the cat, boat, puddle, storefront and foreground from melting into visual soup, and to expose silhouettes that vanish across a cut.

Segmentation-ID output of the same 30-second previz, using flat colours to separate object groups; 24fps with no audio track.

board_2x4.png takes one representative view from each camera and arranges all eight in a 2x4 grid. The video reveals movement. The board lays all shot sizes on the table, where repetition, direction jumps and the close-up-to-long-shot rhythm have fewer places to hide.

Eight-panel Blender blockout showing the cat, paper boat, box, storefront, street and camera positions as simple grey geometry.

board_2x4.png places P01–P08 in one 3D space to check blocking, direction and the final exit. The eight panels are separated below so they remain readable when the overview is reduced.

P01–P02

P01 Blender blockout: the box-shaped cat curls beside the cardboard box. P02 Blender blockout: the cat looks up towards the distant paper boat.

P03–P04

P03 Blender blockout: the cat approaches the boat and extends a forepaw. P04 Blender blockout: close view of the boat with the cat behind it.

P05–P06

P05 Blender blockout: medium view of the cat, box and paper boat. P06 Blender blockout: frontal close-up of the cat's nose and the paper boat.

P07–P08

P07 Blender blockout: the boat lifts as the cat and store door enter the wide view. P08 Blender blockout: view through the doorway as the cat and boat leave.

Blockout, formal storyboard, finished film: three layers, three jobs

Artifact What it checked in this project
Blender blockout Position, eyelines, occlusion, cameras and the exit route in one space
Formal eight-panel boards Shot order and ratio-specific composition for 16:9 and 9:16
Finished films The actions, visual treatment, timing and edited result the model actually produced

The surviving blockout output is 16:9. A separate 9:16 formal storyboard handled portrait composition. Both final films contain the visual beats from P01 through P08. Door states and some action timings changed, and we kept those differences. The cat, boat and store still occupy the same basic spatial relationship; both versions also send the cat and boat away together.

Every P01–P08 comparison, the Act 2 tail-frame handoff, Act 3 first frame and final joins are covered in Does Seedance 2.5 Read Storyboards?. The full story, reference images and three-act generation process are in Making an AI Video: A Paper Boat Came Back for a Stray Cat.

The blockout never went straight into Seedance

The video model received separate formal storyboards and reference images for the cat, boat, box and convenience store. What we can verify is simple: we watched the Blender blockout before video generation and used it to review position, shots and movement.

We found no evidence that the .blend, clay render or segmentation-ID render went directly into Seedance. We will not invent that step. The previz shaped our decisions; the formal boards and reference images carried the visual instructions into generation.

When an AI video deserves a 3D blockout

A single subject against a fixed background, with little camera movement, may only need a simple storyboard. Previz becomes easier to justify when:

  • three or more shots take place in one shared space;
  • the subject, prop and camera all move;
  • one shot's screen direction constrains the next;
  • a close-up must connect to a wide shot, a door must move or a character must exit;
  • landscape and portrait versions need to share the same spatial logic.

Stop when the blockout has answered those questions. Our box cat had no fur and the paper boat had no neon. Before the first video generation, they already knew where they would meet and how they would leave.

SERIESCreate Video with AI4 articles
  1. 01Create Video with AI: your prompt is your video
  2. 02Create Video with AI: start with text, an image or a script?
  3. 03Create Video with AI: use a simple storyboard to connect every shot
  4. 04Making an AI Video: A Paper Boat Came Back for a Stray Cat
RECOMMENDED
TutorialSeedance 2.5 Prompt Skill: Every Reference Needs a JobTutorialHow to make a short film with AI: the full workflow, with real costs