How to write Seedance 2.0 prompts: a practical guide from subject and action to camera work

The official Seedance 2.0 prompt structure, turned into a field-by-field workflow: what each piece of text controls, and which field to fix when a generation fails.

Tutorial
Published
Updated

When writing Seedance 2.0 prompts, what actually helps is not memorising a "magic formula" — it is giving every word a job. The official guide breaks a full prompt into subject, action, scene, light and colour, camera, style, image quality and constraints; for complex videos, it recommends writing each segment out in shot order, covering the subject, position, action and how it is filmed.

You do not have to fill all eight fields every time. For a first pass, remember the short version:

who + doing what + where + how it's filmed + what it feels like + what must not appear

This is not a fixed sentence pattern to fill out every time. A simple image may need one sentence; only when the model starts guessing wrong do you walk this order looking for the missing information.

First decide what each piece of text controls

The most common beginner problem is not that the prompt is too short — it is that one sentence mixes in many adjectives with no clear purpose. Dividing the information up works better than piling words on:

Field The question it answers Example
Subject Who is the frame following? A young woman in a dark blue raincoat, carrying an old suitcase
Action What exactly is she doing? Stops, turns her head towards the platform entrance
Scene Where is she? An old railway station late at night, puddles reflecting on the ground
Light and colour Where does the light come from — does the image lean cold or warm? Overall cold blue, tungsten platform lamps adding a little warm yellow from the side
Camera How does the audience see it? Medium tracking shot, pushing in slowly when she stops
Style What rule does the whole image share? Realistic cinematic feel
Image quality How much detail? 4K, crisp detail
Constraints What absolutely must not happen? No extra passers-by, no subtitles or logos

The field beginners drop most often is light and colour — it usually gets folded into style, so a phrase like "realistic cinematic feel" ends up responsible for the medium, the lighting and the colour temperature all at once, and the model can only guess. Pull light out into its own field, and when a generation fails you know where to look.

When the result is wrong, first ask which field is wrong. If the character is unrecognisable, fix the subject or the reference image; if the action order is wrong, fix how the events are broken up; only when the frame is simply chaotic should you consider deleting camera or style instructions that compete with each other.

Give the subject only two or three stable identifiers

The official Seedance 2.0 guide recommends locking down "who is doing what" first, and identifying the subject with a few clear, static features that do not contradict each other. In other words, there is no need to write the character head-to-toe like a novel; two or three features that will still be visible later are usually more useful.

For example:

A young woman in a dark blue raincoat, with short black hair, carrying an old suitcase.

When you mention her again, call her "the raincoat woman" every time. Do not rotate through traveller, girl, protagonist and stranger. The more her name changes, the more likely the model treats one character as several people.

When using reference material, do not just write "follow the reference". Spell it out:

Use the woman in @Image 1 as the lead, keeping the dark blue raincoat, short black hair and old suitcase; she stands at the centre of the platform in Image 2.

The official documents refer to uploads in order — "Image 1, Image 2" or "Video 1, Video 2". The number only points at the material; the sentence still has to say whether you are taking the character, the composition, the action or the camera movement from it.

Use shots to structure your idea

Cram arrival, running, falling, getting up and a farewell into one sentence, and the model has to guess the order, the rhythm and the camera for you. For complex videos, the official Seedance 2.0 guidance is to split along the timeline into shots, with each shot covering at least "who, where, doing what, filmed how".

For example:

Shot 1: Old railway station platform. The raincoat woman stands behind the yellow safety line, both hands on the suitcase. Static wide shot.

Shot 2: Hearing the train, she turns her head towards the entrance, right hand still on the suitcase. Medium shot from the side, pushing in slowly.

Shot 3: The train pulls in, wind lifting the hem of her raincoat; she takes a step back. Close-up holds on her hesitation.

This is not prose chopped into three lines. It is working with the model in a form the model understands.

This article only demonstrates the Seedance way of writing. If your problem is "every clip looks fine on its own but they jump when cut together", that is shot continuity, not prompting.

The official Seedance 2.0 tutorial notes that second-level timing such as "0–2s, 2–4s" can drift, and the 2.5 tutorial does provide second-level control — but from a creative standpoint, shot order is what decides what the audience gets to see. Even if you are only making something for fun, be clear about how you intend the model to shoot it.

You are still the distinguished director here — not the sucker whose credit card happens to be on file.

Write action as change the camera can see

"Very sad", "running powerfully" and "interacting naturally" all leave the model to interpret. Rewriting emotion as physical change makes the instruction executable:

  • "Head down, shoulders trembling slightly, fingers clutching the hem" is more concrete than "very sad".
  • "First turns her head towards the entrance, then releases the left hand holding the suitcase" is clearer than "hesitates for a moment".
  • "The right foot steps back half a pace first, body still facing the train" adds the body part, the direction and the extent.

Order multiple actions by when they happen. If a character turns, picks something up and walks forward, write three connected events — not three bare verbs.

Explosive running, jumping, tumbling, collisions and multi-person contact are inherently harder than slow single-person action. On a first generation, reduce how much happens at once, confirm the subject and scene are right, then raise the difficulty step by step. That is not lowering your ambition — it is making every re-roll answer one specific question.

Camera language should say what the audience sees

Seedance's official notes on camera language list the basic moves — tracking, panning left and right, trucking left and right — and point out that advanced users can chain several moves into a single take. Framing runs from wide to full, medium and close-up; viewpoints include aerial, underwater, high angle, low angle, macro, and "with something as foreground". The terms are not the point; their value is telling the model how visual information should arrive.

Effect you want Camera writing to try first
Establish the person's relationship to the space Static wide shot, the character at the centre of the platform
Draw attention to the face Push slowly from medium shot to a facial close-up
Follow the character through the space Steady side tracking shot, holding a medium frame
Reveal a clue that was out of frame Pan right to reveal the suitcase by the entrance
Emphasise pressure or vulnerability Low-angle shot up, or high-angle shot down

The official examples also show combined moves and single takes, but those suit people who already control basic shots. Starting out, pick one main move per shot; if a single sentence demands push, pull, orbit, zoom and a viewpoint change at once, the model receives competing filming assignments.

A simple test: delete the terminology — can you still say what the audience sees first, and what they see next? If not, the camera words were decoration, not design.

Style should be a rule the whole video shares

The official examples show explicit styles — 2D, 3D, pixel, felt, clay, illustration — and demonstrate steering the look through genre, light, colour and mood. The stable way to write it is not stacking words that pull against each other ("cinematic, epic, dreamlike, hyper-real, retro"), but choosing one visual direction and adding light and colour.

For example:

Realistic cinematic style, cold blue late-night rain, tungsten platform lamps adding a little warm yellow reflection, low saturation.

That one line controls the sense of medium, the time and weather, the dominant colour temperature, a local light source and the saturation. If you have a style reference image, treat it as the rule for the whole video — don't let a second composition reference smuggle in a different material and palette.

When the style drifts, first check whether the reference materials conflict with each other, then tighten the written style constraints. The official Seedance 2.0 guide also suggests converting a reference image into the target style first, then generating video from it — usually clearer than asking the model to keep a photo's composition while simultaneously repainting the whole frame in another medium.

Images, video and audio each take exactly one job

Seedance 2.0 can inherit a character's appearance, a visual style or a composition from images, and reference a subject, motion, effects or camera movement from video; audio can supply a voice, a melody or dialogue. More material is not automatically more accurate — what matters is whether the prompt states each asset's responsibility.

Asset Assignment
Image 1 The lead's appearance and clothing
Image 2 Platform composition and the warm–cool light ratio
Video 1 The speed and path of the slow side tracking shot
Audio 1 Background rhythm for the whole video, not repurposed as dialogue

Written out:

Character appearance follows Image 1; scene composition follows Image 2. Use the side tracking from Video 1, but the character only turns her head and steps back once. Audio 1 is the background music for the whole video.

When spatial relationships matter, the official guide prefers an image that directly demonstrates where people, objects and the scene sit, over a long string of positional text. If you need several angles, you can supply multi-view references of the same subject — but every image must stay consistent; you cannot lock the character with one image while another quietly changes the outfit.

Constraints are only for the errors you truly cannot accept

Constraints are best spent on the most important no-go zones:

  • No new characters; keep the same outfit and hairstyle;
  • No subtitles, text, watermarks or brand logos;
  • Do not change the platform layout from Image 2;
  • Keep the same realistic cold-toned style throughout.

The official Seedance 2.0 guide specifically advises that if you don't want subtitles or text, say "no subtitles, no text" explicitly — and remove any text already present in your reference material first. This lowers the odds of stray text, though it is not a hard guarantee. Logos and watermarks belong on the constraints list too.

Do not let the constraints grow into a second essay. A dozen "don'ts" can contradict each other and dilute the ones that matter. Pick the two or three errors that make an output unusable, and leave the rest until you have seen the first version.

Run a final check with one complete template

Assembling the rainy-station example gives you this editable skeleton:

Subject: use the woman in Image 1 as the lead, keeping the dark blue raincoat, short black hair and old suitcase.

Scene and style: the old railway station platform from Image 2, raining late at night; realistic cinematic style, cold blue environment, tungsten lamps adding a little warm yellow reflection, low saturation.

Shot 1: The lead stands behind the yellow safety line, both hands on the suitcase. Static wide shot.

Shot 2: Hearing the train, she first turns towards the entrance, then releases her left hand; her right hand stays on the suitcase. Medium shot from the side, pushing in slowly.

Shot 3: The train pulls in, wind lifting the hem of her raincoat; she takes a step back. Close-up holds on her hesitation.

Constraints: no passers-by added; clothing, hairstyle and suitcase unchanged; no subtitles, no text, no logos, no watermarks.

A prompt is not a writing contest. It is a miniature storyboard — for the model to execute, and for you to debug. When a result fails, you should be able to point at the field to change, instead of rewriting the whole paragraph longer.

Official sources

This article uses the Seedance 2.0 guide as its primary basis for prompting and multimodal references, supplemented by the camera-language and style examples from the Seedance 1.0 Pro / Pro Fast guide.

Last checked: 2026-07-28. The specific claims follow the two publicly readable BytePlus English documents — the eight prompt fields map item by item to their "Advanced prompt formula". The two Volcano Engine Chinese pages are linked pending review; their body text has not been obtained.

SERIESVideo Generation Models4 articles
  1. 01Seedance 2.5 Prompt Skill: Every Reference Needs a Job
  2. 02How to write Seedance 2.0 prompts: a practical guide from subject and action to camera workYou are here
  3. 03One idea, two models: Seedance writes structure into the text, MiniMax writes it into the protocol
  4. 04How to prompt MiniMax H3: the formula isn't in the guide, it's in the API
RECOMMENDED
TutorialExtended an AI video and the join stutters? No frame repeats — the motion breaks for half a secondTutorialAI Video Previsualization: We Built a Box Cat in Blender