A structure that works
Think of a prompt as a tiny shot list. One or two sentences of action, one phrase of camera, one line of sound. Long poetic descriptions add little; specific verbs and order matter more.
- Subject and action. Who is on screen and what they do, in order. “The chef tosses vegetables and a burst of flame rises.”
- Cuts or camera. “Cut to a close-up of the pan, then cut back to the chef.” Or “Slow push in”, “Handheld, quick cuts”, “Gentle tracking shot”.
- Dialogue. Quote the exact words and describe the delivery: “says in a deep gravelly voice: “Storm’s coming.””
- Sound. A final “Sound:” line: “crashing waves, howling wind”. Image to video always returns audio, so say what it should contain.
Image to video: describe change, not the picture
Your start image already defines the subject, the setting and the framing. Spend the prompt on what changes: movement, expression, a line of dialogue, a cut. Repeating what the image shows rarely helps.
The old lighthouse keeper turns toward the camera as a wave crashes on the rocks behind him and says in a deep gravelly voice: "Storm's coming. Light the lamp." The lighthouse beam sweeps across the dark clouds. Slow push in. Sound: crashing waves, howling wind.
Requesting cuts inside one clip
Vidu Q4 can produce several shots in one render. Name each shot in order with “cut to” and keep the count realistic for the length: three beats fit in 5 seconds, more need 8 to 16 seconds.
The chef tosses vegetables in the pan and a burst of flame rises from it. Cut to a close-up of the sizzling pan, then cut back to the chef, who grins at the camera. Sound: loud sizzle, the whoosh of the flame, busy kitchen clatter.
Reference to video: tag every upload
Uploads are numbered by order. Vidu Q4’s own syntax is [@reference_image_1], [@reference_image_2] … up to 12, and [reference_audio_1] to [reference_audio_3] for voice clips. In this studio you can type the shorter @Image1 and @Audio1; they are converted to Vidu’s tags before the request is sent, and prompts already written in Vidu’s syntax pass through unchanged. Use the tags where the subject appears in the action so each reference gets a clear role.
A single untagged portrait was still respected in our test, but with two or more references tags remove the guesswork about who is who.
[@reference_image_1] sits on a bench on [@reference_image_3] and laces up [@reference_image_2]. She stands, dribbles a basketball and makes a jump shot at golden hour. She turns to the camera and says: "New season, new kicks." Handheld camera, quick cuts. Sound: ball bouncing, sneakers squeaking, city ambience.
Common mistakes
- Asking for more shots than the length allows: five cuts in 3 seconds will blur together.
- Forgetting sound: image to video always has audio, so an unspecified track may not be what you want.
- Describing the reference instead of tagging it: “a woman with red hair” may produce a new woman; “@Image1” uses yours.
- Long dialogue: keep lines to one short sentence per few seconds of video.