How we tested
All renders went through fal’s public endpoints fal-ai/vidu/q4/image-to-video and fal-ai/vidu/q4/reference-to-video, the same routes this studio uses. Start and reference images were generated by us for the test, so every frame shown on this site is ours to publish. We checked each output file with ffprobe (resolution, frame rate, audio stream, loudness) and transcribed the audio with Whisper to check spoken lines.
- Two quick 3-second 540p probes: one image to video, one reference to video with a single portrait.
- Three 5-second 720p image-to-video clips: a talking lighthouse keeper, a chef with requested cuts, a red panda with forest ambience.
- One 5-second 720p, 9:16 reference-to-video ad from three references: a person, a sneaker and a rooftop court.
- One reference-to-video request with no references, to see whether text-only generation is possible.
Dialogue and native sound
Every Vidu Q4 image-to-video file we received had an AAC audio track, even though the route has no audio switch. For the lighthouse keeper we asked for the line “Storm’s coming. Light the lamp.” Whisper transcribed the generated audio as “Storms coming. Light the lamp.”, with waves and wind underneath. The panda clip, which asked only for ambience, came back as quiet forest sound (mean loudness −37.8 dB) rather than music or speech.
The old lighthouse keeper turns toward the camera as a wave crashes on the rocks behind him and says in a deep gravelly voice: "Storm's coming. Light the lamp." The lighthouse beam sweeps across the dark clouds. Slow push in. Sound: crashing waves, howling wind.
Camera cuts and multi-shot prompts
ShengShu says Vidu Q4 keeps camera moves and cuts coherent with the action. In our chef test we asked for three beats in five seconds. The clip shows the flame burst, a close-up of the vegetables sizzling in the pan, and a cut back to the chef grinning. One honest flaw: the chef’s apron changed from black in our start image to brown in the opening shot, and back to black after the cut, so check wardrobe details when continuity matters.
The chef tosses vegetables in the pan and a burst of flame rises from it. Cut to a close-up of the sizzling pan, then cut back to the chef, who grins at the camera. Sound: loud sizzle, the whoosh of the flame, busy kitchen clatter.
Reference consistency
Reference to video accepts up to 12 images and 3 MP3 voice clips on fal (ShengShu announces up to 15 images for Vidu Q4 itself). In a first probe we gave a single portrait and did not tag it in the prompt; the woman in the result still matched the reference. For the sneaker ad we tagged three uploads by position, and the person, the white sneaker with orange laces and the rooftop court all carried through five shots, and Whisper transcribed her line as “New season, new kicks.”
[@reference_image_1] sits on a bench on [@reference_image_3] and laces up [@reference_image_2]. She stands, dribbles a basketball and makes a jump shot at golden hour. She turns to the camera and says: "New season, new kicks." Handheld camera, quick cuts. Sound: ball bouncing, sneakers squeaking, city ambience.
Limits we hit
- No text-only route: reference to video without a reference image or voice clip is rejected with “At least one reference image or reference audio is required”.
- Image to video follows the start image’s shape and takes no aspect-ratio setting; reference to video offers 16:9, 9:16, 1:1, 4:3 and 3:4.
- There is no end-frame input and no video-to-video input on either route.
- Renders are not instant: 2 to 8 minutes for 3- to 5-second clips in our tests, and reference clips took longest.
What it cost
fal charges Vidu Q4 per output second, the same on both routes, with sound included. Until November 30, 2026 fal runs a 30% launch discount, so a 5-second 720p clip costs $0.3325 instead of $0.475. All the Vidu Q4 renders in this review together cost about $1.50 in fal usage at launch rates. On ViduQ4.org the same 5-second 720p clip is 15 credits; see the pricing guide for every resolution.
Verdict
For a preview, Vidu Q4 is strongest where it says it is: performances with spoken lines, and short multi-shot clips that follow the action. Reference to video is the more useful route for ads and recurring characters. The gaps are the missing text-only mode, the occasional wardrobe drift between shots, and render times measured in minutes. We will update this review as the model leaves preview.