Performances, not just motion
Vidu Q4 was built to coordinate facial expression, emotion, body movement and voice. Write the line in quotes and say how it is delivered; the character speaks it with matching lip movement and sound.
Start Creating NowOnly live routes are selectable. Your prompt stays in place when you switch.
Your video will appear here
Vidu Q4 is the newest video generation model from ShengShu Technology, released as a public preview on October 7, 2026. Vidu Q4 focuses on performance: facial expression, emotion, body movement and voice are generated together, and camera moves and cuts follow fast action. On ViduQ4.org you can run the two public Vidu Q4 routes: image to video, which animates your start frame with dialogue and native sound, and reference to video, which keeps up to 12 characters, products or places and up to 3 voices consistent in one scene. Clips run 3 to 16 seconds at 540p, 720p, 1080p, 2K or 4K, the exact credit cost is shown before every Vidu Q4 render, and every result stays in your history.
Three ways Vidu Q4 gives you more control over the result you walk away with.
Vidu Q4 was built to coordinate facial expression, emotion, body movement and voice. Write the line in quotes and say how it is delivered; the character speaks it with matching lip movement and sound.
Start Creating NowReference to video keeps up to 12 uploaded characters, outfits, products or locations consistent, plus up to 3 MP3 voice clips. Cite them as @Image1, @Image2 … to give each one a role.
Ask for a cut, a close-up or a tracking shot inside one prompt. In our tests Vidu Q4 cut from a flaming pan to a close-up and back to the chef in a single 5-second clip.
Vidu Q4 has two routes on this site. Use image to video when you already have the opening frame, and reference to video when the scene must keep the same people, products, places or voices from your uploads.
Keeps your image as the first frame and its aspect ratio. Sound is always on for this route.
Builds a new scene in 16:9, 9:16, 1:1, 4:3 or 3:4. Sound can be turned off; it is on whenever a voice clip is attached.
2 credits per second at 540p up to 10 at 4K, the same on both routes.
The public Vidu Q4 routes need a start image or at least one reference. There is no end-frame or video-to-video input.
One image is animated as the start frame (image to video). Two to 12 images, or any MP3 voice clip, switch to reference to video, which builds a new scene around them.
Describe what happens in order, then the camera, then the sound. Put dialogue in quotes; with references, cite uploads as @Image1, @Image2 and so on.
Any length from 3 to 16 seconds and 540p, 720p, 1080p, 2K or 4K. The credit quote updates before you submit.
Watch the clip with sound, download the MP4 and change one thing at a time for the next take.
Working prompt structures with the result they produced. Swap the subject, keep the direction concrete, and adapt them to your own brief.
Upload the person, the product and the location as references, cite each one and describe the beat. Vidu Q4 kept the same face, jacket and sneaker across every cut of this vertical ad.
[@reference_image_1] sits on a bench on [@reference_image_3] and laces up [@reference_image_2]. She stands, dribbles a basketball and makes a jump shot at golden hour. She turns to the camera and says: "New season, new kicks." Handheld camera, quick cuts. Sound: ball bouncing, sneakers squeaking, city ambience.
Start from a portrait or film still, write one short line in quotes and add the sound around it. Vidu Q4 animates the face, voices the line and keeps the original frame.
The old lighthouse keeper turns toward the camera as a wave crashes on the rocks behind him and says in a deep gravelly voice: "Storm's coming. Light the lamp." The lighthouse beam sweeps across the dark clouds. Slow push in. Sound: crashing waves, howling wind.
See the credit cost before every generation, then choose a subscription or a one-time top-up when you need more.
For regular individual creative work.
For campaigns and higher-volume production.
For larger workloads, with adjustable credit capacity.
Plans are shown for launch preparation. Checkout opens after production checks are complete.
Vidu Q4 is the latest video generation model from ShengShu Technology, the company behind Vidu. It launched as a public preview on October 7, 2026, with a focus on lifelike performances: expression, emotion, movement and voice generated together, plus camera work that follows action scenes.
Vidu Q4 is a paid model; every render costs provider compute. On ViduQ4.org credits start at 6 for a 3-second 540p clip, and you see the exact cost before you submit. Plans and one-time credit packs are on the pricing page.
Not through the public routes. Vidu Q4 image to video needs a start image, and reference to video needs at least one reference image or voice clip; a text-only request is rejected by the model API. Upload any image you have rights to and describe the shot.
Yes. Image to video always returns a soundtrack, and reference to video adds dialogue and sound effects unless you switch sound off. Write spoken lines in quotes and describe the sounds you want.
ShengShu announced up to 15 image references for Vidu Q4. The route this studio uses accepts up to 12 images and 3 MP3 voice clips of 3 to 12 seconds each, so those are the limits here.
Any whole number of seconds from 3 to 16, at 540p, 720p, 1080p, 2K or 4K. Higher resolutions cost more credits per second.
In our October 8, 2026 tests a 5-second 720p image-to-video clip took about 2.5 minutes and a 5-second 720p clip with three references about 8 minutes. You can leave the page; finished videos are saved to your history.
Credits are reserved when a request starts. Confirmed failed jobs return them automatically; a job still being checked stays pending until it is reconciled.
No. ViduQ4.org is an independent studio that runs the public Vidu Q4 model through fal. It is not affiliated with or endorsed by ShengShu Technology or Vidu. The official product is at vidu.com.
Use Vidu Q4 online: animate a photo or keep up to 12 references consistent, with dialogue and native sound, 3–16 seconds, 540p to 4K. Exact credits shown first.
Start Creating Now