Multi-Shot Storytelling
Build connected shots from one idea. Describe shot progression, camera changes, subjects, and environments in one prompt to create a more complete short-form sequence.
Create multi-shot AI videos from text, images, or reference clips with native audio, consistent storytelling, and up to 1080p output.







Wan 2.6 is an AI video generation model family for text-to-video, image-to-video, and reference-to-video creation. It supports multi-shot generation, native audio, and 720p or 1080p output.
Text-to-video and image-to-video workflows can generate clips up to 15 seconds, while reference-to-video supports shorter clips up to 10 seconds. Use it to turn an idea into a sequence, animate an image, or guide a new video with a reference clip.
Wan 2.6 is useful for creators, marketing teams, e-commerce brands, and teams that need short-form video with stronger continuity across shots.

Build connected shots from one idea. Describe shot progression, camera changes, subjects, and environments in one prompt to create a more complete short-form sequence.
Generate dialogue, music, and sound effects with the video so visual movement and sound can be planned together instead of assembled afterward.
Use an existing video to guide a new prompt-driven result while changing the action, setting, or story and retaining relevant reference cues.
Create connected stories where the same subject appears across shots. Reference quality, scene complexity, and prompt structure still affect consistency.
| Mode | Start With | Maximum Model Duration | Best For |
|---|---|---|---|
| Text to Video | Text prompt | Up to 15 seconds | Story ideas, concepts, social videos, narrative sequences |
| Image to Video | Image + prompt | Up to 15 seconds | Product images, portraits, artwork, existing visual concepts |
| Reference to Video | Reference video + prompt | Up to 10 seconds | Character guidance, reference-driven scenes, identity or motion continuity |
Choose Text to Video when you have an idea but no starting visual. Describe the subject, scene, action, camera, audio, and shot flow.
Choose Image to Video when you already have the visual you want to build from. Upload it and describe the motion you want to add.
Choose Reference to Video when an existing clip should guide the new result while your prompt changes the setting, action, or story.
| Feature | Wan 2.6 Model Capability | |
|---|---|---|
| Text to Video | Yes | |
| Image to Video | Yes | |
| Reference to Video | Yes | |
| Multi-Shot Generation | Supported | |
| Native Audio | Supported | |
| Resolution | 720p / 1080p | |
| T2V Maximum Duration | Up to 15 seconds | |
| I2V Maximum Duration | Up to 15 seconds | |
| R2V Maximum Duration | Up to 10 seconds | |
| Browser-Based Generation | Platform dependent | |
| Watermark | Platform dependent | |
| Commercial Usage | Terms dependent |
Start with Text to Video, Image to Video, or Reference to Video based on the material you already have. Use text for an idea, an image for a starting visual, or a reference clip for stronger guidance.
Describe the subject, action, environment, camera movement, lighting, and audio. For multi-shot video, explain how the sequence should progress from one shot to the next.
Review the subject, motion, transitions, audio, and continuity. Adjust the prompt, simplify difficult interactions, or improve the reference material before generating again.
A useful prompt can follow this structure: Subject + Action + Environment + Camera + Lighting + Audio or Dialogue + Shot Progression. Add only the details that matter to your scene.
“A black ceramic coffee cup sits beside an open notebook in a quiet studio. Slow camera push-in, soft morning window light, shallow depth of field, subtle room ambience.”
“Create a three-shot launch video for a compact travel camera. Shot one: wide shot of a traveler walking through a train station. Shot two: close-up of the camera in their hand. Shot three: over-the-shoulder view as they photograph the arriving train. Natural station ambience and smooth visual continuity.”
“Two designers stand beside a prototype in a bright studio. The first asks, ‘Ready to test it?’ The second smiles and replies, ‘Let’s see what it can do.’ Use natural timing, medium close-ups, subtle hand movement, and quiet studio ambience.”
“Keep the main character from the reference clip recognizable. Move the scene to a rooftop at sunset. The character walks toward the camera, stops near the edge, looks across the city, and says one short sentence. Keep the pacing calm and cinematic.”

Create short-form concepts for TikTok, Instagram Reels, YouTube Shorts, personal vlogs, music visuals, storytelling clips, and creator campaigns.

Develop campaign hooks, product reveals, social ads, and short branded stories from text, product images, or reference clips, then refine promising ideas before final production.

Turn product photography and campaign ideas into product demos, launch teasers, lifestyle concepts, Shopify visuals, and short social advertisements.

Prototype training scenes, concept demonstrations, internal communications, presentations, and visual simulations before committing to larger production work.
A newer model does not automatically make an existing workflow irrelevant. Choose based on the project you need to complete.
| Consider Wan 2.6 When... | Compare Wan 3.0 When... |
|---|---|
| You already have a tested Wan 2.6 workflow | You are evaluating newer Wan model options |
| Existing prompts already give you useful results | You are willing to retest prompts and settings |
| You need to reproduce an earlier 2.6 project | You are starting a new workflow from scratch |
| Your current provider, speed, and cost fit the project | You need a capability your 2.6 workflow does not provide |
Do not switch versions based only on the version number. Compare actual output, available controls, generation cost, and the needs of your project.
AI video generation is not deterministic, and some scenes may need several attempts.
Scenes with several people, overlapping movement, or close physical interaction give the model more details to manage at once. Simpler blocking can make the intended action easier to communicate.
Reference-driven generation can still vary across frames or shots. Start with clear reference material and avoid unnecessary visual changes when identity retention matters.
Small prompt changes can affect camera movement, scene composition, dialogue, and motion. Change one important instruction at a time when refining a result.
More shots create more opportunities for details to change. Clearly describe what must stay consistent, including wardrobe, environment, subject appearance, and lighting.
A usable video may require more than one generation. Consider the cost of retries, not only the price of a single clip, when comparing generation options.
Choose a generation mode, duration, and resolution based on your project. Review the displayed credit cost and current terms before generating.
$9.9
100 credits · $0.099/credit
A lightweight one-time credit pack for trying the video generator.
$29.9
330 credits · $0.091/credit
A one-time credit pack with a lower per-credit price than Starter.
$49.9
600 credits · $0.083/credit
A larger one-time credit pack with a lower per-credit price.
$99.9
1,250 credits · $0.080/credit
The largest one-time credit pack with the lowest listed per-credit price.
One-time credit packs. Any applicable taxes and credit-validity terms are shown at checkout or in the current service terms.
Wan 2.6 is an AI video generation model family that supports Text-to-Video, Image-to-Video, and Reference-to-Video workflows. It also supports multi-shot generation, native audio, and 720p or 1080p output.
Start with a prompt, an image, or a reference clip. Choose the mode that fits your project, generate a first version, then refine the shot flow, motion, and audio.