What Wan 3.0 is good at
- Up to 30 seconds: Generates long clips with coherent motion and audio in a single pass
- Native audio: Every output includes a synchronized audio track by default
- Text, image, and reference input: Pure text prompts, first-frame animation, or up to 10 reference images, 5 reference videos, and 5 reference audio clips
- Reference tags: Mention connected media directly in the prompt as
@Image1,@Video1, or@Audio1 - Flexible output: 480p, 720p, or 1080p resolution with adaptive, 16:9, 9:16, 1:1, 4:3, or 3:4 aspect ratios
- WAN3-Prime model: The
modeloption also offerswan3.0-video-primefor higher-fidelity output at a higher per-second rate - Bilingual prompts: Prompt in English or Chinese
Example outputs
Text-to-video generation from a prompt alone, with synchronized audio included by default: Image-to-video generation, animating a single image into a full clip: Reference-to-video generation, keeping a product and character consistent from reference images:Use it in ComfyUI
Wan 3.0 workflows
Run the text-to-video, image-to-video, and reference-to-video workflows in ComfyUI, locally or on Comfy Cloud