Wan 3.0

Cinematic video with references and audio

Describe the video to generate, including motion, camera, and audio direction. Required for Text-to-Video; optional when an image or references drive the shot.

Token cost updates automatically by duration and resolution. 1080p is the sharpest and costs more per second.

5

Choose any whole-second duration from 2 to 30 seconds.

Adaptive lets the model choose the frame from your prompt or references. Pick a ratio manually for a guaranteed vertical or landscape result.

Sign in to spend tokens

Prompt is required.

Your generated video will appear here