Beta notice: Wan 3.0 is currently in internal beta and may be unstable. Generated sound, visible text, and detailed instruction accuracy can vary.
BetaMay Be Unstable

Wan 3.0 Video Generator

Build longer narratives with up to 30 seconds of output, more realistic detail, and cross-modal guidance from creative media or office materials.

Up to 30sMultimodal Reference480p / 720p / 1080pSound GenerationFrom 20 credits

Wan 3.0 Video capabilities

Three generation modes cover direct prompting, controlled frame animation, and broad multimodal reference workflows.

Text-to-Video up to 30 seconds

Turn a prompt into a complete short sequence with selectable resolution, aspect ratio, sound, watermark, and seed controls.

  • Fixed 2–30 second output
  • 480p, 720p, and 1080p tiers
  • Automatic or explicit aspect ratio

Smart duration is intentionally unavailable during the beta so billing stays predictable.

First and last frame animation

Animate one image or guide the transition with an optional second image as the final frame.

  • One or two frame images
  • Optional prompt for motion direction
  • Sound on by default

Useful for product reveals, portrait motion, and directed transitions.

Multimodal all-reference workflow

Combine images, short videos, audio, and one document or webpage to guide content and visual direction.

  • Up to 10 images, 5 videos, and 5 audio files
  • One document or one public webpage
  • Reference video and output seconds are billed together

Each video or audio reference must be 1–15 seconds, with a 15-second total per media type.

Office material to video

Use presentations, spreadsheets, PDFs, text, Markdown, and Apple iWork files as creative context.

  • Documents up to 95 MiB
  • Provider limit of up to 50 pages
  • No local document parsing or content retention

Document page-count validation is handled by the model provider.

Wan 3.0 Video Showcase

A real beta image-to-video result generated through the same Wan 3.0 workflow available above.

Luminous Fashion Gallery

Preserve the same elegant fashion subject and visual identity across the sequence. She steps into a luminous futuristic gallery as the camera makes a slow cinematic orbit; reflective surfaces, flowing fabric, precise facial detail, natural motion, refined commercial lighting, coherent colors, polished premium fashion-film finish.

image-to-video720pauto5.038s

Designed for longer, reference-rich stories

Thirty-second narrative room

Use the expanded duration for multi-beat product stories, cinematic establishing shots, educational sequences, and social videos that need more than a single moment.

Realism, detail, and sound

Wan 3.0 aims for detailed motion and audiovisual output, while this beta clearly surfaces that voices, sound, visible text, and fine instructions may still vary between generations.

Wan 3.0 Video FAQ

What is Wan 3.0 Video?

Wan 3.0 Video is a beta video generation model supporting text, first and last frames, and multimodal references including images, videos, audio, documents, and webpages.

How many credits does Wan 3.0 use?

Wan 3.0 costs 10 credits per second at 480p, 18 at 720p, and 36 at 1080p. Text and image modes bill output seconds; all-reference mode bills output plus reference-video seconds.

Why is Wan 3.0 marked Beta?

The integration is available for testing but may be less stable than production models. Sound quality, text rendering, instruction accuracy, and provider availability can still vary.

Which document formats are supported?

Supported formats are DOC, DOCX, XLS, XLSX, PPT, PPTX, PDF, TXT, Markdown, Keynote, Pages, and Numbers. Use one document or one webpage, not both.

Can Wan 3.0 generate sound?

Yes. Sound is enabled by default and may be disabled before generation. Audio quality and spoken or on-screen text accuracy may still fluctuate during the beta.

Explore the Wan AI model family

Compare beta multimodal creation, flagship production, LoRA-tuned character video, open-source workflows, multi-shot generation, and motion transfer without leaving the Wan family.

Try Wan 3.0 Video Beta

Start with text, a pair of frames, or a multimodal reference pack. Review the estimated credits before every submission.