New Model

🎉 Wan 3.0 is live — Alibaba's newest video model. Native 30-second clips at 1080P with synced audio, guided by up to 20 image, video, and audio references!

Wan 3.0 Video Generator

Generate videos using cutting-edge AI models including Veo 3, Sora 2, and more - with optional reference images

Wan 3.0Wan 3.0
Upload Image
+
Optional
+
Optional

Note: Images will be used as reference materials to guide the video generation.

Upload up to 5 reference audio clips (total ≤15s) to guide audio generation.

16:9
9:16
1:1
4:3
3:4
Adaptive
480P
720P
1080P
4s2s - 30s
37/5000

✨ Please login to get free credits ✨

Wan 3.0 — 30-Second AI Video with Native Audio and Omni-Reference

Tell longer stories with Wan 3.0, Alibaba's latest all-in-one video model. Generate up to 30 seconds in a single pass at 1080P, guide every shot with up to 20 image, video, and audio references, and get dialogue, music, and sound effects synced from the first frame.
Steps

How to Use Wan 3.0 on LuminaMind

Turn a prompt and a few references into a 30-second, audio-synced clip with Wan 3.0.

Upload images, video clips, or audio tracks to anchor characters, products, locations, and voices. Wan 3.0 keeps them consistent across every shot of the clip.

Add reference assets for Wan 3.0
Enter a prompt for Wan 3.0 generation
Select the Wan 3.0 model
Download your Wan 3.0 video

Why Choose Wan 3.0?

Alibaba's most capable video model yet—built for longer stories, richer references, and audio that arrives with the picture.

Native 30-Second Generation

Wan 3.0 doubles the 15-second ceiling of Wan 2.7. Generate a complete 30-second scene with multi-shot narrative and director-level camera moves in one pass—no stitching.

Omni-Reference Inputs

Combine up to 10 images, 5 video clips, and 5 audio tracks in a single generation to lock in characters, props, locations, and voices. Documents and web pages can be references too.

Pixel-Perfect Consistency

Faces, wardrobe, and products stay recognizable from the first shot to the last, with fewer deformations and clipping artifacts than earlier Wan releases.

Native Synced Audio

Dialogue, background music, and sound effects are generated together with the video, so lips, beats, and impacts land on the right frame without post-production dubbing.

Reality-Grade Rendering

Sharper hair, glass reflections, water, and lighting, plus more expressive faces and micro-expressions for performances that read as real.

Precision Editing & Text

Edit generated clips by instruction or reference, extend them forward or backward, and render legible on-screen text, charts, and formulas in 12 languages.

Stats

Wan 3.0 at a Glance

Longer, more controllable, audio-native

Max Length

30s

Single pass

Resolution

1080P

30 fps

Native Audio

Synced

Dialogue, music, SFX

Use Cases

Where Wan 3.0 Shines

Official Wan 3.0 showcases—single-take rides, surreal commercials, comedy, cinematic realism, music, and sci-fi action.

30-Second Single-Take Ride

A kid rides a shopping cart down a palm-lined Los Angeles boulevard, weaving through traffic and launching skyward, all in one unbroken 30-second shot with the camera locked on the action.

Fantasy Product Commercial

A white-haired girl opens the fridge and shrinks into a frozen world of juice cartons, jelly monsters, and broccoli forests, then returns with dessert. One continuous clip, five surreal set pieces.

Deadpan Desert Comedy

A Route 66 mechanic lounges outside a sun-bleached gas station, toothpick in mouth, as a cowgirl rides up on horseback. Dry timing, natural light, and expressive close-ups carry the joke.

Cinematic Realism

A woman in a crimson gown stands in a misty river as a herd of black horses charges past her. Splashing water, flowing fabric, and muscle motion all hold up under a slow cinematic push-in.

Music & Performance

A violinist plays through pouring rain under cold teal light, water streaming off the instrument. Bow strokes stay in sync with the generated score for a moody, music-video-ready single take.

Sci-Fi Action Sequence

Combat robots battle through a burning city street while a single white flower survives in the rubble. Explosions, debris, and metal impacts land with synced sound across the full 30 seconds.

Pricing

Credits can be used for video generation with multiple AI models including Wan 3.0 and more.

Save 30%

700 Credits

Popular
$59.9$35/ month

Most popular for individual creators!

Includes

  • 700 credits / month
  • Credits never expire
  • 4K Video Resolution
  • Text/Image/Video to Video:
    Veo 3.1Veo 3.1
    Sora 2Sora 2
    Seedance 2.5Seedance 2.5
    Wan 3.0Wan 3.0
  • Text/Image to Image:
    GPT Image 2GPT Image 2
    Nano Banana 2Nano Banana 2
  • No Watermark
  • Private Generation
  • Commercial License

cancel anytime

400 Credits

$39.9$21/ month

Perfect for trying out.

Includes

  • 400 credits / month
  • Credits never expire
  • 4K Video Resolution
  • Text/Image/Video to Video:
    Veo 3.1Veo 3.1
    Sora 2Sora 2
    Seedance 2.5Seedance 2.5
    Wan 3.0Wan 3.0
  • Text/Image to Image:
    GPT Image 2GPT Image 2
    Nano Banana 2Nano Banana 2
  • No Watermark
  • Private Generation
  • Commercial License

cancel anytime

1500 Credits

Most Cost-Effective
$119.9$70/ month

Best for professional creators!

Includes

  • 1500 credits / month
  • Credits never expire
  • 4K Video Resolution
  • Text/Image/Video to Video:
    Veo 3.1Veo 3.1
    Sora 2Sora 2
    Seedance 2.5Seedance 2.5
    Wan 3.0Wan 3.0
  • Text/Image to Image:
    GPT Image 2GPT Image 2
    Nano Banana 2Nano Banana 2
  • No Watermark
  • Private Generation
  • Commercial License
  • Priority Support

cancel anytime

Secure, encrypted checkoutPowered byStripe

Made with Wan 3.0

See what creators are making with Wan 3.0

The Technology Behind Wan 3.0

How Alibaba's unified multimodal transformer reaches 30 seconds.

One DiT for Every Task

Text-to-video, keyframes, references, editing, and extension run through one flow-matching diffusion transformer. A router reads media types and prompt intent, and the whole 30-second latent is denoised as one sequence.

Joint Audio-Video Latents

Frames go through the 3D causal Wan-VAE and sound through a 1D audio VAE; both token streams are denoised together in one hybrid-MMDiT sequence, so dialogue, music, and effects emerge with the picture, not as a later dub.

Reference Token Binding

Up to 20 images, clips, and audio tracks enter as numbered reference tokens (Image 1, Video 1, Audio 1) that every frame attends to, while a thinking stage rewrites the prompt or a document into a timestamped shot script.

FAQ

Wan 3.0 FAQ

Common questions about Wan 3.0

1

What is Wan 3.0?

Wan 3.0 is Alibaba's latest all-in-one AI video model from Tongyi Lab, released in public beta on August 6, 2026 and generally available since August 24, 2026. It generates video and native audio together from text, image, video, audio, and even document inputs, in clips up to 30 seconds.

2

How is Wan 3.0 different from Wan 2.5?

Wan 3.0 raises the single-pass limit from 10 seconds in Wan 2.5 to 30 seconds, merges reference-to-video, keyframes, editing, and extension into one model, accepts up to 20 multimodal references, and delivers steadier characters, richer physical detail, and cleaner audio-visual sync.

3

What resolution and length does Wan 3.0 support?

Wan 3.0 outputs 480P, 720P, or 1080P at 30 fps in 16:9, 9:16, 1:1, 4:3, 3:4, or adaptive aspect ratios, with any length from 2 to 30 seconds in a single generation. When you add a reference video, the input and output durations together can add up to 30 seconds in total.

4

Does Wan 3.0 generate audio and lip-sync?

Yes. Wan 3.0 produces dialogue, background music, and sound effects in the same pass as the picture, so speech stays lip-synced and effects land on the right frame. Put spoken lines in quotes in your prompt, or supply up to five reference audio clips to drive voices and rhythm.

5

What inputs can I use with Wan 3.0?

Text prompts, first and last frames, reference images, reference video clips, and reference audio can all guide one generation. Wan 3.0 takes up to 20 references in total—10 images, 5 videos, and 5 audio clips—and the model can also read documents, spreadsheets, slides, and web pages.

6

Is Wan 3.0 suitable for commercial use?

Yes. Native 30-second clips, 1080P output, precise multi-reference control, and synced audio make Wan 3.0 a strong fit for advertising, product videos, short drama, tourism marketing, and music videos, from quick concept tests to polished, brand-ready final deliverables.

Start Creating with Wan 3.0

Generate 30-second, audio-synced AI videos with Wan 3.0 today.