New — access 450+ AI models through one unified API.Start building free

wan-3.0/text-to-video

Wan 3.0 (Text-to-Video) turns a prompt alone into native video up to 30 seconds with synchronized dialogue, music, and sound effects — no reference media required. Supports multi-scene narratives at up to 1080p with strong instruction-following.

text-to-video$0.2000/ run

Prompt

0/2000
PreviewJSON

Your result will appear here

Write a prompt and hit Generate.

Examples

Close-up portrait of a woman pressing her palm against a rain-soaked window at night, shot from outside. Heavy rain streaks distort her face through the glass, neon signs from the street below cast fragmented magenta and cyan reflections across her skin. Camera slowly pushes in at 0.3x speed. Shallow depth of field, bokeh rain drops, cinematic 4K, anamorphic lens flare, 24fps.

Related models

wan-2.2-spicy / image-to-videoimage-to-video
wan-2.2-spicy / image-to-video

WAN 2.2 Spicy converts images into unlimited high-quality videos with smooth animations optimized for scalable content generation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Wan$0.1500
wan-2.2-spicy / image-to-video-loralora-support
wan-2.2-spicy / image-to-video-lora

Generate AI videos with personalized styles using LoRA. Upload images and apply a trained style model to WAN 2.2 — create unique, stylized videos with consistent visual identity.

Wan$0.2000
wan-3.0 / image-to-videoimage-to-video
wan-3.0 / image-to-video

Wan 3.0 (Image-to-Video) animates a reference image into video up to 30 seconds at up to 1080p, with first and last frame control and native audio-video synthesis. Alibaba's latest Wan generation supports multi-scene narratives with synchronized dialogue, BGM and sound effects generated automatically.

Wan$0.2000
wan-3.0 / reference-to-videoreference-to-video
wan-3.0 / reference-to-video

Wan 3.0 (Reference-to-Video) generates video guided by first/last frame images plus multiple reference images, videos, and audio clips, producing native 30-second scenes with synchronized audio at up to 1080p — Alibaba's most capable Wan model for multimodal creative control.

Wan$0.2000
wan-2.7 / image-to-videoimage-to-video
wan-2.7 / image-to-video

WAN 2.7 converts images into videos (480p/720p) with optional audio, supporting first and last frame control. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Wan$0.0900
wan-2.7 / text-to-videotext-to-video
wan-2.7 / text-to-video

WAN 2.7 Text-to-Video turns plain prompts into coherent, cinematic clips with crisp detail, stable motion, and strong instruction-following—great for ads, explainers, and social posts. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Wan$0.0900

Wan 3.0 Text To Video

Wan 3.0 (Text-to-Video) turns a prompt alone into native video up to 30 seconds with synchronized dialogue, music, and sound effects — no reference media required. Supports multi-scene narratives at up to 1080p with strong instruction-following.

Key Features

  • Generates video from a prompt alone — no reference image required, up to 30 seconds.
  • Native audio-video synthesis — synchronized dialogue, background music and sound effects generated automatically.
  • Multi-scene narrative support — a single generation can carry more than one beat or camera setup.
  • Strong instruction-following — accurately renders detailed, multi-element prompts.
  • Multiple resolution tiers — generate at 480p, 720p, or 1080p.
  • Bilingual prompt support — supports both English and Chinese prompts natively.

Parameters

ParameterRequiredDescription
promptYesDetailed description of the scene, action, and dialogue (English or Chinese).
durationNoVideo length in seconds: 2-30 (default: 5).
resolutionNoOutput resolution: 480p, 720p (default), or 1080p.
aspect_ratioNoOutput format: 16:9 (default), 9:16, 4:3, 3:4, 1:1, or adaptive.

How to Use

  1. Write a prompt — describe subject, action, camera movement, lighting and mood in English or Chinese.
  2. Select your aspect ratio (e.g., 16:9 for widescreen).
  3. Choose a duration up to 30 seconds.
  4. Generate and download your video with synchronized dialogue, music, and sound effects.

Code Examples

import os
import requests

response = requests.post(
    "https://aircube.ai/api/v3/wan-3.0/text-to-video",
    headers={
        "Authorization": "Bearer " + os.environ["AIRCUBE_API_KEY"],
        "Content-Type": "application/json",
    },
    json={
    "prompt": "A chef plates a dessert in a busy restaurant kitchen, steam rising, camera pushes in slowly, ambient kitchen noise",
    "duration": 8,
    "aspect_ratio": "16:9",
    "resolution": "1080p"
},
    timeout=300,
)
data = response.json()

if data["success"]:
    print("ID:", data["data"]["id"], "Status:", data["data"]["status"])
else:
    print("Error:", data["error"]["message"])

Pricing

ResolutionDurationCost
480p4s$0.20
480p5s$0.25
480p6s$0.30
480p8s$0.40
480p10s$0.50
480p12s$0.60
480p15s$0.75
480p20s$1.00
480p25s$1.25
480p30s$1.50
720p4s$0.40
720p5s$0.50
720p6s$0.60
720p8s$0.80
720p10s$1.00
720p12s$1.20
720p15s$1.50
720p20s$2.00
720p25s$2.50
720p30s$3.00
1080p4s$0.80
1080p5s$1.00
1080p6s$1.20
1080p8s$1.60
1080p10s$2.00
1080p12s$2.40
1080p15s$3.00
1080p20s$4.00
1080p25s$5.00
1080p30s$6.00

Billing rules

  • 480p / 4s: $0.20.
  • 720p / 4s: $0.40.
  • 1080p / 4s: $0.80.
  • Longer durations scale proportionally.
  • Failed generations are not charged.

Best Use Cases

  • Short-form video content — create engaging clips for social media platforms with native audio.
  • Storyboarding — visualize written concepts as animated multi-scene sequences.
  • Promotional content — generate video ads with spoken dialogue directly from a brief.
  • Multi-scene narrative — script scenes with more than one camera setup in a single 30-second generation.
  • Chinese-language projects — native Chinese prompt support.

Pro Tips

  • Include specific camera directions: tracking shot, dolly zoom, static wide angle.
  • Describe lighting explicitly: golden hour, dramatic shadows, soft diffused light.
  • Write dialogue in quotes for spoken lines — Wan 3.0 generates synchronized audio automatically.
  • Use the 30-second ceiling for scripts with multiple scenes or beats.
  • Wan 3.0 supports both English and Chinese prompts natively.

Notes

  • Bilingual prompt support: English and Chinese.
  • Output in MP4 format with native synchronized audio.
  • Duration range: 2-30 seconds.
  • Successor to Wan 2.7 with a longer duration ceiling and native multi-scene narratives.

Wan 3.0 Text To Video — Frequently asked questions