New — access 450+ AI models through one unified API.Start building free

wan-3.0/image-to-video

Wan 3.0 (Image-to-Video) animates a reference image into video up to 30 seconds at up to 1080p, with first and last frame control and native audio-video synthesis. Alibaba's latest Wan generation supports multi-scene narratives with synchronized dialogue, BGM and sound effects generated automatically.

image-to-video$0.2000/ run

Prompt

0/2000

Image

PreviewJSON

Your result will appear here

Upload an image, write a prompt and hit Generate.

Examples

A father and young daughter flying a green diamond kite together in a sunlit park. Both arms raised holding the kite string, both gazing upward with joy. Camera slowly circles them, warm golden afternoon light, soft lens flare, cinematic handheld feel.

Related models

wan-2.2-spicy / image-to-videoimage-to-video
wan-2.2-spicy / image-to-video

WAN 2.2 Spicy converts images into unlimited high-quality videos with smooth animations optimized for scalable content generation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Wan$0.1500
wan-2.2-spicy / image-to-video-loralora-support
wan-2.2-spicy / image-to-video-lora

Generate AI videos with personalized styles using LoRA. Upload images and apply a trained style model to WAN 2.2 — create unique, stylized videos with consistent visual identity.

Wan$0.2000
wan-3.0 / text-to-videotext-to-video
wan-3.0 / text-to-video

Wan 3.0 (Text-to-Video) turns a prompt alone into native video up to 30 seconds with synchronized dialogue, music, and sound effects — no reference media required. Supports multi-scene narratives at up to 1080p with strong instruction-following.

Wan$0.2000
wan-3.0 / reference-to-videoreference-to-video
wan-3.0 / reference-to-video

Wan 3.0 (Reference-to-Video) generates video guided by first/last frame images plus multiple reference images, videos, and audio clips, producing native 30-second scenes with synchronized audio at up to 1080p — Alibaba's most capable Wan model for multimodal creative control.

Wan$0.2000
wan-2.7 / image-to-videoimage-to-video
wan-2.7 / image-to-video

WAN 2.7 converts images into videos (480p/720p) with optional audio, supporting first and last frame control. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Wan$0.0900
wan-2.7 / text-to-videotext-to-video
wan-2.7 / text-to-video

WAN 2.7 Text-to-Video turns plain prompts into coherent, cinematic clips with crisp detail, stable motion, and strong instruction-following—great for ads, explainers, and social posts. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Wan$0.0900

Wan 3.0 Image To Video

Wan 3.0 (Image-to-Video) animates a reference image into video up to 30 seconds at up to 1080p, with first and last frame control and native audio-video synthesis. Alibaba's latest Wan generation supports multi-scene narratives with synchronized dialogue, BGM and sound effects generated automatically.

Key Features

  • Alibaba's latest Wan generation — native video up to 30 seconds, up from 15s on Wan 2.7.
  • First and last frame control — pin the start and end of the clip with two reference images for precise pixel-level continuity.
  • Native audio-video synthesis — synchronized dialogue, background music and sound effects generated automatically alongside the picture.
  • Multi-scene narrative support — a single generation can carry more than one beat or camera setup.
  • Multiple resolution tiers — generate at 480p, 720p, or 1080p.
  • Bilingual prompt support — supports both English and Chinese prompts natively.

Parameters

ParameterRequiredDescription
promptYesDescription of the motion and scene dynamics (English or Chinese).
imageYesFirst-frame image URL to animate.
last_imageNoLast-frame image URL for precise start/end-frame control.
durationNoVideo length in seconds: 2-30 (default: 5).
resolutionNoOutput resolution: 480p, 720p (default), or 1080p.
aspect_ratioNoOutput format: 16:9 (default), 9:16, 4:3, 3:4, 1:1, or adaptive.

How to Use

  1. Upload your first-frame image (clear, well-lit frame).
  2. Optionally add a last-frame image to control exactly how the clip ends.
  3. Write a prompt describing the desired motion — supports English and Chinese.
  4. Select duration (up to 30s) and resolution.
  5. Generate and download your animated video with synchronized audio.

Code Examples

import os
import requests

response = requests.post(
    "https://aircube.ai/api/v3/wan-3.0/image-to-video",
    headers={
        "Authorization": "Bearer " + os.environ["AIRCUBE_API_KEY"],
        "Content-Type": "application/json",
    },
    json={
    "prompt": "Gentle waves lapping at the shore, clouds drifting across the sky",
    "image": "https://example.com/beach-sunset.jpg",
    "duration": 5,
    "resolution": "720p"
},
    timeout=300,
)
data = response.json()

if data["success"]:
    print("ID:", data["data"]["id"], "Status:", data["data"]["status"])
else:
    print("Error:", data["error"]["message"])

Pricing

ResolutionDurationCost
480p4s$0.20
480p5s$0.25
480p6s$0.30
480p8s$0.40
480p10s$0.50
480p12s$0.60
480p15s$0.75
480p20s$1.00
480p25s$1.25
480p30s$1.50
720p4s$0.40
720p5s$0.50
720p6s$0.60
720p8s$0.80
720p10s$1.00
720p12s$1.20
720p15s$1.50
720p20s$2.00
720p25s$2.50
720p30s$3.00
1080p4s$0.80
1080p5s$1.00
1080p6s$1.20
1080p8s$1.60
1080p10s$2.00
1080p12s$2.40
1080p15s$3.00
1080p20s$4.00
1080p25s$5.00
1080p30s$6.00

Billing rules

  • 480p / 4s: $0.20.
  • 720p / 4s: $0.40.
  • 1080p / 4s: $0.80.
  • Longer durations scale proportionally.
  • Failed generations are not charged.

Best Use Cases

  • Product demos — animate product shots with precise first/last-frame control.
  • Social media content — animate photos into longer, narrative-driven posts.
  • Portrait animation — bring still portraits to life with synchronized audio.
  • Chinese-language projects — native Chinese prompt support.
  • Multi-beat storytelling — use the extended 30-second ceiling for scenes with more than one moment.

Pro Tips

  • Wan 3.0 supports both English and Chinese prompts natively.
  • Use clear, well-lit source images for best motion quality.
  • Add a last-frame image whenever the clip needs to land on a specific pose or composition.
  • Describe physics-based motion for natural results.
  • Combine with premium models like Seedance for hero content, Wan for volume.

Notes

  • Bilingual prompt support: English and Chinese.
  • Output in MP4 format.
  • Duration range: 2-30 seconds.
  • Successor to Wan 2.7 with a longer duration ceiling and first/last-frame control.

Wan 3.0 Image To Video — Frequently asked questions