New — access 450+ AI models through one unified API.Start building free

wan-3.0/reference-to-video

Wan 3.0 (Reference-to-Video) generates video guided by first/last frame images plus multiple reference images, videos, and audio clips, producing native 30-second scenes with synchronized audio at up to 1080p — Alibaba's most capable Wan model for multimodal creative control.

reference-to-video$0.2000/ run

Prompt

0/2000
Reference MediaImages, videos, and audio supported
PreviewJSON

Your result will appear here

Write a prompt and hit Generate.

Related models

wan-2.2-spicy / image-to-videoimage-to-video
wan-2.2-spicy / image-to-video

WAN 2.2 Spicy converts images into unlimited high-quality videos with smooth animations optimized for scalable content generation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Wan$0.1500
wan-2.2-spicy / image-to-video-loralora-support
wan-2.2-spicy / image-to-video-lora

Generate AI videos with personalized styles using LoRA. Upload images and apply a trained style model to WAN 2.2 — create unique, stylized videos with consistent visual identity.

Wan$0.2000
wan-3.0 / image-to-videoimage-to-video
wan-3.0 / image-to-video

Wan 3.0 (Image-to-Video) animates a reference image into video up to 30 seconds at up to 1080p, with first and last frame control and native audio-video synthesis. Alibaba's latest Wan generation supports multi-scene narratives with synchronized dialogue, BGM and sound effects generated automatically.

Wan$0.2000
wan-3.0 / text-to-videotext-to-video
wan-3.0 / text-to-video

Wan 3.0 (Text-to-Video) turns a prompt alone into native video up to 30 seconds with synchronized dialogue, music, and sound effects — no reference media required. Supports multi-scene narratives at up to 1080p with strong instruction-following.

Wan$0.2000
wan-2.7 / image-to-videoimage-to-video
wan-2.7 / image-to-video

WAN 2.7 converts images into videos (480p/720p) with optional audio, supporting first and last frame control. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Wan$0.0900
wan-2.7 / text-to-videotext-to-video
wan-2.7 / text-to-video

WAN 2.7 Text-to-Video turns plain prompts into coherent, cinematic clips with crisp detail, stable motion, and strong instruction-following—great for ads, explainers, and social posts. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Wan$0.0900

Wan 3.0 Reference To Video

Wan 3.0 (Reference-to-Video) generates video guided by first/last frame images plus multiple reference images, videos, and audio clips, producing native 30-second scenes with synchronized audio at up to 1080p — Alibaba's most capable Wan model for multimodal creative control.

Key Features

  • Multimodal reference inputs — combine first/last frame images with additional reference images, videos, and audio in a single request.
  • Motion and camera reference — extract choreography and camera movement from a reference video and apply it to a new scene.
  • Style and character consistency — reference images preserve appearance and style across the generated clip.
  • Native audio-video synthesis — synchronized dialogue, music and sound effects generated automatically alongside the picture.
  • Native 30-second clips — orchestrate longer, multi-scene sequences from a single reference set.
  • Bilingual prompt support — supports both English and Chinese prompts natively.

Parameters

ParameterRequiredDescription
promptYesScene description guiding how the reference assets should be used (English or Chinese).
imageNoFirst-frame reference image URL.
last_imageNoLast-frame reference image URL for precise start/end-frame control.
reference_imagesNoAdditional reference image URLs for style or character guidance (up to 10).
reference_videosNoReference video URLs for motion and camera guidance (up to 5, 15s combined).
reference_audiosNoReference audio URLs for soundtrack or voice guidance (up to 5, 15s combined).
durationNoVideo length in seconds: 2-30 (default: 5).
resolutionNoOutput resolution: 480p, 720p (default), or 1080p.
aspect_ratioNoOutput format: 16:9 (default), 9:16, 4:3, 3:4, 1:1, or adaptive.

How to Use

  1. Upload reference assets — first/last frame images anchor the start and end, additional images define style/character, videos define motion/camera, audio drives soundtrack.
  2. Write a prompt describing how the references should come together into a scene.
  3. Choose duration (up to 30s), resolution, and aspect ratio.
  4. Generate — submit and wait, then download the finished video with synchronized audio.
  5. Iterate — start with 480p/5s for quick iterations, then render the final version at 1080p.

Code Examples

import os
import requests

response = requests.post(
    "https://aircube.ai/api/v3/wan-3.0/reference-to-video",
    headers={
        "Authorization": "Bearer " + os.environ["AIRCUBE_API_KEY"],
        "Content-Type": "application/json",
    },
    json={
    "prompt": "A dancer performs the same choreography as the reference video, but on a neon-lit rooftop at night",
    "image": "https://example.com/dancer-first-frame.jpg",
    "reference_videos": [
        "https://example.com/choreography-reference.mp4"
    ],
    "duration": 8,
    "resolution": "1080p",
    "aspect_ratio": "16:9"
},
    timeout=300,
)
data = response.json()

if data["success"]:
    print("ID:", data["data"]["id"], "Status:", data["data"]["status"])
else:
    print("Error:", data["error"]["message"])

Pricing

ResolutionDurationCost
480p4s$0.20
480p5s$0.25
480p6s$0.30
480p8s$0.40
480p10s$0.50
480p12s$0.60
480p15s$0.75
480p20s$1.00
480p25s$1.25
480p30s$1.50
720p4s$0.40
720p5s$0.50
720p6s$0.60
720p8s$0.80
720p10s$1.00
720p12s$1.20
720p15s$1.50
720p20s$2.00
720p25s$2.50
720p30s$3.00
1080p4s$0.80
1080p5s$1.00
1080p6s$1.20
1080p8s$1.60
1080p10s$2.00
1080p12s$2.40
1080p15s$3.00
1080p20s$4.00
1080p25s$5.00
1080p30s$6.00

Billing rules

  • 480p / 4s: $0.20.
  • 720p / 4s: $0.40.
  • 1080p / 4s: $0.80.
  • Longer durations scale proportionally.
  • First 5 reference images are free; each additional image adds $0.04.
  • Reference video input is billed per second at $0.05/s (480p), $0.10/s (720p), $0.20/s (1080p), added on top of the output price.
  • Failed generations are not charged.

Best Use Cases

  • Motion transfer — extract choreography or camera movement from a reference video and apply it to new subjects.
  • Style transfer — apply a reference image's look to the generated scene.
  • Multi-scene narrative — combine multiple references to maintain character and camera continuity across a 30-second cut.
  • Precise start/end control — pin the exact first and last frame while the model fills in the motion between them.
  • Soundtrack-guided generation — use a reference audio clip to drive the pacing and mood of the scene.

Pro Tips

  • Combine a first-frame and last-frame image whenever the clip needs to land on a specific pose or composition.
  • Keep reference videos short and focused — Wan 3.0 reads up to 15 seconds combined across all reference videos.
  • Describe each reference's role explicitly in the prompt to avoid ambiguity.
  • Use one primary camera instruction per prompt for the cleanest motion transfer.
  • Start with 480p/5s for rapid iteration, then render the final version at 1080p.

Notes

  • Up to 10 reference images, 5 reference videos (15s combined), and 5 reference audio clips (15s combined) per request.
  • Native audio generation is enabled by default.
  • Duration range: 2-30 seconds (continuous selection).
  • Successor to Wan 2.7 Reference-to-Video with a longer duration ceiling and first/last-frame control.

Wan 3.0 Reference To Video — Frequently asked questions