Wan 3.0 Reference To Video
Wan 3.0 (Reference-to-Video) generates video guided by first/last frame images plus multiple reference images, videos, and audio clips, producing native 30-second scenes with synchronized audio at up to 1080p — Alibaba's most capable Wan model for multimodal creative control.
Key Features
- Multimodal reference inputs — combine first/last frame images with additional reference images, videos, and audio in a single request.
- Motion and camera reference — extract choreography and camera movement from a reference video and apply it to a new scene.
- Style and character consistency — reference images preserve appearance and style across the generated clip.
- Native audio-video synthesis — synchronized dialogue, music and sound effects generated automatically alongside the picture.
- Native 30-second clips — orchestrate longer, multi-scene sequences from a single reference set.
- Bilingual prompt support — supports both English and Chinese prompts natively.
Parameters
| Parameter | Required | Description |
|---|---|---|
| prompt | Yes | Scene description guiding how the reference assets should be used (English or Chinese). |
| image | No | First-frame reference image URL. |
| last_image | No | Last-frame reference image URL for precise start/end-frame control. |
| reference_images | No | Additional reference image URLs for style or character guidance (up to 10). |
| reference_videos | No | Reference video URLs for motion and camera guidance (up to 5, 15s combined). |
| reference_audios | No | Reference audio URLs for soundtrack or voice guidance (up to 5, 15s combined). |
| duration | No | Video length in seconds: 2-30 (default: 5). |
| resolution | No | Output resolution: 480p, 720p (default), or 1080p. |
| aspect_ratio | No | Output format: 16:9 (default), 9:16, 4:3, 3:4, 1:1, or adaptive. |
How to Use
- Upload reference assets — first/last frame images anchor the start and end, additional images define style/character, videos define motion/camera, audio drives soundtrack.
- Write a prompt describing how the references should come together into a scene.
- Choose duration (up to 30s), resolution, and aspect ratio.
- Generate — submit and wait, then download the finished video with synchronized audio.
- Iterate — start with 480p/5s for quick iterations, then render the final version at 1080p.
Code Examples
import os
import requests
response = requests.post(
"https://aircube.ai/api/v3/wan-3.0/reference-to-video",
headers={
"Authorization": "Bearer " + os.environ["AIRCUBE_API_KEY"],
"Content-Type": "application/json",
},
json={
"prompt": "A dancer performs the same choreography as the reference video, but on a neon-lit rooftop at night",
"image": "https://example.com/dancer-first-frame.jpg",
"reference_videos": [
"https://example.com/choreography-reference.mp4"
],
"duration": 8,
"resolution": "1080p",
"aspect_ratio": "16:9"
},
timeout=300,
)
data = response.json()
if data["success"]:
print("ID:", data["data"]["id"], "Status:", data["data"]["status"])
else:
print("Error:", data["error"]["message"])Pricing
| Resolution | Duration | Cost |
|---|---|---|
| 480p | 4s | $0.20 |
| 480p | 5s | $0.25 |
| 480p | 6s | $0.30 |
| 480p | 8s | $0.40 |
| 480p | 10s | $0.50 |
| 480p | 12s | $0.60 |
| 480p | 15s | $0.75 |
| 480p | 20s | $1.00 |
| 480p | 25s | $1.25 |
| 480p | 30s | $1.50 |
| 720p | 4s | $0.40 |
| 720p | 5s | $0.50 |
| 720p | 6s | $0.60 |
| 720p | 8s | $0.80 |
| 720p | 10s | $1.00 |
| 720p | 12s | $1.20 |
| 720p | 15s | $1.50 |
| 720p | 20s | $2.00 |
| 720p | 25s | $2.50 |
| 720p | 30s | $3.00 |
| 1080p | 4s | $0.80 |
| 1080p | 5s | $1.00 |
| 1080p | 6s | $1.20 |
| 1080p | 8s | $1.60 |
| 1080p | 10s | $2.00 |
| 1080p | 12s | $2.40 |
| 1080p | 15s | $3.00 |
| 1080p | 20s | $4.00 |
| 1080p | 25s | $5.00 |
| 1080p | 30s | $6.00 |
Billing rules
- 480p / 4s: $0.20.
- 720p / 4s: $0.40.
- 1080p / 4s: $0.80.
- Longer durations scale proportionally.
- First 5 reference images are free; each additional image adds $0.04.
- Reference video input is billed per second at $0.05/s (480p), $0.10/s (720p), $0.20/s (1080p), added on top of the output price.
- Failed generations are not charged.
Best Use Cases
- Motion transfer — extract choreography or camera movement from a reference video and apply it to new subjects.
- Style transfer — apply a reference image's look to the generated scene.
- Multi-scene narrative — combine multiple references to maintain character and camera continuity across a 30-second cut.
- Precise start/end control — pin the exact first and last frame while the model fills in the motion between them.
- Soundtrack-guided generation — use a reference audio clip to drive the pacing and mood of the scene.
Pro Tips
- Combine a first-frame and last-frame image whenever the clip needs to land on a specific pose or composition.
- Keep reference videos short and focused — Wan 3.0 reads up to 15 seconds combined across all reference videos.
- Describe each reference's role explicitly in the prompt to avoid ambiguity.
- Use one primary camera instruction per prompt for the cleanest motion transfer.
- Start with 480p/5s for rapid iteration, then render the final version at 1080p.
Notes
- Up to 10 reference images, 5 reference videos (15s combined), and 5 reference audio clips (15s combined) per request.
- Native audio generation is enabled by default.
- Duration range: 2-30 seconds (continuous selection).
- Successor to Wan 2.7 Reference-to-Video with a longer duration ceiling and first/last-frame control.




