# wan-3.0/reference-to-video

> Wan 3.0 (Reference-to-Video) generates video guided by first/last frame images plus multiple reference images, videos, and audio clips, producing native 30-second scenes with synchronized audio at up to 1080p — Alibaba's most capable Wan model for multimodal creative control.

- **Provider:** Alibaba
- **Category:** reference-to-video
- **Price:** $0.2000 per run

## Key Features

- Multimodal reference inputs — combine first/last frame images with additional reference images, videos, and audio in a single request.
- Motion and camera reference — extract choreography and camera movement from a reference video and apply it to a new scene.
- Style and character consistency — reference images preserve appearance and style across the generated clip.
- Native audio-video synthesis — synchronized dialogue, music and sound effects generated automatically alongside the picture.
- Native 30-second clips — orchestrate longer, multi-scene sequences from a single reference set.
- Bilingual prompt support — supports both English and Chinese prompts natively.

## Parameters

| Parameter | Required | Description |
| --- | --- | --- |
| `prompt` | Yes | Scene description guiding how the reference assets should be used (English or Chinese). |
| `image` | No | First-frame reference image URL. |
| `last_image` | No | Last-frame reference image URL for precise start/end-frame control. |
| `reference_images` | No | Additional reference image URLs for style or character guidance (up to 10). |
| `reference_videos` | No | Reference video URLs for motion and camera guidance (up to 5, 15s combined). |
| `reference_audios` | No | Reference audio URLs for soundtrack or voice guidance (up to 5, 15s combined). |
| `duration` | No | Video length in seconds: 2-30 (default: 5). |
| `resolution` | No | Output resolution: 480p, 720p (default), or 1080p. |
| `aspect_ratio` | No | Output format: 16:9 (default), 9:16, 4:3, 3:4, 1:1, or adaptive. |

## How to Use

1. Upload reference assets — first/last frame images anchor the start and end, additional images define style/character, videos define motion/camera, audio drives soundtrack.
2. Write a prompt describing how the references should come together into a scene.
3. Choose duration (up to 30s), resolution, and aspect ratio.
4. Generate — submit and wait, then download the finished video with synchronized audio.
5. Iterate — start with 480p/5s for quick iterations, then render the final version at 1080p.

## Code Examples

### Python

```python
import os
import requests

response = requests.post(
    "https://aircube.ai/api/v3/wan-3.0/reference-to-video",
    headers={
        "Authorization": "Bearer " + os.environ["AIRCUBE_API_KEY"],
        "Content-Type": "application/json",
    },
    json={
    "prompt": "A dancer performs the same choreography as the reference video, but on a neon-lit rooftop at night",
    "image": "https://example.com/dancer-first-frame.jpg",
    "reference_videos": [
        "https://example.com/choreography-reference.mp4"
    ],
    "duration": 8,
    "resolution": "1080p",
    "aspect_ratio": "16:9"
},
    timeout=300,
)
data = response.json()

if data["success"]:
    print("ID:", data["data"]["id"], "Status:", data["data"]["status"])
else:
    print("Error:", data["error"]["message"])
```

### Node.js

```javascript
const response = await fetch("https://aircube.ai/api/v3/wan-3.0/reference-to-video", {
  method: "POST",
  headers: {
    "Authorization": "Bearer " + process.env.AIRCUBE_API_KEY,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
  "prompt": "A dancer performs the same choreography as the reference video, but on a neon-lit rooftop at night",
  "image": "https://example.com/dancer-first-frame.jpg",
  "reference_videos": [
    "https://example.com/choreography-reference.mp4"
  ],
  "duration": 8,
  "resolution": "1080p",
  "aspect_ratio": "16:9"
}),
});

const data = await response.json();

if (data.success) {
  console.log("ID:", data.data.id, "Status:", data.data.status);
} else {
  console.error("Error:", data.error.message);
}
```

### cURL

```curl
curl -X POST "https://aircube.ai/api/v3/wan-3.0/reference-to-video" \
  -H "Authorization: Bearer $AIRCUBE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "prompt": "A dancer performs the same choreography as the reference video, but on a neon-lit rooftop at night",
  "image": "https://example.com/dancer-first-frame.jpg",
  "reference_videos": [
    "https://example.com/choreography-reference.mp4"
  ],
  "duration": 8,
  "resolution": "1080p",
  "aspect_ratio": "16:9"
}'
```

### Python (Async)

```python
import os
import time
import requests

# 1. Submit
response = requests.post(
    "https://aircube.ai/api/v3/wan-3.0/reference-to-video",
    headers={
        "Authorization": "Bearer " + os.environ["AIRCUBE_API_KEY"],
        "Content-Type": "application/json",
    },
    json={
    "prompt": "A dancer performs the same choreography as the reference video, but on a neon-lit rooftop at night",
    "image": "https://example.com/dancer-first-frame.jpg",
    "reference_videos": [
        "https://example.com/choreography-reference.mp4"
    ],
    "duration": 8,
    "resolution": "1080p",
    "aspect_ratio": "16:9"
},
    timeout=300,
)
data = response.json()

if not data["success"]:
    print("Error:", data["error"]["message"])
    exit(1)

generation_id = data["data"]["id"]
print(f"Submitted: {generation_id}")

# 2. Poll until completed
while True:
    time.sleep(5)
    r = requests.get(
        f"https://aircube.ai/api/v3/status/{generation_id}",
        headers={"Authorization": "Bearer " + os.environ["AIRCUBE_API_KEY"]},
    )
    result = r.json()["data"]

    if result["status"] == "completed":
        print(result["output_url"])
        break
    elif result["status"] == "failed":
        print("Generation failed")
        break
```

### Node.js (Async)

```javascript
// 1. Submit
const response = await fetch("https://aircube.ai/api/v3/wan-3.0/reference-to-video", {
  method: "POST",
  headers: {
    "Authorization": "Bearer " + process.env.AIRCUBE_API_KEY,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
  "prompt": "A dancer performs the same choreography as the reference video, but on a neon-lit rooftop at night",
  "image": "https://example.com/dancer-first-frame.jpg",
  "reference_videos": [
    "https://example.com/choreography-reference.mp4"
  ],
  "duration": 8,
  "resolution": "1080p",
  "aspect_ratio": "16:9"
}),
});
const data = await response.json();

if (!data.success) {
  console.error("Error:", data.error.message);
  process.exit(1);
}

const generationId = data.data.id;
console.log("Submitted:", generationId);

// 2. Poll until completed
while (true) {
  await new Promise((r) => setTimeout(r, 5000));
  const res = await fetch(
    `https://aircube.ai/api/v3/status/${generationId}`,
    { headers: { "Authorization": "Bearer " + process.env.AIRCUBE_API_KEY } },
  );
  const result = (await res.json()).data;

  if (result.status === "completed") {
    console.log(result.output_url);
    break;
  } else if (result.status === "failed") {
    console.error("Generation failed");
    break;
  }
}
```

### cURL (Async)

```curl
# 1. Submit
RESPONSE=$(curl -s -X POST "https://aircube.ai/api/v3/wan-3.0/reference-to-video" \
  -H "Authorization: Bearer $AIRCUBE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "prompt": "A dancer performs the same choreography as the reference video, but on a neon-lit rooftop at night",
  "image": "https://example.com/dancer-first-frame.jpg",
  "reference_videos": [
    "https://example.com/choreography-reference.mp4"
  ],
  "duration": 8,
  "resolution": "1080p",
  "aspect_ratio": "16:9"
}')

ID=$(echo "$RESPONSE" | jq -r '.data.id')
echo "Submitted: $ID"

# 2. Poll until completed
while true; do
  sleep 5
  STATUS_RES=$(curl -s "https://aircube.ai/api/v3/status/$ID" \
    -H "Authorization: Bearer $AIRCUBE_API_KEY")
  STATUS=$(echo "$STATUS_RES" | jq -r '.data.status')

  if [ "$STATUS" = "completed" ]; then
    echo "$STATUS_RES" | jq -r '.data.output_url'
    break
  elif [ "$STATUS" = "failed" ]; then
    echo "Generation failed"; break
  fi
done
```

## Pricing

| Resolution | Duration | Cost |
| --- | --- | --- |
| 480p | 4s | $0.20 |
| 480p | 5s | $0.25 |
| 480p | 6s | $0.30 |
| 480p | 8s | $0.40 |
| 480p | 10s | $0.50 |
| 480p | 12s | $0.60 |
| 480p | 15s | $0.75 |
| 480p | 20s | $1.00 |
| 480p | 25s | $1.25 |
| 480p | 30s | $1.50 |
| 720p | 4s | $0.40 |
| 720p | 5s | $0.50 |
| 720p | 6s | $0.60 |
| 720p | 8s | $0.80 |
| 720p | 10s | $1.00 |
| 720p | 12s | $1.20 |
| 720p | 15s | $1.50 |
| 720p | 20s | $2.00 |
| 720p | 25s | $2.50 |
| 720p | 30s | $3.00 |
| 1080p | 4s | $0.80 |
| 1080p | 5s | $1.00 |
| 1080p | 6s | $1.20 |
| 1080p | 8s | $1.60 |
| 1080p | 10s | $2.00 |
| 1080p | 12s | $2.40 |
| 1080p | 15s | $3.00 |
| 1080p | 20s | $4.00 |
| 1080p | 25s | $5.00 |
| 1080p | 30s | $6.00 |

### Billing Rules

- 480p / 4s: $0.20.
- 720p / 4s: $0.40.
- 1080p / 4s: $0.80.
- Longer durations scale proportionally.
- First 5 reference images are free; each additional image adds $0.04.
- Reference video input is billed per second at $0.05/s (480p), $0.10/s (720p), $0.20/s (1080p), added on top of the output price.
- Failed generations are not charged.

## Best Use Cases

- Motion transfer — extract choreography or camera movement from a reference video and apply it to new subjects.
- Style transfer — apply a reference image's look to the generated scene.
- Multi-scene narrative — combine multiple references to maintain character and camera continuity across a 30-second cut.
- Precise start/end control — pin the exact first and last frame while the model fills in the motion between them.
- Soundtrack-guided generation — use a reference audio clip to drive the pacing and mood of the scene.

## Pro Tips

- Combine a first-frame and last-frame image whenever the clip needs to land on a specific pose or composition.
- Keep reference videos short and focused — Wan 3.0 reads up to 15 seconds combined across all reference videos.
- Describe each reference's role explicitly in the prompt to avoid ambiguity.
- Use one primary camera instruction per prompt for the cleanest motion transfer.
- Start with 480p/5s for rapid iteration, then render the final version at 1080p.

## Notes

- Up to 10 reference images, 5 reference videos (15s combined), and 5 reference audio clips (15s combined) per request.
- Native audio generation is enabled by default.
- Duration range: 2-30 seconds (continuous selection).
- Successor to Wan 2.7 Reference-to-Video with a longer duration ceiling and first/last-frame control.

## FAQ

**Q: What is the Wan 3.0 Reference-to-Video API?**

Alibaba's most capable Wan model for multimodal video generation — combine first/last frame images with additional reference images, videos, and audio to guide a native 30-second scene, available via AirCube.

**Q: What is the difference between Reference-to-Video and Image-to-Video?**

Image-to-Video animates a single first/last-frame pair. Reference-to-Video additionally accepts style/character reference images, motion reference videos, and soundtrack reference audio.

**Q: How many reference assets can I use?**

Up to 10 reference images, 5 reference videos (15 seconds combined), and 5 reference audio clips (15 seconds combined) in a single request.

**Q: Does it preserve motion from a reference video?**

Yes — the model can extract choreography and camera movement from a reference video and apply it to a new scene.

**Q: How much does it cost?**

Starting at $0.25 for a 480p 5-second clip, scaling with resolution and duration. The number of reference assets does not affect pricing.

**Q: Can I use outputs commercially?**

Yes — generated videos are yours to use commercially.
