Seedance 2.5 Reference To Video
Seedance 2.5 (Reference-to-Video) generates video guided by up to 50 full-modal reference assets spanning images, videos, and audio clips. Use @Image1, @Video1, @Audio1 tags in your prompt to assign roles — style transfer, lip-sync, motion transfer, and multi-scene composition across native 30-second clips at up to 1080p.
Key Features
- Up to 50 full-modal reference assets — combine images, videos and audio in a single request, the richest reference budget in the Seedance family.
- @ mention syntax in prompts — use @image1, @video1, @audio1 tags to assign each reference asset a specific role with '= instruction' format.
- Native 30-second clips — orchestrate longer, multi-scene sequences from a single reference set.
- Style transfer & character consistency — preserves facial features, clothing, and artistic style across frames and scenes without drift.
- Motion transfer — extracts choreography, action sequences, and camera movement from reference videos and applies them to new scenes.
- Native audio-visual synchronization — generates video with matched sound effects, dialogue, and ambient audio in a single pass.
Parameters
| Parameter | Required | Description |
|---|---|---|
| prompt | Yes | Scene description followed by @image1, @video1, @audio1 tags with '= instruction' to assign each asset a role (e.g. '@image1 = her: keep the exact face, she is the lead dancer'). |
| image_urls | No | Reference image URLs for character, style, or scene guidance (up to 50 combined reference assets). |
| video_urls | No | Reference video URLs (each 1.8-30.2s) for motion and camera guidance. |
| audio_urls | No | Reference audio URLs for lip-sync or soundtrack guidance. |
| duration | No | Video length in seconds: 4-30 (default: 5). |
| aspect_ratio | No | Output format: 16:9 (default), 9:16, 4:3, 3:4, 1:1, 21:9. |
| resolution | No | Output resolution: 480p, 720p (default), or 1080p. |
| bitrate_mode | No | Bitrate quality: standard (default) or high. |
| generate_audio | No | Generate synchronized audio (default: true). |
| camera | No | Camera movement: static, pan-left, pan-right, zoom-in, zoom-out, tilt-up, tilt-down. |
| seed | No | Random seed for reproducibility (-1 = random). |
How to Use
- Upload reference assets — images define appearance/style, videos define motion/camera, audio drives lip-sync/rhythm.
- Write a prompt — describe the scene first, then use @image1 = instruction, @video1 = instruction to tell the model what to do with each asset (e.g. keep face, copy choreography).
- Choose parameters — duration (4-30s), resolution (up to 1080p), aspect ratio, and camera movement.
- Generate — submit and wait, then download the finished video with synchronized audio.
- Iterate — start with 480p/5s for quick iterations, then render final output at 1080p.
Code Examples
import os
import requests
response = requests.post(
"https://aircube.ai/api/v3/seedance-2.5/reference-to-video",
headers={
"Authorization": "Bearer " + os.environ["AIRCUBE_API_KEY"],
"Content-Type": "application/json",
},
json={
"prompt": "A woman walks through a neon-lit alley at night, moody cinematic atmosphere.\n@image1 = her: keep the exact face and outfit, she walks through the alley\n@audio1 = use as the ambient background soundtrack",
"image_urls": [
"https://example.com/character.jpg"
],
"audio_urls": [
"https://example.com/ambient.mp3"
],
"duration": 8,
"resolution": "1080p",
"aspect_ratio": "16:9",
"generate_audio": true
},
timeout=300,
)
data = response.json()
if data["success"]:
print("ID:", data["data"]["id"], "Status:", data["data"]["status"])
else:
print("Error:", data["error"]["message"])Pricing
| Resolution | Duration | Cost |
|---|---|---|
| 480p | 4s | $0.76 |
| 480p | 5s | $0.95 |
| 480p | 6s | $1.14 |
| 480p | 8s | $1.52 |
| 480p | 10s | $1.90 |
| 480p | 12s | $2.28 |
| 480p | 15s | $2.85 |
| 480p | 20s | $3.80 |
| 480p | 25s | $4.75 |
| 480p | 30s | $5.70 |
| 720p | 4s | $1.76 |
| 720p | 5s | $2.20 |
| 720p | 6s | $2.64 |
| 720p | 8s | $3.52 |
| 720p | 10s | $4.40 |
| 720p | 12s | $5.28 |
| 720p | 15s | $6.60 |
| 720p | 20s | $8.80 |
| 720p | 25s | $11.00 |
| 720p | 30s | $13.20 |
| 1080p | 4s | $4.20 |
| 1080p | 5s | $5.25 |
| 1080p | 6s | $6.30 |
| 1080p | 8s | $8.40 |
| 1080p | 10s | $10.50 |
| 1080p | 12s | $12.60 |
| 1080p | 15s | $15.75 |
| 1080p | 20s | $21.00 |
| 1080p | 25s | $26.25 |
| 1080p | 30s | $31.50 |
Billing rules
- 480p / 4s: $0.76.
- 720p / 4s: $1.76.
- 1080p / 4s: $4.20.
- Longer durations scale proportionally.
- First 5 reference images are free; each additional image adds $0.04.
- Reference video input is billed per second at $0.12/s (480p), $0.26/s (720p), $0.63/s (1080p), added on top of the output price.
- Failed generations are not charged.
Best Use Cases
- Style transfer — apply a reference image's artistic style to generated video.
- Action / choreography cloning — extract motion from a reference video and apply it to new characters or scenes.
- Lip-sync — audio-driven mouth synchronization.
- Multi-scene narrative — combine up to 50 references to maintain character, scene, and camera continuity across a 30-second cut.
- Camera replication — extract dolly, tracking, or orbit camera techniques from a reference video.
Pro Tips
- With 1-2 images the model uses keyframe mode (faster); with 3+ assets or any video it switches to reference mode.
- Use the '@asset = instruction' format to assign each asset a clear role — e.g. '@image1 = her: keep the exact face and outfit, she is the lead dancer'.
- Use one primary camera instruction per prompt; add 'slow', 'smooth', or 'gentle' to control pacing.
- Audio references must be paired with at least one image or video.
- Take advantage of the 50-asset budget and 30-second ceiling for scenes with multiple characters or cuts.
Notes
- Up to 50 combined reference files per request (images + videos + audio).
- Native audio generation is enabled by default.
- Duration range: 4-30 seconds (continuous selection).
- Successor to Seedance 2.0 Reference-to-Video with a larger reference budget and longer duration ceiling.




