New — access 450+ AI models through one unified API.Start building free

seedance-2.0/reference-to-video

15% OFF

Seedance 2.0 (Reference-to-Video) generates cinematic videos guided by up to 12 reference files spanning images, videos, and audio clips. Use @Image1, @Video1, @Audio1 tags in your prompt to assign roles — style transfer, lip-sync, motion transfer, character consistency, and multi-scene composition. Supports 480P / 720P / 1080P / 4K output, 4-15s duration, and flexible aspect ratios. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

reference-to-video$0.4080$0.4800/ run

Prompt

0/2000
Reference MediaImages, videos, and audio supported
PreviewJSON

Your result will appear here

Write a prompt and hit Generate.

Examples

Seedance 2.0 example 1

Seedance 2.0 example 2

Seedance 2.0 example 3

Seedance 2.0 example 4

Related models

seedance-2.0 / image-to-video15% OFFimage-to-video
seedance-2.0 / image-to-video

Seedance 2.0 (Image-to-Video) generates Hollywood-grade cinematic videos from reference images and text prompts with native audio-visual synchronization, director-level camera and lighting control, and exceptional motion stability. Built on Seed's unified multimodal architecture, it preserves the input image's subject and composition while adding expressive, physically accurate motion.

Seedance$0.4080$0.4800
seedance-2.0 / text-to-video15% OFFtext-to-video
seedance-2.0 / text-to-video

Seedance 2.0 (Text-to-Video) generates Hollywood-grade cinematic videos from text prompts with native audio-visual synchronization, director-level camera and lighting control, and exceptional motion stability. Built on Seed's unified multimodal architecture, it leads on instruction adherence, motion quality, and visual aesthetics.

Seedance$0.4080$0.4800
seedance-2.0 / image-to-video-spicy15% OFFimage-to-video
seedance-2.0 / image-to-video-spicy

Seedance 2.0 Spicy Image to Video is a fast AI image-to-video generation model that creates high-quality cinematic clips from images, optimized for scalable content generation with smooth animations and stable aesthetics. Ready-to-use REST inference API for animating images, social media clips, product videos, advertising creatives, visual storytelling, and professional image-to-video workflows with simple integration, no coldstarts, and affordable pricing.

Seedance$0.5100$0.6000
seedance-2.0 / text-to-video-spicy15% OFFtext-to-video
seedance-2.0 / text-to-video-spicy

Seedance 2.0 Spicy Text to Video generates high-quality cinematic clips from text prompts, optimized for scalable content generation with smooth animations and stable aesthetics. Ready-to-use REST inference API for creating social media clips, product videos, advertising creatives, visual storytelling, and professional text-to-video workflows with simple integration, no coldstarts, and affordable pricing.

Seedance$0.5100$0.6000
seedance-2.0 / reference-to-video-spicy15% OFFreference-to-video
seedance-2.0 / reference-to-video-spicy

Seedance 2.0 Spicy Reference to Video generates cinematic videos guided by up to 12 reference files spanning images, videos, and audio clips. Use @Image1, @Video1, @Audio1 tags in your prompt to assign roles — style transfer, lip-sync, motion transfer, character consistency, and multi-scene composition. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Seedance$0.5100$0.6000
seedance-2.0-fast / image-to-video15% OFFimage-to-video
seedance-2.0-fast / image-to-video

Seedance 2.0 Fast (Image-to-Video) generates cinematic videos from reference images and text prompts with native audio-visual synchronization, director-level control, and exceptional motion stability — optimized for faster generation at lower cost. Built on Seed's unified multimodal architecture.

Seedance$0.2210$0.2600

Seedance 2.0 Reference To Video

Seedance 2.0 (Reference-to-Video) generates cinematic videos guided by up to 12 reference files spanning images, videos, and audio clips. Use @Image1, @Video1, @Audio1 tags in your prompt to assign roles — style transfer, lip-sync, motion transfer, character consistency, and multi-scene composition. Supports 480P / 720P / 1080P / 4K output, 4-15s duration, and flexible aspect ratios. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Key Features

  • Multimodal reference inputs — combine up to 12 reference files (9 images + 3 videos + 3 audio clips) to orchestrate a unified video.
  • @ mention syntax in prompts — use @image1, @video1, @audio1 tags to assign each reference asset a specific role with '= instruction' format.
  • Style transfer & character consistency — preserves facial features, clothing, and artistic style across frames and scenes without drift.
  • Motion transfer — extracts choreography, action sequences, and camera movement from reference videos and applies them to new scenes.
  • Audio lip-sync — audio-driven mouth synchronization supporting 8+ languages.
  • Native audio-visual synchronization — generates video with matched sound effects, dialogue, and ambient audio in a single pass.

Parameters

ParameterRequiredDescription
promptYesScene description followed by @image1, @video1, @audio1 tags with '= instruction' to assign each asset a role (e.g. '@image1 = her: keep the exact face, she is the lead dancer').
image_urlsNoUp to 9 reference image URLs for character, style, or scene guidance.
video_urlsNoUp to 3 reference video URLs (MP4/MOV, 480p-720p) for motion and camera guidance.
audio_urlsNoUp to 3 audio files (MP3/WAV, total ≤15s, each ≤15MB) for lip-sync or soundtrack.
durationNoVideo length in seconds: 4-15 (default: 5).
aspect_ratioNoOutput format: 16:9 (default), 9:16, 4:3, 3:4, 1:1, 21:9.
resolutionNoOutput resolution: 480p, 720p (default), 1080p, or 4k.
bitrate_modeNoBitrate quality: standard (default) or high.
generate_audioNoGenerate synchronized audio (default: true).
cameraNoCamera movement: static, pan-left, pan-right, zoom-in, zoom-out, tilt-up, tilt-down.
seedNoRandom seed for reproducibility (-1 = random).

How to Use

  1. Upload reference assets — images define appearance/style, videos define motion/camera, audio drives lip-sync/rhythm.
  2. Write a prompt — describe the scene first, then use @image1 = instruction, @video1 = instruction to tell the model what to do with each asset (e.g. keep face, copy choreography).
  3. Choose parameters — duration (4-15s), resolution (up to 4K), aspect ratio, and camera movement.
  4. Generate — submit and wait, then download the finished video with synchronized audio.
  5. Iterate — start with 480p/5s for quick iterations, then render final output at 1080p or 4K.

Code Examples

import os
import requests

response = requests.post(
    "https://aircube.ai/api/v3/seedance-2.0/reference-to-video",
    headers={
        "Authorization": "Bearer " + os.environ["AIRCUBE_API_KEY"],
        "Content-Type": "application/json",
    },
    json={
    "prompt": "A woman walks through a neon-lit alley at night, moody cinematic atmosphere.\n@image1 = her: keep the exact face and outfit, she walks through the alley\n@audio1 = use as the ambient background soundtrack",
    "image_urls": [
        "https://example.com/character.jpg"
    ],
    "audio_urls": [
        "https://example.com/ambient.mp3"
    ],
    "duration": 8,
    "resolution": "1080p",
    "aspect_ratio": "16:9",
    "generate_audio": true
},
    timeout=300,
)
data = response.json()

if data["success"]:
    print("ID:", data["data"]["id"], "Status:", data["data"]["status"])
else:
    print("Error:", data["error"]["message"])

Pricing

ResolutionDurationCost
480p4s$0.48
480p5s$0.60
480p6s$0.72
480p8s$0.96
480p10s$1.20
480p12s$1.44
480p15s$1.80
720p4s$0.96
720p5s$1.20
720p6s$1.44
720p8s$1.92
720p10s$2.40
720p12s$2.88
720p15s$3.60
1080p4s$2.40
1080p5s$3.00
1080p6s$3.60
1080p8s$4.80
1080p10s$6.00
1080p12s$7.20
1080p15s$9.00
4k4s$4.80
4k5s$6.00
4k6s$7.20
4k8s$9.60
4k10s$12.00
4k12s$14.40
4k15s$18.00

Billing rules

  • 480p / 4s: $0.48.
  • 720p / 4s: $0.96.
  • 1080p / 4s: $2.40.
  • 4k / 4s: $4.80.
  • Longer durations scale proportionally.
  • Failed generations are not charged.

Best Use Cases

  • Style transfer — apply a reference image's artistic style to generated video.
  • Action / choreography cloning — extract motion from a reference video and apply it to new characters or scenes.
  • Lip-sync — audio-driven mouth synchronization in 8+ languages.
  • Multi-scene narrative — combine multiple references to maintain character, scene, and camera continuity across cuts.
  • Camera replication — extract dolly, tracking, or orbit camera techniques from a reference video.

Pro Tips

  • With 1-2 images the model uses keyframe mode (faster); with 3+ assets or any video it switches to reference mode.
  • Use the '@asset = instruction' format to assign each asset a clear role — e.g. '@image1 = her: keep the exact face and outfit, she is the lead dancer'.
  • Use one primary camera instruction per prompt; add 'slow', 'smooth', or 'gentle' to control pacing.
  • Audio references must be paired with at least one image or video.
  • Start with 480p/5s for rapid iteration, then render the final version at 1080p or 4K.

Notes

  • Maximum 12 reference files per request (9 images + 3 videos + 3 audio clips).
  • Total audio duration ≤15 seconds, each audio file ≤15MB.
  • Native audio generation is enabled by default.
  • Duration range: 4-15 seconds (continuous selection).
  • Median generation time: approximately 170-250 seconds.

Seedance 2.0 Reference To Video — Frequently asked questions