New — access 450+ AI models through one unified API.Start building free

seedance-2.5/text-to-video

Seedance 2.5 (Text-to-Video) generates native cinematic video up to 30 seconds directly from a prompt, with native audio-visual synchronization and stronger precision in editing and generation control than Seedance 2.0. Built on ByteDance's Seed architecture for production-ready creative workflows.

text-to-video$0.7600/ run

Prompt

0/2000
PreviewJSON

Your result will appear here

Write a prompt and hit Generate.

Examples

A realistic cinematic close-up shot of a beautiful young blonde European woman blowing a bubble gum bubble. She has long blonde hair, soft natural makeup, fair skin, and a stylish casual outfit. Warm natural daylight, shallow depth of field, film grain texture.

A cinematic ocean wave at sunrise, highly detailed.

Related models

seedance-2.5 / image-to-videoimage-to-video
seedance-2.5 / image-to-video

Seedance 2.5 (Image-to-Video) turns a reference image and prompt into native cinematic video up to 30 seconds, with up to 50 full-modal reference assets, richer editing control, and native audio-visual synchronization. Built on ByteDance's next-generation Seed architecture, it preserves the input image's subject and composition while adding expressive, physically accurate motion.

Seedance$0.7600
seedance-2.5 / reference-to-videoreference-to-video
seedance-2.5 / reference-to-video

Seedance 2.5 (Reference-to-Video) generates video guided by up to 50 full-modal reference assets spanning images, videos, and audio clips. Use @Image1, @Video1, @Audio1 tags in your prompt to assign roles — style transfer, lip-sync, motion transfer, and multi-scene composition across native 30-second clips at up to 1080p.

Seedance$0.7600
seedance-2.0 / image-to-video15% OFFimage-to-video
seedance-2.0 / image-to-video

Seedance 2.0 (Image-to-Video) generates Hollywood-grade cinematic videos from reference images and text prompts with native audio-visual synchronization, director-level camera and lighting control, and exceptional motion stability. Built on Seed's unified multimodal architecture, it preserves the input image's subject and composition while adding expressive, physically accurate motion.

Seedance$0.4080$0.4800
seedance-2.0 / text-to-video15% OFFtext-to-video
seedance-2.0 / text-to-video

Seedance 2.0 (Text-to-Video) generates Hollywood-grade cinematic videos from text prompts with native audio-visual synchronization, director-level camera and lighting control, and exceptional motion stability. Built on Seed's unified multimodal architecture, it leads on instruction adherence, motion quality, and visual aesthetics.

Seedance$0.4080$0.4800
seedance-2.0 / reference-to-video15% OFFreference-to-video
seedance-2.0 / reference-to-video

Seedance 2.0 (Reference-to-Video) generates cinematic videos guided by up to 12 reference files spanning images, videos, and audio clips. Use @Image1, @Video1, @Audio1 tags in your prompt to assign roles — style transfer, lip-sync, motion transfer, character consistency, and multi-scene composition. Supports 480P / 720P / 1080P / 4K output, 4-15s duration, and flexible aspect ratios. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Seedance$0.4080$0.4800
seedance-2.0 / image-to-video-spicy15% OFFimage-to-video
seedance-2.0 / image-to-video-spicy

Seedance 2.0 Spicy Image to Video is a fast AI image-to-video generation model that creates high-quality cinematic clips from images, optimized for scalable content generation with smooth animations and stable aesthetics. Ready-to-use REST inference API for animating images, social media clips, product videos, advertising creatives, visual storytelling, and professional image-to-video workflows with simple integration, no coldstarts, and affordable pricing.

Seedance$0.5100$0.6000

Seedance 2.5 Text To Video

Seedance 2.5 (Text-to-Video) generates native cinematic video up to 30 seconds directly from a prompt, with native audio-visual synchronization and stronger precision in editing and generation control than Seedance 2.0. Built on ByteDance's Seed architecture for production-ready creative workflows.

Key Features

  • Native 30-second clips generated directly from text — no reference image required.
  • Up to 50 full-modal reference assets — combine images, videos and audio to steer style, character and motion without leaving text-to-video mode.
  • Native audio-visual synchronization — generates video with synchronized sound effects, dialogue and ambience in a single pass.
  • Sharper editing and generation control than Seedance 2.0, courtesy of ByteDance's next-generation Seed architecture.
  • Director-level control over camera movement, lighting, shadows and character performance.
  • Strong instruction adherence — accurately renders complex multi-element, multi-scene scripts from detailed prompts.

Parameters

ParameterRequiredDescription
promptYesDetailed cinematic description of the scene to generate.
aspect_ratioNoOutput format: 16:9 (default), 9:16, 4:3, 3:4, 1:1, 21:9.
durationNoVideo length in seconds: 4-30 (default: 5).
resolutionNoOutput resolution: 480p, 720p (default), or 1080p.
reference_imagesNoUp to 50 reference image URLs for style/subject guidance.
reference_videosNoReference video URLs (each 1.8-30.2s) for motion guidance.
reference_audiosNoReference audio URLs for audio style guidance.
bitrate_modeNoBitrate quality: standard (default) or high.
generate_audioNoGenerate synchronized audio (default: true).

How to Use

  1. Write a cinematic prompt — describe subject, action, camera movement, lighting and mood.
  2. Choose aspect ratio: 16:9 for widescreen, 9:16 for vertical, or others.
  3. Set duration from 4 to 30 seconds.
  4. Optionally add reference images, videos, or audio for style guidance.
  5. Generate and download your video with synchronized audio.

Code Examples

import os
import requests

response = requests.post(
    "https://aircube.ai/api/v3/seedance-2.5/text-to-video",
    headers={
        "Authorization": "Bearer " + os.environ["AIRCUBE_API_KEY"],
        "Content-Type": "application/json",
    },
    json={
    "prompt": "Aerial shot of a coastal city at golden hour, camera slowly descending toward the waterfront, waves crashing against the pier",
    "duration": 8,
    "aspect_ratio": "16:9",
    "resolution": "1080p",
    "generate_audio": true
},
    timeout=300,
)
data = response.json()

if data["success"]:
    print("ID:", data["data"]["id"], "Status:", data["data"]["status"])
else:
    print("Error:", data["error"]["message"])

Pricing

ResolutionDurationCost
480p4s$0.76
480p5s$0.95
480p6s$1.14
480p8s$1.52
480p10s$1.90
480p12s$2.28
480p15s$2.85
480p20s$3.80
480p25s$4.75
480p30s$5.70
720p4s$1.76
720p5s$2.20
720p6s$2.64
720p8s$3.52
720p10s$4.40
720p12s$5.28
720p15s$6.60
720p20s$8.80
720p25s$11.00
720p30s$13.20
1080p4s$4.20
1080p5s$5.25
1080p6s$6.30
1080p8s$8.40
1080p10s$10.50
1080p12s$12.60
1080p15s$15.75
1080p20s$21.00
1080p25s$26.25
1080p30s$31.50

Billing rules

  • 480p / 4s: $0.76.
  • 720p / 4s: $1.76.
  • 1080p / 4s: $4.20.
  • Longer durations scale proportionally.
  • Failed generations are not charged.

Best Use Cases

  • Film production — generate longer cinematic footage from screenplays and treatments.
  • Commercials — create professional, multi-beat ad content directly from creative briefs.
  • Music videos — generate visual sequences from lyrical descriptions.
  • Premium social media — produce high-quality short-form video content.
  • Film visualization — prototype longer scenes and camera movements before production.

Pro Tips

  • Write prompts like a screenplay — subject, action, setting, camera, lighting.
  • Specify camera movements explicitly: 'slow dolly in', 'handheld tracking shot', 'aerial crane up'.
  • Include temporal language: 'gradually', 'suddenly', 'the camera slowly reveals'.
  • Use the extended 30-second ceiling and up to 50 reference assets for multi-scene narratives.
  • Start with 5s at 480p for fast iterations, then render final at 1080p.

Notes

  • Audio is generated natively by default — no need for separate audio tools.
  • Duration range: 4-30 seconds (continuous selection).
  • Supports up to 50 combined reference images, videos, and audio clips.
  • Successor to Seedance 2.0 with a longer duration ceiling and richer reference support.

Seedance 2.5 Text To Video — Frequently asked questions