AirCubeAirCube
API Reference

Submit Generation

Submit an AI generation request to the AirCube API.

Endpoint

POST https://aircube.ai/api/v3/{model-slug}/{task}

Submits an asynchronous generation request. Returns immediately with a generation ID that you use to poll for the result.

Authentication

Authorization: Bearer YOUR_API_KEY
Content-Type: application/json

URL format

The URL path determines which model and task to run:

POST /api/v3/{model-slug}/{task}
SegmentDescriptionExample
model-slugModel identifier (from GET /models)seedream-v5.0-pro
taskWhat to generatetext-to-image, image-to-video

The task segment determines the output type:

Task containsOutput
videoVideo file
audio, music, speech, ttsAudio file
anything elseImage file

Example URLs

POST /api/v3/seedream-v5.0-pro/text-to-image
POST /api/v3/seedance-2.0/image-to-video
POST /api/v3/veo3.1/text-to-video
POST /api/v3/qwen3-tts/text-to-speech
POST /api/v3/music-2.5/text-to-music

Parameters by task

Different tasks accept different parameters. Below is the reference for each task type.

text-to-image

Generate an image from a text prompt.

{
  "prompt": "a photorealistic glass cube on a marble surface",
  "aspect_ratio": "1:1",
  "resolution": "2k",
  "quality": "high",
  "output_format": "png"
}
ParameterTypeRequiredDescription
promptstringYesText description of the image to generate
aspect_ratiostringNoOutput aspect ratio: "1:1", "16:9", "9:16", "4:3", "3:4"
resolutionstringNoOutput resolution: "1k", "2k", "4k". Default: "1k"
qualitystringNoRendering quality: "low", "medium", "high". GPT Image models only
countnumberNoNumber of images to generate (1-4). Default: 1
output_formatstringNoImage format: "png", "jpeg", "webp". Default varies by model

image-to-image

Transform or edit an existing image.

{
  "prompt": "make the background a sunset beach",
  "image": "https://example.com/input.jpg",
  "resolution": "2k"
}
ParameterTypeRequiredDescription
promptstringYesEdit instruction or target description
imagestringYesURL of the input image to transform
aspect_ratiostringNoOutput aspect ratio
resolutionstringNoOutput resolution: "1k", "2k", "4k"
qualitystringNo"low", "medium", "high". GPT Image models only
output_formatstringNo"png", "jpeg", "webp"

text-to-video

Generate a video from a text prompt.

{
  "prompt": "a drone shot flying over a coastal city at golden hour",
  "duration": "5",
  "aspect_ratio": "16:9",
  "resolution": "720p"
}
ParameterTypeRequiredDescription
promptstringYesText description of the video to generate
durationstringNoVideo length: "4", "5", "6", "8", "10", "12", "15" (seconds). Default: "5"
aspect_ratiostringNoOutput aspect ratio: "16:9", "9:16", "1:1"
resolutionstringNoOutput resolution: "480p", "720p", "1080p", "4k". Default: "720p"

image-to-video

Animate a still image into a video.

{
  "prompt": "the flowers sway gently in the wind",
  "image": "https://example.com/photo.jpg",
  "duration": "5",
  "resolution": "720p"
}
ParameterTypeRequiredDescription
imagestringYesURL of the source image to animate
promptstringNoMotion description (recommended for better results)
durationstringNoVideo length in seconds. Default: "5"
aspect_ratiostringNoOutput aspect ratio
resolutionstringNo"480p", "720p", "1080p", "4k". Default: "720p"

reference-to-video

Generate a video guided by reference images, videos, or audio clips.

{
  "prompt": "@Image1 walks through a garden",
  "reference_images": ["https://example.com/character.jpg"],
  "reference_videos": ["https://example.com/motion-ref.mp4"],
  "reference_audios": ["https://example.com/voice.mp3"],
  "duration": "5",
  "resolution": "720p"
}
ParameterTypeRequiredDescription
promptstringYesUse @Image1, @Video1, @Audio1 tags to reference uploaded media
reference_imagesstring[]NoUp to 9 reference image URLs
reference_videosstring[]NoUp to 3 reference video URLs
reference_audiosstring[]NoUp to 3 reference audio URLs
durationstringNoVideo length in seconds
aspect_ratiostringNoOutput aspect ratio
resolutionstringNo"480p", "720p", "1080p", "4k"

At least one reference media array must be provided.

video-to-video

Transform an existing video.

{
  "prompt": "convert to anime style",
  "image": "https://example.com/source-video.mp4",
  "resolution": "720p"
}
ParameterTypeRequiredDescription
promptstringYesTransformation description
imagestringYesURL of the source video
resolutionstringNoOutput resolution

video-extend

Extend the duration of an existing video.

{
  "prompt": "continue the scene naturally",
  "image": "https://example.com/source-video.mp4",
  "duration": "5"
}
ParameterTypeRequiredDescription
imagestringYesURL of the video to extend
promptstringNoDescription of how to continue
durationstringNoAdditional seconds to generate

text-to-speech

Synthesize speech from text.

{
  "prompt": "Hello, welcome to our platform.",
  "voice": "alloy"
}
ParameterTypeRequiredDescription
promptstringYesText to convert to speech
voicestringNoVoice ID. Available voices depend on the model

text-to-music

Generate music from a text description.

{
  "prompt": "upbeat electronic track with synth pads, 120 BPM",
  "duration": "60"
}
ParameterTypeRequiredDescription
promptstringYesMusic description (genre, mood, instruments, tempo)
durationstringNoDuration: "30", "60", "120" (seconds). Default: "60"

Input media handling

When you provide URLs in image, images, reference_images, reference_videos, or reference_audios, AirCube automatically downloads the media and stores it for processing. Supported formats:

  • Images: jpg, jpeg, png, webp, gif
  • Videos: mp4, webm, mkv
  • Audio: mp3, wav, m4a

URLs must be publicly accessible. Private or authenticated URLs will fail.


Response

Status: 202 Accepted

{
  "success": true,
  "data": {
    "id": "cm5abc123def456",
    "status": "processing",
    "model": "Seedream 5.0 Pro",
    "type": "image",
    "output_url": null,
    "created_at": "2025-01-15T12:00:00.000Z"
  }
}
FieldTypeDescription
idstringGeneration ID — use this to poll for results
statusstringAlways "processing" on submission
modelstringResolved display name of the model
typestring"image", "video", or "audio"
output_urlnullPopulated when status becomes "completed"
created_atstringISO 8601 timestamp

Credits are charged immediately. If the generation fails, credits are automatically refunded.


Error responses

StatusCodeWhen
400VALIDATION_ERRORInvalid JSON body, missing required fields, or invalid media URL
401UNAUTHORIZEDMissing or invalid API key
402PAYMENT_REQUIREDInsufficient credit balance
403FORBIDDENAPI key is disabled or expired
404NOT_FOUNDModel not found or task not supported
429RATE_LIMITEDRate limit exceeded

See Error Codes for the full error reference.


Examples

Text to image (cURL)

curl -X POST https://aircube.ai/api/v3/seedream-v5.0-pro/text-to-image \
  -H "Authorization: Bearer $AIRCUBE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "prompt": "a cyberpunk cityscape at night, neon reflections on wet streets",
    "aspect_ratio": "16:9",
    "resolution": "2k"
  }'

Image to video (Python)

import os
import requests

API_KEY = os.environ["AIRCUBE_API_KEY"]
BASE = "https://aircube.ai/api/v3"
HEADERS = {
    "Authorization": f"Bearer {API_KEY}",
    "Content-Type": "application/json",
}

response = requests.post(
    f"{BASE}/seedance-2.0/image-to-video",
    headers=HEADERS,
    json={
        "prompt": "the camera slowly zooms in as petals fall",
        "image": "https://example.com/flower.jpg",
        "duration": "5",
        "resolution": "720p",
    },
)
generation_id = response.json()["data"]["id"]

Text to speech (JavaScript)

const res = await fetch(
  "https://aircube.ai/api/v3/qwen3-tts/text-to-speech",
  {
    method: "POST",
    headers: {
      Authorization: `Bearer ${process.env.AIRCUBE_API_KEY}`,
      "Content-Type": "application/json",
    },
    body: JSON.stringify({
      prompt: "Welcome to AirCube. The unified API for AI generation.",
    }),
  }
);
const { data } = await res.json();
console.log(`Generation ID: ${data.id}`);

What happens next

Use the generation id from the response to poll for results. Typical generation times:

TypeTypical time
Image3–15 seconds
Video30 seconds – 5 minutes
Audio3–15 seconds

On this page