Submit Generation
Submit an AI generation request to the AirCube API.
Endpoint
POST https://aircube.ai/api/v3/{model-slug}/{task}Submits an asynchronous generation request. Returns immediately with a generation ID that you use to poll for the result.
Authentication
Authorization: Bearer YOUR_API_KEY
Content-Type: application/jsonURL format
The URL path determines which model and task to run:
POST /api/v3/{model-slug}/{task}| Segment | Description | Example |
|---|---|---|
model-slug | Model identifier (from GET /models) | seedream-v5.0-pro |
task | What to generate | text-to-image, image-to-video |
The task segment determines the output type:
| Task contains | Output |
|---|---|
video | Video file |
audio, music, speech, tts | Audio file |
| anything else | Image file |
Example URLs
POST /api/v3/seedream-v5.0-pro/text-to-image
POST /api/v3/seedance-2.0/image-to-video
POST /api/v3/veo3.1/text-to-video
POST /api/v3/qwen3-tts/text-to-speech
POST /api/v3/music-2.5/text-to-musicParameters by task
Different tasks accept different parameters. Below is the reference for each task type.
text-to-image
Generate an image from a text prompt.
{
"prompt": "a photorealistic glass cube on a marble surface",
"aspect_ratio": "1:1",
"resolution": "2k",
"quality": "high",
"output_format": "png"
}| Parameter | Type | Required | Description |
|---|---|---|---|
prompt | string | Yes | Text description of the image to generate |
aspect_ratio | string | No | Output aspect ratio: "1:1", "16:9", "9:16", "4:3", "3:4" |
resolution | string | No | Output resolution: "1k", "2k", "4k". Default: "1k" |
quality | string | No | Rendering quality: "low", "medium", "high". GPT Image models only |
count | number | No | Number of images to generate (1-4). Default: 1 |
output_format | string | No | Image format: "png", "jpeg", "webp". Default varies by model |
image-to-image
Transform or edit an existing image.
{
"prompt": "make the background a sunset beach",
"image": "https://example.com/input.jpg",
"resolution": "2k"
}| Parameter | Type | Required | Description |
|---|---|---|---|
prompt | string | Yes | Edit instruction or target description |
image | string | Yes | URL of the input image to transform |
aspect_ratio | string | No | Output aspect ratio |
resolution | string | No | Output resolution: "1k", "2k", "4k" |
quality | string | No | "low", "medium", "high". GPT Image models only |
output_format | string | No | "png", "jpeg", "webp" |
text-to-video
Generate a video from a text prompt.
{
"prompt": "a drone shot flying over a coastal city at golden hour",
"duration": "5",
"aspect_ratio": "16:9",
"resolution": "720p"
}| Parameter | Type | Required | Description |
|---|---|---|---|
prompt | string | Yes | Text description of the video to generate |
duration | string | No | Video length: "4", "5", "6", "8", "10", "12", "15" (seconds). Default: "5" |
aspect_ratio | string | No | Output aspect ratio: "16:9", "9:16", "1:1" |
resolution | string | No | Output resolution: "480p", "720p", "1080p", "4k". Default: "720p" |
image-to-video
Animate a still image into a video.
{
"prompt": "the flowers sway gently in the wind",
"image": "https://example.com/photo.jpg",
"duration": "5",
"resolution": "720p"
}| Parameter | Type | Required | Description |
|---|---|---|---|
image | string | Yes | URL of the source image to animate |
prompt | string | No | Motion description (recommended for better results) |
duration | string | No | Video length in seconds. Default: "5" |
aspect_ratio | string | No | Output aspect ratio |
resolution | string | No | "480p", "720p", "1080p", "4k". Default: "720p" |
reference-to-video
Generate a video guided by reference images, videos, or audio clips.
{
"prompt": "@Image1 walks through a garden",
"reference_images": ["https://example.com/character.jpg"],
"reference_videos": ["https://example.com/motion-ref.mp4"],
"reference_audios": ["https://example.com/voice.mp3"],
"duration": "5",
"resolution": "720p"
}| Parameter | Type | Required | Description |
|---|---|---|---|
prompt | string | Yes | Use @Image1, @Video1, @Audio1 tags to reference uploaded media |
reference_images | string[] | No | Up to 9 reference image URLs |
reference_videos | string[] | No | Up to 3 reference video URLs |
reference_audios | string[] | No | Up to 3 reference audio URLs |
duration | string | No | Video length in seconds |
aspect_ratio | string | No | Output aspect ratio |
resolution | string | No | "480p", "720p", "1080p", "4k" |
At least one reference media array must be provided.
video-to-video
Transform an existing video.
{
"prompt": "convert to anime style",
"image": "https://example.com/source-video.mp4",
"resolution": "720p"
}| Parameter | Type | Required | Description |
|---|---|---|---|
prompt | string | Yes | Transformation description |
image | string | Yes | URL of the source video |
resolution | string | No | Output resolution |
video-extend
Extend the duration of an existing video.
{
"prompt": "continue the scene naturally",
"image": "https://example.com/source-video.mp4",
"duration": "5"
}| Parameter | Type | Required | Description |
|---|---|---|---|
image | string | Yes | URL of the video to extend |
prompt | string | No | Description of how to continue |
duration | string | No | Additional seconds to generate |
text-to-speech
Synthesize speech from text.
{
"prompt": "Hello, welcome to our platform.",
"voice": "alloy"
}| Parameter | Type | Required | Description |
|---|---|---|---|
prompt | string | Yes | Text to convert to speech |
voice | string | No | Voice ID. Available voices depend on the model |
text-to-music
Generate music from a text description.
{
"prompt": "upbeat electronic track with synth pads, 120 BPM",
"duration": "60"
}| Parameter | Type | Required | Description |
|---|---|---|---|
prompt | string | Yes | Music description (genre, mood, instruments, tempo) |
duration | string | No | Duration: "30", "60", "120" (seconds). Default: "60" |
Input media handling
When you provide URLs in image, images, reference_images, reference_videos, or reference_audios, AirCube automatically downloads the media and stores it for processing. Supported formats:
- Images: jpg, jpeg, png, webp, gif
- Videos: mp4, webm, mkv
- Audio: mp3, wav, m4a
URLs must be publicly accessible. Private or authenticated URLs will fail.
Response
Status: 202 Accepted
{
"success": true,
"data": {
"id": "cm5abc123def456",
"status": "processing",
"model": "Seedream 5.0 Pro",
"type": "image",
"output_url": null,
"created_at": "2025-01-15T12:00:00.000Z"
}
}| Field | Type | Description |
|---|---|---|
id | string | Generation ID — use this to poll for results |
status | string | Always "processing" on submission |
model | string | Resolved display name of the model |
type | string | "image", "video", or "audio" |
output_url | null | Populated when status becomes "completed" |
created_at | string | ISO 8601 timestamp |
Credits are charged immediately. If the generation fails, credits are automatically refunded.
Error responses
| Status | Code | When |
|---|---|---|
| 400 | VALIDATION_ERROR | Invalid JSON body, missing required fields, or invalid media URL |
| 401 | UNAUTHORIZED | Missing or invalid API key |
| 402 | PAYMENT_REQUIRED | Insufficient credit balance |
| 403 | FORBIDDEN | API key is disabled or expired |
| 404 | NOT_FOUND | Model not found or task not supported |
| 429 | RATE_LIMITED | Rate limit exceeded |
See Error Codes for the full error reference.
Examples
Text to image (cURL)
curl -X POST https://aircube.ai/api/v3/seedream-v5.0-pro/text-to-image \
-H "Authorization: Bearer $AIRCUBE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"prompt": "a cyberpunk cityscape at night, neon reflections on wet streets",
"aspect_ratio": "16:9",
"resolution": "2k"
}'Image to video (Python)
import os
import requests
API_KEY = os.environ["AIRCUBE_API_KEY"]
BASE = "https://aircube.ai/api/v3"
HEADERS = {
"Authorization": f"Bearer {API_KEY}",
"Content-Type": "application/json",
}
response = requests.post(
f"{BASE}/seedance-2.0/image-to-video",
headers=HEADERS,
json={
"prompt": "the camera slowly zooms in as petals fall",
"image": "https://example.com/flower.jpg",
"duration": "5",
"resolution": "720p",
},
)
generation_id = response.json()["data"]["id"]Text to speech (JavaScript)
const res = await fetch(
"https://aircube.ai/api/v3/qwen3-tts/text-to-speech",
{
method: "POST",
headers: {
Authorization: `Bearer ${process.env.AIRCUBE_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
prompt: "Welcome to AirCube. The unified API for AI generation.",
}),
}
);
const { data } = await res.json();
console.log(`Generation ID: ${data.id}`);What happens next
Use the generation id from the response to poll for results. Typical generation times:
| Type | Typical time |
|---|---|
| Image | 3–15 seconds |
| Video | 30 seconds – 5 minutes |
| Audio | 3–15 seconds |