Minimax H3 Text To Video
MiniMax H3 (Official) generates 2K resolution video clips up to 15 seconds with native stereo audio. Supports text-to-video with multiple aspect ratios including 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16.
Key Features
- Omni-modal 2K video generation with native stereo audio — no separate TTS step needed.
- Up to 15-second video clips from text prompts with native stereo audio.
- Multiple aspect ratio support: 21:9, 16:9, 4:3, 1:1, 3:4, 9:16.
- Advanced prompt comprehension — up to 7,000 character prompts for detailed scene descriptions.
- 768P and 2K resolution output.
Parameters
| Parameter | Required | Description |
|---|---|---|
| prompt | Yes | Text description of the video to generate (max 7,000 characters). Describe both visual scenes and audio/sound effects. |
| duration | No | Video length in seconds: 4-15, integer only (default: 5). |
| aspect_ratio | No | Output ratio: 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 (default: 16:9). Cannot be 'adaptive' in text-to-video mode. |
| resolution | No | Output resolution: 768P or 2K (default: 2K). |
How to Use
- Write a detailed prompt describing the desired video scene, motion, and atmosphere.
- Select an aspect ratio and duration.
- Click Generate — the API returns a task ID for polling.
- Poll the status endpoint until the video is ready, then download the MP4 with audio.
Code Examples
import os
import requests
response = requests.post(
"https://aircube.ai/api/v3/minimax-h3/text-to-video",
headers={
"Authorization": "Bearer " + os.environ["AIRCUBE_API_KEY"],
"Content-Type": "application/json",
},
json={
"prompt": "A cinematic wide shot of a lighthouse on a rocky cliff at sunset, waves crashing below, seagulls calling overhead",
"duration": "5s",
"aspect_ratio": "16:9",
"resolution": "2K"
},
timeout=300,
)
data = response.json()
if data["success"]:
print("ID:", data["data"]["id"], "Status:", data["data"]["status"])
else:
print("Error:", data["error"]["message"])Pricing
| Resolution | Duration | Cost |
|---|---|---|
| 768p | 4s | $0.36 |
| 768p | 5s | $0.45 |
| 768p | 6s | $0.54 |
| 768p | 8s | $0.72 |
| 768p | 10s | $0.90 |
| 768p | 12s | $1.08 |
| 768p | 15s | $1.35 |
| 2k | 4s | $0.52 |
| 2k | 5s | $0.65 |
| 2k | 6s | $0.78 |
| 2k | 8s | $1.04 |
| 2k | 10s | $1.30 |
| 2k | 12s | $1.56 |
| 2k | 15s | $1.95 |
Billing rules
- 768p / 4s: $0.36.
- 2k / 4s: $0.52.
- Longer durations scale proportionally.
- Failed generations are not charged.
Best Use Cases
- Cinematic scene generation with synchronized audio and dialogue.
- Social media content with native sound (no separate audio workflow needed).
- Concept visualization and storyboarding for film and advertising.
- Music videos and promotional content with built-in stereo audio.
Pro Tips
- Use detailed, descriptive prompts for best results — H3 supports up to 7,000 characters.
- For text-to-video, always specify an aspect ratio (it cannot be 'adaptive').
- Start with shorter durations (4-5s) to iterate on prompts before generating longer clips.
- The model generates native stereo audio — describe sounds and dialogue in your prompt for better results.
Notes
- Video URLs are time-limited — download promptly after generation completes.
- Task results are retained for 7 days.


