Kling v3.0 Image-to-Video
Kling v3.0 Image-to-Video is Kuaishou's next-generation multimodal AI video model. Built on the unified Omni architecture, it takes static images or subject references to generate up to 15-second cinematic videos in up to 4K resolution. It features native audio-visual synchronization, multilingual lip-sync, and enhanced subject consistency to prevent visual drift, along with intelligent multi-shot control—ideal for commercial ads, film VFX, and narrative short videos.
Kling v3.0 Image-to-Video
Kling v3.0 Image-to-Video is Kuaishou's next-generation multimodal AI video model. Built on the unified Omni architecture, it takes static images or subject references to generate up to 15-second cinematic videos in up to 4K resolution. It features native audio-visual synchronization, multilingual lip-sync, and enhanced subject consistency to prevent visual drift, along with intelligent multi-shot control—ideal for commercial ads, film VFX, and narrative short videos.
Base URL
https://api.icreat.aiAuthentication
All API requests must be authenticated with an API Key. You can obtain an API Key from the console.
export ICREAT_API_KEY="your-api-key-here"HTTP Request Headers
import os
API_KEY = os.environ.get("ICREAT_API_KEY")
headers = {
"Content-Type": "application/json",
"Authorization": "Bearer " + API_KEY,
}Protect your API Key
Never expose your API Key in client-side code or public repositories. Use environment variables or a backend proxy.
Code Examples
Image and video generation uses a three-step async flow: submit to get task_id, poll status, then fetch results when status is SUCCEEDED.
1. Submit Task
Send a generation request to the submit endpoint.
2. Poll Status
Poll with task_id. Response contains only status.
3. Get Result
Fetch output when the task succeeds.
Input Schema
Submit Task — Input
Total: 4 Required: 4 Optional: 0
URL of the input image used as the visual reference or initial frame.
The prompt used to generate the video. Chinese prompts must not exceed 2,000 characters, and English prompts must not exceed 2,000 words.
Output resolution mode. Supported values: std, pro, and 4k.
Video duration in seconds. Accepted range: 4–15.
Poll Status — Input
Total: 1 Required: 1 Optional: 0
Task ID from the submit endpoint.
Get Result — Input
Total: 1 Required: 1 Optional: 0
Task ID from the submit endpoint.
Output Schema
Submit Task — Output
Total: 1
Async task identifier.
Poll Status — Output
Total: 1
Current task status. Fetch results when SUCCEEDED.
Get Result — Output
Top-level array of resource objects.
Total: 3
Resource type; Video for this model.
Preview/access URL.
Download URL.
LLM Prompt
The Markdown below is an LLM-friendly prompt you can paste into AI assistants (e.g. Cursor, ChatGPT) to help them understand this model's API, call flow, and key parameters. Use Copy for AI or copy from the code block below.
# kwaivgi/kling-v3-0/image-to-video
> Kling v3.0 Image-to-Video is Kuaishou's next-generation multimodal AI video model.
## Overview
Use the iCreat three-step async task API: submit a generation request, poll status, then fetch the video URL on success.
## API Info
- **Base URL**:`https://api.icreat.ai`
- **Submit endpoint (POST)**:`/v1/task/submit/kwaivgi/kling-v3-0/image-to-video`
- **Poll endpoint (POST)**:`/v1/task/query-status`
- **Get result endpoint (POST)**:`/v1/task/get-result`
- **Model ID**:`kwaivgi/kling-v3-0/image-to-video`
- **Auth**:`Authorization: Bearer ${ICREAT_API_KEY}`
## Call Flow
1. **Submit**: POST submit path with body per Input Notes; response `{ "task_id": "..." }`
2. **Poll**: POST `/v1/task/query-status` with `{ "task_id": "..." }`; response contains **only** `{ "status": "SUBMITTED|SUCCEEDED|FAILED" }`
3. **Get result**: when `status` is `SUCCEEDED`, POST `/v1/task/get-result`; top-level array; read `url` or `download_url` when `type` is `Video`
### Input Notes
- Request body uses top-level flat fields
- Required: `image`, `prompt`, `mode`, `duration`
- `image` (required): URL of the input image used as the visual reference or initial frame.
- `prompt` (required): The prompt used to generate the video. Chinese prompts must not exceed 2,000 characters, and English prompts must not exceed 2,000 words.
- `mode` (required): Output resolution mode. Supported values: `std`, `pro`, and `4k`.
- `duration` (required): Video duration in seconds. Accepted range: `4`–`15`.
### Output Notes
- Poll: read `status` only
- Result: top-level `[{ "type": "Video", "url": "...", "download_url": "..." }]`
## Notes
- Use the same `task_id` across all three steps; do not skip polling
- `FAILED` is terminal — check request parameters or reference media