Kling V3.0
Gain production-grade API access to Kuaishou's flagship audiovisual generation suite, Kling 3.0. Built on a unified, native multimodal (All-in-One) training architecture, Kling 3.0 seamlessly integrates text-to-video, image-to-video, reference conditioning, and in-video editing into a streamlined workflow. Breaking previous video length barriers, it generates continuous 3 to 15-second cinematic sequences with native intelligent multi-shot storyboarding.
All Models

Kling Video O1
Kling Video O1 (Omni One) is Kuaishou's industry-first unified multimodal video model that merges video generation and editing into a single engine. Integrating text-to-video, image-to-video, element referencing, localized inpainting, and video restyling, it supports referencing up to 7 subjects simultaneously to lock character and prop consistency. Creators can perform conversational video editing via natural language prompts—ideal for advertising, VFX, and end-to-end film production.

Kling v3.0 Image-to-Video
Kling v3.0 Image-to-Video is Kuaishou's next-generation multimodal AI video model. Built on the unified Omni architecture, it takes static images or subject references to generate up to 15-second cinematic videos in up to 4K resolution. It features native audio-visual synchronization, multilingual lip-sync, and enhanced subject consistency to prevent visual drift, along with intelligent multi-shot control—ideal for commercial ads, film VFX, and narrative short videos.

Kling v3.0 Text-to-Video
Kling v3.0 Text-to-Video is Kuaishou's next-generation AI video model. Powered by the native Omni architecture, it accurately parses complex prompt text to generate up to 4K cinematic-grade videos up to 15 seconds long. It natively supports integrated audio-video generation (ambient audio, music, and multilingual lip-sync) alongside exceptional visual realism, multi-shot coherence, and complex physical simulation—ideal for commercial advertising, film VFX, and content creation.

Kling 3.0 Omni
Kling 3.0 Omni delivers high-quality text-to-video generation with smooth motion, cinematic visuals, accurate prompt adherence, and native audio for ready-to-share clips. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
Kling V3.0 Models API Pricing Details
| Model | Pricing (USD) | Our Pricing (USD) | Discount | |
|---|---|---|---|---|
| Kling Video O1 | $0.084/SEC | Start from$0.084/SEC | — | |
| Kling v3.0 Image-to-Video | $0.084/SEC | Start from$0.084/SEC | — | |
| Kling v3.0 Text-to-Video | $0.084/SEC | Start from$0.084/SEC | — | |
| Kling 3.0 Omni | $0.084/SEC | Start from$0.084/SEC | — |