Kling V3.0

Text-to-VideoImage-to-Video

Gain production-grade API access to Kuaishou's flagship audiovisual generation suite, Kling 3.0. Built on a unified, native multimodal (All-in-One) training architecture, Kling 3.0 seamlessly integrates text-to-video, image-to-video, reference conditioning, and in-video editing into a streamlined workflow. Breaking previous video length barriers, it generates continuous 3 to 15-second cinematic sequences with native intelligent multi-shot storyboarding.

All Models

Kling Video O1
OfficialText-to-Video

Kling Video O1

Kling Video O1 (Omni One) is Kuaishou's industry-first unified multimodal video model that merges video generation and editing into a single engine. Integrating text-to-video, image-to-video, element referencing, localized inpainting, and video restyling, it supports referencing up to 7 subjects simultaneously to lock character and prop consistency. Creators can perform conversational video editing via natural language prompts—ideal for advertising, VFX, and end-to-end film production.

Kling v3.0 Image-to-Video
OfficialImage-to-Video

Kling v3.0 Image-to-Video

Kling v3.0 Image-to-Video is Kuaishou's next-generation multimodal AI video model. Built on the unified Omni architecture, it takes static images or subject references to generate up to 15-second cinematic videos in up to 4K resolution. It features native audio-visual synchronization, multilingual lip-sync, and enhanced subject consistency to prevent visual drift, along with intelligent multi-shot control—ideal for commercial ads, film VFX, and narrative short videos.

Kling v3.0 Text-to-Video
OfficialText-to-Video

Kling v3.0 Text-to-Video

Kling v3.0 Text-to-Video is Kuaishou's next-generation AI video model. Powered by the native Omni architecture, it accurately parses complex prompt text to generate up to 4K cinematic-grade videos up to 15 seconds long. It natively supports integrated audio-video generation (ambient audio, music, and multilingual lip-sync) alongside exceptional visual realism, multi-shot coherence, and complex physical simulation—ideal for commercial advertising, film VFX, and content creation.

Kling 3.0 Omni
OfficialText-to-Video

Kling 3.0 Omni

Kling 3.0 Omni delivers high-quality text-to-video generation with smooth motion, cinematic visuals, accurate prompt adherence, and native audio for ready-to-share clips. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Kling V3.0 Models API Pricing Details

ModelPricing (USD)Our Pricing (USD)Discount
Kling Video O1$0.084/SECStart from$0.084/SEC
Kling v3.0 Image-to-Video$0.084/SECStart from$0.084/SEC
Kling v3.0 Text-to-Video$0.084/SECStart from$0.084/SEC
Kling 3.0 Omni$0.084/SECStart from$0.084/SEC