curl-X POST 'https://pollo.ai/api/platform/v1/generation/kling-ai/kling-v3-omni/video'\-H'Content-Type: application/json'\-H'x-api-key: YOUR_API_KEY'\-d'{ "input": { "prompt": "A cinematic shot of a golden retriever running through a field of sunflowers at sunset, warm rim light, shallow depth of field", "aspectRatio": "16:9", "duration": 5, "mode": "pro", "generateAudio": true }}'
{"input":{"prompt":"A cinematic shot of a golden retriever running through a field of sunflowers at sunset, warm rim light, shallow depth of field","aspectRatio":"16:9","duration":5,"mode":"pro","generateAudio":true},"webhookUrl":"https://example.com/webhooks/pollo"}
模式
輸入所選端點和模式的欄位。
場地
類型
必需的
描述
prompt
string
支持
Text prompt describing the content to generate.
aspectRatio
string
不
Output aspect ratio, as width:height (e.g. 16:9).
duration
integer
不
The duration of the generated video, in seconds.
mode
string
不
Generation mode of the model (e.g. a faster or higher-quality tier).
generateAudio
boolean
不
Generate natural-sounding, fitting audio for the output video.
Kling 3.0 Omni generates video from multiple input types including text, images, subject sets, source video, and audio in a single request up to 4K.
With Kling 3.0 Omni on Pollo API, developers can drive reference-led video generation from a mix of guidance types, choosing clip length, aspect ratio, and optional matching audio per request.
Key Features of Kling 3.0 Omni API
Omni-Modal Reference Input
The model accepts up to seven references at once, combining an image, a source video, a subject set, and an audio cue to steer a single generation.
Subject Consistency
Provide multiple images of a character or object. The model maintains subject identity across the shot as the scene moves.
Audio-Guided Motion
Supply an audio reference and the model times movement, rhythm, and pacing to the sound, aligning action with beats, speech, or ambient cues in the clip.
Environmental Simulation
Spatial simulation renders single-shot scenes with coherent light, depth, and environmental reaction as camera and subjects move through space.
4K Volumetric Output
Generations reach up to 4K, holding fine texture and depth cues at high resolution for broadcast and large-screen use cases.
Flexible Duration And Framing
Clips run from 3 to 15 seconds in landscape, portrait, or square, letting one integration serve social, product, and full-frame delivery from the same model.
Use Cases of Kling 3.0 Omni API
Recurring Character Content: Reuse a subject set to keep a mascot, host, or avatar consistent across episodic clips without redrawing identity each time.
Audio-Synced Performance Clips: Drive a dancer, musician, or speaker so movement locks to an uploaded track, useful for music and lip-motion tools.
Video-Referenced Restyling: Feed an existing clip as reference and regenerate it with new direction while keeping the original motion and staging intact.
Scene Continuation Pipelines: Combine image and video references to extend a shot forward, keeping environment and subject continuity across a longer sequence.
Multi-Reference Storyboarding: Blend several images and a subject into one prompt to preview complex scenes before committing to a full production pass.
High-Resolution Broadcast Assets: Generate 4K single-shot scenes for signage, streaming intros, or large displays where resolution and depth matter.
Localized Ad Variants: Reuse one subject and swap audio references to produce region-specific versions with matched pacing and consistent branding.
Immersive Environment Demos: Show a space filling with fog, water, or crowd motion where surroundings respond believably for real-estate or event previews.
How to Use Kling 3.0 Omni API
Get an API Key: Create a Pollo API account and generate your API key from the developer dashboard.
Choose the Model: Send requests to the Kling 3.0 Omni endpoint at /generation/kling-ai/kling-v3-omni/ref2video.
Add Your Inputs: Provide a prompt and one to seven refs (image, subject, video, or audio), then set duration (3-15), aspectRatio, and resolution (720P, 1080P, or 4K).
Generate and Retrieve: Submit the task, poll the returned task status, and download the finished video, optionally with generateAudio producing a matching soundtrack.
Prompting Best Practices for Kling 3.0 Omni API
Think of the prompt as directing references rather than describing from scratch. Name each reference's role, the action you want, and how the scene should behave around your subject.
A Simple Prompt Formula
Subject and its reference role + specific action + scene and environment behavior + camera movement + audio or pacing cue
What You Should Notice
Assign Reference Roles: Tell the model which reference is the subject, which is the setting, and which drives audio, so guidance does not compete.
Keep Subjects Clear: Provide clean, varied images of one character so identity holds steady across the generated shot.
Describe Environment Response: State how light, weather, or surroundings should shift, since the model simulates space rather than a flat backdrop.
Match Audio To Action: When using an audio reference, name the movement it should sync with so pacing follows the sound.
Pick Resolution By Purpose: Choose 4K for broadcast and large screens, lower settings for faster drafts and iteration.
Hold One Scene: Keep each generation to a single coherent moment so the clip stays continuous end to end.
Example Prompts
Consistent Mascot Episode
"Use the subject reference of an illustrated red fox in a scout uniform. The fox walks through a sunlit library, pauses to pull a glowing book from a high shelf, and grins. Slow dolly-in, warm dust-filled light."
Audio-Synced Street Dancer
"Animate the subject dancer to the uploaded drum track. Sharp footwork and arm hits land on each beat inside a concrete underpass at dusk, cool blue shadows, handheld camera circling as bass pulses build."
Environment Simulation Reveal
"A stone courtyard at dawn slowly fills with rolling ground fog that curls around old columns and catches first light. Camera cranes up to reveal the full scene, quiet ambient tones, calm realistic style."
Video-Referenced Restyle
"Take the reference clip of a cyclist on a coastal road and regenerate it as a hand-painted look. Keep the same route and pedaling motion, add windblown grass and passing gulls, soft afternoon light."
Kling 3.0 Omni suits products that build from assets, letting you fuse subjects, source clips, and audio into consistent, high-resolution scenes rather than prompting each shot blind.
Through Pollo API, you reach Kling 3.0 Omni with one API key alongside 300+ leading video and image models, with clean docs, task polling and webhooks.
Start with the Kling 3.0 Omni endpoint on Pollo API to ship reference-driven, 4K-ready video features, then compare it against Veo, Sora, and other Kling models without touching your integration.
Kling 3.0 Omni API FAQs
What is Kling 3.0 Omni?
Kling 3.0 Omni is a Kling AI video model that generates from combined text, image, subject, video, and audio references, producing consistent single-shot scenes up to 4K.
What inputs can Kling 3.0 Omni API combine?
It accepts up to seven references in one request, mixing an image, a source video, a subject image set, and an audio cue alongside your text prompt.
How does subject consistency work?
You supply a small set of images of one character or object, and the model keeps that subject recognizable throughout the generated clip as the scene moves.
What resolutions and lengths are supported?
Kling 3.0 Omni API outputs 720P, 1080P, or 4K, with clip durations selectable from 3 to 15 seconds and landscape, portrait, or square framing.
Can it generate matching audio?
Yes. It can produce natural-sounding audio for the output video, and an audio reference can also guide the timing and rhythm of on-screen motion.
Why run Kling 3.0 Omni through Pollo API?
Pollo API gives one integration for Kling 3.0 Omni and 300+ other models, with documentation, task tracking and webhooks.