curl-X POST 'https://pollo.ai/api/platform/v1/generation/kling-ai/kling-v2-6/video'\-H'Content-Type: application/json'\-H'x-api-key: YOUR_API_KEY'\-d'{ "input": { "prompt": "A cinematic shot of a golden retriever running through a field of sunflowers at sunset, warm rim light, shallow depth of field", "aspectRatio": "16:9", "duration": 5, "generateAudio": true }}'
Optionally include webhookUrl to receive a callback when the task succeeds or fails.
Optional request field
{"input":{"prompt":"A cinematic shot of a golden retriever running through a field of sunflowers at sunset, warm rim light, shallow depth of field","aspectRatio":"16:9","duration":5,"generateAudio":true},"webhookUrl":"https://example.com/webhooks/pollo"}
Schema
Input fields for the selected endpoint and mode.
Field
Type
Required
Description
prompt
string
Yes
Text prompt describing the content to generate.
aspectRatio
string
No
Output aspect ratio, as width:height (e.g. 16:9).
duration
number
No
The duration of the generated video, in seconds.
generateAudio
boolean
No
Generate natural-sounding, fitting audio for the output video.
Kling 2.6 generates video with synchronized audio. Dialogue, tone, and sound effects are produced with the footage in a single request.
With Kling 2.6 on Pollo API, developers can add text-to-video and image-to-video generation to their products, guide start and end frames, choose 5 or 10 second clips, and toggle matching audio per call.
Key Features of Kling 2.6 API
Synchronized Audio and Visuals
Speech, ambient tone, and sound effects are produced synchronized with the action, with lips and gestures timed to the audio.
Directed Sound Control
Prompts can steer dialogue lines, emotional tone, and specific effects, letting a talking character or an impact moment carry the exact sound you describe rather than generic filler.
Start and Tail Frame Guidance
Supplying a first image and a tail image lets the model interpolate a controlled path between them, so transitions and product reveals resolve on the frame you intend.
Text and Image Guidance
Generation starts from a written prompt or an image URL, so idea-first scripts and asset-first shots both become moving clips within the same model.
Flexible Framing and Length
Clips render in 16:9, 9:16, or 1:1 for landscape, vertical, and square feeds, at 5 or 10 seconds, matching short-form placements without post-crop reframing.
Use Cases of Kling 2.6 API
Talking Character Clips: Generate a subject who speaks a scripted line with matching mouth movement and tone, useful for avatar and dialogue-driven apps.
Sound-Effect Scenes: Produce moments where the audio is the point, such as a bell ringing, rain on a tin roof, or footsteps on gravel.
Vertical Social Videos: Render 9:16 clips with built-in audio for feeds where sound-on autoplay decides whether a viewer stays.
Guided Transitions: Use a first and tail image to control how a scene starts and ends, ideal for reveals, before-and-after shots, and loops.
Image-Led Animation: Animate a still photo into a short moving shot with fitting ambience, preserving the original subject and framing.
Narrated Explainers: Turn a written script into a spoken short where voice and on-screen action stay aligned for tutorials and product walkthroughs.
Ambient Mood Pieces: Create atmospheric clips where wind, water, or crowd sound sets the tone alongside the visuals.
How to Use Kling 2.6 API
Get an API Key: Create a Pollo API account and generate your API key from the developer dashboard.
Choose the Model: Send requests to the Kling 2.6 endpoint at /generation/kling-ai/kling-v2-6.
Add Your Inputs: Provide a prompt for text-to-video, or an image URL for image-to-video, optionally an imageTail, then set length (5 or 10), aspectRatio, and generateAudio.
Generate and Retrieve: Submit the task, poll the returned task status, and download the finished clip once processing succeeds.
Prompting Best Practices for Kling 2.6 API
Write for both channels at once. Describe what happens on screen, then name the spoken line, tone, or effect you want, so the model has a clear target for audio and motion together.
A Simple Prompt Formula
Subject and setting + specific action + spoken line or sound cue + tone or mood + camera and framing
What You Should Notice
Name the Sound: State the dialogue, effect, or ambience directly so the generated audio matches your intent instead of defaulting.
Keep Speech Short: Fit spoken lines to 5 or 10 seconds so delivery and lip movement stay natural within the clip length.
Set the Two Frames: When using a tail image, make the start and end states clearly different so the interpolation has a direction to follow.
Pick Framing Early: Choose 16:9, 9:16, or 1:1 up front so composition suits the target feed from the first render.
One Audio Focus: Center each clip on a single voice or effect so the soundtrack reads clearly rather than competing.
Example Prompts
Scripted Dialogue
"A librarian looks up from a book and says warmly, 'You made it just in time.' Soft page-turning and a distant clock tick in the background. Medium shot, quiet indoor light, gentle push-in, realistic style."
Sound-Driven Moment
"A blacksmith strikes glowing metal on an anvil, sparks scattering with each rhythmic clang. The hammer rings sharp and metallic. Side-on shot, dim forge glow, slight camera shake, gritty realistic style."
Guided Frame Transition
"Start on a folded paper crane on a windowsill; end with it mid-flight across a bright room. Soft flutter of paper and a light breeze. Static camera, morning light, clean realistic style."
Ambient Vertical Clip
"A street musician plays a slow melody on an accordion at dusk, coins clinking as a passerby drops one in. Vertical 9:16 framing, warm streetlight glow, camera slowly rising, natural realistic style."
Kling 2.6 vs Veo 3.1 vs Sora 2
Capability
Kling 2.6
Veo 3.1
Sora 2
Synchronized audio and video
✅ Dialogue, tone, and effects
✅ Native audio
✅ Native audio
Directed sound control
✅ Prompt-steered lines and effects
✅
✅
Start and tail frame guidance
✅ First and last image
✅
❌
Text and image inputs
✅
✅
✅
Clip length
5 or 10 seconds
Short directed clips
Short directed clips
Framing options
16:9, 9:16, 1:1
Landscape and vertical
Landscape and vertical
Recommended For
Audio-synced clips with frame control
Prompt-faithful audio shots
Complex multi-shot scenes
Why Choose Kling 2.6 API?
Kling 2.6 fits products that need voice and picture to land together, from talking characters and narrated explainers to sound-led moments where the audio carries the scene.
Through Pollo API, you reach Kling 2.6 with one API key alongside 300+ leading video and image models, with clean docs and task status polling.
Start with the Kling 2.6 endpoint on Pollo API to ship audio-synced video features, then compare it against Veo, Sora, and other models without touching your integration.
Kling 2.6 API FAQs
What is Kling 2.6?
Kling 2.6 is a Kling AI video generation model with synchronized audio and visuals, supporting text-to-video and image-to-video, plus prompt control over dialogue, tone, and sound effects.
Does Kling 2.6 API generate audio?
Yes. You can enable matching audio per request, and Kling 2.6 produces dialogue, ambience, and effects timed to the on-screen action rather than added afterward.
What inputs does Kling 2.6 API support?
It accepts a text prompt or an image URL, and you can add a tail image to guide the ending frame for controlled transitions and reveals.
What clip lengths and framing are available?
Kling 2.6 renders 5 or 10 second clips in 16:9, 9:16, or 1:1, covering landscape, vertical, and square placements per request.
How does the tail image work?
Providing a first image and a tail image lets Kling 2.6 interpolate between them, so a clip begins and ends on the frames you choose.
Why run Kling 2.6 through Pollo API?
Pollo API gives one integration for Kling 2.6 and 300+ other models, with documentation and task tracking.