curl-X POST 'https://pollo.ai/api/platform/v1/generation/google/veo-3-1/video'\-H'Content-Type: application/json'\-H'x-api-key: YOUR_API_KEY'\-d'{ "input": { "prompt": "A close-up of raindrops rippling across a still pond in slow motion, natural daylight", "aspectRatio": "16:9", "duration": 4, "resolution": "720p", "generateAudio": true }}'
Optionally include webhookUrl to receive a callback when the task succeeds or fails.
Optional request field
{"input":{"prompt":"A close-up of raindrops rippling across a still pond in slow motion, natural daylight","aspectRatio":"16:9","duration":4,"resolution":"720p","generateAudio":true},"webhookUrl":"https://example.com/webhooks/pollo"}
Schema
Input fields for the selected endpoint and mode.
Field
Type
Required
Description
prompt
string
Yes
Text prompt describing the content to generate.
aspectRatio
string
No
Output aspect ratio, as width:height (e.g. 16:9).
duration
number
No
The duration of the generated video, in seconds.
resolution
string
No
Output resolution of the generated media (e.g. 720p, 1080p, 2K, 4K). Values are case-insensitive.
seed
integer
No
Random seed for reproducible generation.
generateAudio
boolean
No
Generate natural-sounding, fitting audio for the output video.
Google Veo 3.1 brings cinematic video generation, stronger prompt following, reference-based visual control, native audio, and first-to-last frame continuity into one advanced model.
With Pollo API, you can integrate Veo 3.1 into apps for product videos, ads, storytelling, social content, and creative production workflows without managing separate model access or fragmented infrastructure.
Key Features
Cinematic Video Quality
Generate polished videos with realistic motion, rich lighting, and strong composition, ideal for ads, product clips, branded content, and cinematic storytelling.
First and Last Frame Control
Define the opening and ending frames to create smoother transitions, product reveals, transformations, and story moments with clearer visual direction.
Reference-Guided Generation
Use reference images to keep characters, products, objects, and visual styles more consistent across generated videos.
Native Audio Generation
Generate synchronized dialogue, ambience, music, and sound effects directly within the video generation workflow.
Scene Extension
Extend existing clips into longer sequences while preserving scene continuity, visual style, and narrative flow.
Use Cases
Product Showcase Videos: Turn product photos or descriptions into polished clips for stores, ads, and campaign pages.
Cinematic Ad Creatives: Generate high-quality visual ads with stronger control over scene style, motion, and audio.
Social Media Clips: Create reels, shorts, and branded posts that feel more complete with synchronized sound and natural movement.
Brand Launch Teasers: Build short cinematic videos for product launches, event announcements, and campaign reveals.
Character-Based Stories: Use reference images to keep characters, objects, or visual styles more consistent across generated clips.
Extended Multi-Scene Videos: Continue existing clips into longer sequences while keeping the scene direction more coherent.
How to Use Google Veo 3.1 API
Get an API Key: Create your Pollo API key from your account or developer dashboard.
Set Veo 3.1 as the Model: Choose Google Veo 3.1 in your video generation request.
Add Your Inputs: Send a prompt, reference images, first/last frames, or a source video depending on your workflow.
Generate and Retrieve: Submit the request, check the task status or webhook, then download the final video
Prompting Best Practices for Google Veo 3.1 API
Think of your Veo 3.1 prompt as direction for a video shoot. The clearer you are about what should appear, how it should move, how it should be filmed, and what it should sound like, the easier it is to get a polished result.
A Simple Prompt Formula
Scene + Subject + Action + Camera + Lighting + Style + Audio
What You Should Notice
Use clear camera language: Mention close-up, wide shot, drone shot, slow push-in, handheld follow, or tracking shot.
Describe the main action: Keep the scene focused on one strong movement or event.
Add lighting and mood: Words like golden hour, neon-lit, soft studio light, dramatic, playful, or suspenseful help shape the result.
Specify audio: Include ambience, dialogue, music, sound effects, or silence instead of letting the model guess.
Use references when needed: Add reference images for products, characters, or brand visuals when visual accuracy matters.
Example Prompts
Product Launch Video
“A cinematic product launch video of a matte black wireless headphone floating above a reflective glass surface. Soft studio lighting, slow rotating camera movement, subtle blue highlights, premium tech commercial style. Add deep ambient music and a soft electronic whoosh as the product turns toward the camera.”
Fashion Brand Teaser
“A stylish model wearing a beige trench coat walks through a rainy city street at night. Neon signs reflect on the wet pavement, the camera follows from a low angle, slow-motion movement, luxury fashion campaign mood. Include soft rain ambience and quiet cinematic music.”
Character Story Scene
“A young astronaut stands alone on the surface of a red desert planet, looking at a distant glowing city on the horizon. Wide cinematic shot, dust moving in the wind, dramatic sunset lighting, slow camera push-in. Add low atmospheric wind and emotional orchestral music.”
First-to-Last Frame Transition
“Start with a closed luxury perfume box on a marble table. End with the bottle fully revealed, standing upright with golden light reflecting through the glass. Generate a smooth cinematic reveal between the two moments, elegant camera movement, premium beauty ad style, soft sparkling sound effects.”
Veo 3.1 vs Sora 2 vs Kling 3.0
Capability
Veo 3.1
Sora 2
Kling 3.0
Native Audio
✅ Built-in audio generation
✅ Built-in audio generation
—
Reference Images
✅
✅
✅
First/Last Frame Control
✅
—
✅
Scene Extension
✅
✅
✅
Text-to-Video
✅
✅
✅
Image-to-Video
✅
✅
✅
Recommended For
Product marketing, ads, storytelling
Creative scene generation and realism
Dynamic motion and character animation
Why Choose Veo 3.1 API?
Veo 3.1 is a strong choice for projects that require both video quality and controllability. Support for reference images, first/last frame conditioning, native audio generation, and scene extension makes it suitable for advertising, product marketing, storytelling, and content production workflows.
Through Pollo API, Veo 3.1 can be accessed alongside other leading video models through a single API, simplifying evaluation, integration, and deployment.
Google Veo 3.1 API FAQs
What is Google Veo 3.1?
Google Veo 3.1 is an advanced video generation model built for cinematic, controllable outputs, combining prompt following, reference guidance, native audio, frame control, and scene extension.
Can I generate videos from both text and images?
Yes. Veo 3.1 supports text-to-video and image-to-video workflows, making it useful for building flexible tools that turn ideas, assets, or product visuals into finished clips.
What inputs does Veo 3.1 support?
Veo 3.1 can use reference images to better preserve product details, character appearance, visual style, and scene identity, especially when consistency matters more than random variation.
Can I control how the video starts and ends?
Yes. First and last frame control helps you design more intentional transitions, reveals, transformations, and story beats instead of relying on the model to choose the full motion path.
Does Veo 3.1 generate sound with the video?
Yes. Veo 3.1 can create synchronized audio, including dialogue, ambience, music, and sound effects, helping the output feel closer to a complete video asset.
Is Veo 3.1 suitable for commercial video content?
Yes. Its cinematic quality, audio support, and stronger visual control make it well suited for ads, product showcases, brand teasers, social clips, and campaign-ready creative workflows.
Why run Veo 3.1 through Pollo API?
Pollo API lets you access Veo 3.1 alongside other leading models, making it easier to compare outputs, scale generation, and choose the best model per workflow.