curl-X POST 'https://pollo.ai/api/platform/v1/generation/xai/grok-imagine-video-1-5/video'\-H'Content-Type: application/json'\-H'x-api-key: YOUR_API_KEY'\-d'{ "input": { "prompt": "A cinematic shot of a golden retriever running through a field of sunflowers at sunset, warm rim light, shallow depth of field", "image": "https://example.com/first-frame.jpg", "duration": 5, "resolution": "480p" }}'
Optionally include webhookUrl to receive a callback when the task succeeds or fails.
Optional request field
{"input":{"prompt":"A cinematic shot of a golden retriever running through a field of sunflowers at sunset, warm rim light, shallow depth of field","image":"https://example.com/first-frame.jpg","duration":5,"resolution":"480p"},"webhookUrl":"https://example.com/webhooks/pollo"}
Schema
Input fields for the selected endpoint and mode.
Field
Type
Required
Description
prompt
string
Yes
Text prompt describing the content to generate.
image
string
Yes
First-frame / reference image URL (HTTP(S).
duration
integer
No
The duration of the generated video, in seconds.
resolution
string
No
Output resolution of the generated media (e.g. 720p, 1080p, 2K, 4K). Values are case-insensitive.
Grok Imagine 1.5 supports image-to-video generation from a single source image, adding subject motion, object interactions, and synchronized audio while preserving the source subject’s appearance and visual style.
Through Pollo API, you can integrate Grok Imagine 1.5 into image-animation applications, with configurable clip duration, aspect ratio, and output resolution for each request.
Key Features of Grok Imagine 1.5 API
Image-Driven Motion
Generate video from a source image and text prompt while using the image’s subject, composition, framing, and visual style as generation guidance.
Synchronized Audio
Generate synchronized dialogue, ambience, and sound effects alongside the video, aligning audio cues with subject movement and on-screen events.
Object Interaction
Support interactions between subjects, objects, liquids, and environmental elements, including gripping, splashing, contact, and collision-based motion.
Subject Preservation
Maintain recognizable facial features, products, objects, and scene details from the source image throughout the generated motion.
Flexible Clip Length
Support configurable video duration for short loops, individual actions, and longer image-to-video sequences.
Use Cases of Grok Imagine 1.5 API
Still Photo Animation: Turn an archived portrait or product image into a short moving clip while preserving its primary subject, composition, and visual direction.
Sound-Backed Social Clips: Generate vertical videos with synchronized audio for short-form feeds, stories, and mobile content workflows.
Ecommerce Motion Shots: Animate catalog images to show product movement, fabric motion, opening actions, or powered-on states for product demonstrations and listings.
Interactive Scene Beats: Produce short scenes where subjects and objects interact, such as a ball dropping into water or a match striking a surface.
Music and Instrument Clips: Animate performer images with coordinated hand movement and generated instrument audio for music and creator applications.
Food and Sizzle Reels: Animate food or cooking images with steam, movement, sizzling effects, and ambient kitchen audio.
Pet and Character Loops: Add subtle motion and generated audio to pet portraits or illustrated characters for short looping clips.
How to Use Grok Imagine 1.5 API
Get an API Key: Create a Pollo API account and generate your API key from the developer dashboard.
Choose the Model: Send your request to the Grok Imagine 1.5 endpoint at /generation/grok/grok-imagine-video-1.5.
Add Your Inputs: Provide an image URL and a prompt, then set length (up to 15), resolution (480p, 720p, or 1080p), and aspectRatio as needed.
Generate and Retrieve: Submit the task, poll the returned task status, and download the finished video with its audio once processing succeeds.
Prompting Best Practices for Grok Imagine 1.5 API
Since every generation starts from a source image, focus the prompt on the intended motion, object interactions, camera behavior, and synchronized audio.
A Simple Prompt Formula
Subject in the image + the motion it performs + how objects interact + the sound the scene should make + pace or mood
What You Should Notice
Anchor To The Image: Describe motion for the subject already shown in the image and avoid adding unrelated visual elements.
Name The Sound: State the intended sound, such as splashing, sizzling, or a plucked note, to guide the synchronized audio output.
Direct Interactions: Specify which subjects or objects should touch, collide, or react to guide contact and interaction within the scene.
Keep One Action: Describe one clear movement per clip to guide motion and scene progression across different durations.
Match Framing To Source: Select an aspect ratio close to the source image to preserve more of its original composition and framing.
Example Prompts
Street Food Sizzle
"A close-up of skewers on a charcoal grill from the photo start cooking, fat drips and flares, smoke curls upward, and the sizzle builds as a hand turns one skewer. Warm evening light, unhurried pace."
Poured Drink Interaction
"The bottle in the image tips and pours amber liquid into a glass, ice cracks and shifts as it fills, foam settles at the rim, with a clear pouring and fizzing sound. Steady tabletop framing."
Acoustic Moment
"The seated guitarist from the photo begins fingerpicking, fingers moving across the strings while the body of the guitar resonates, a soft melodic sound matching each motion. Gentle indoor light, calm and intimate mood."
Rainy Window
"Raindrops on the window in the still start sliding and merging down the glass, the blurred street lights behind them flicker, with steady rainfall and distant traffic sound. Quiet, contemplative pace, no camera movement."
Grok Imagine 1.5 vs Kling 2.6 vs Hailuo 2.3
Capability
Grok Imagine 1.5
Kling 2.6
Hailuo 2.3
Primary input
✅ Image-to-video with prompt
✅ Text and image-to-video
✅ Text and image-to-video
Synchronized audio
✅ Generated with the clip
❌ Not native
❌ Not native
Object interaction
✅ Strong reactive contact
✅ Good motion
✅ Physics-focused motion
Subject preservation
✅ Faithful to source image
✅
✅
Clip length
Up to 15 seconds
Short directed clips
6 or 10 seconds
Resolution options
480p, 720p, 1080p
Up to 1080p
768P and 1080P
Recommended For
Animating stills with matching sound
Text-first video shots
Physics-accurate motion
Why Choose Grok Imagine 1.5 API?
When an application needs to animate source images with coordinated motion and sound, Grok Imagine 1.5 supports subject preservation, object interaction, and synchronized audio within a single generated clip.
It can support product motion, portrait animation, social videos, music content, and object-interaction scenes.
Pollo API offers you access to Grok Imagine 1.5 and other supported models through the same request and task-management workflow, making it easier to compare outputs, scale generation, and select models within a unified integration.
Grok Imagine 1.5 API FAQs
What is Grok Imagine 1.5?
Grok Imagine 1.5 is a Grok image-to-video model that animates a still image with realistic motion, object interaction, and synchronized audio while preserving the subject's look.
Does Grok Imagine 1.5 API support text-only generation?
No. Each request needs an image plus a prompt, so it animates an existing still rather than generating video from text alone.
Does it generate audio?
Yes. Clips come with audio synced to the on-screen action, so poured liquids, instruments, and other motion carry matching sound.
What clip lengths and resolutions are available?
You can set clip length up to 15 seconds and choose 480p, 720p, or 1080p output, with several aspect ratios per request.
How well does it keep the original subject?
Subject preservation is a core strength, so faces, products, and scene details from your source image stay recognizable through the generated motion.
Why run Grok Imagine 1.5 through Pollo API?
Pollo API lets you access Grok Imagine 1.5 alongside other leading models, making it easier to compare outputs, scale generation, and select the right model for each workflow.