Alibaba Cloud’s flagship multimodal video foundation model, Wanx 2.1 excels at generating cinematic videos with realistic physics, complex motion handling, and bilingual text effects.
curl-X POST 'https://pollo.ai/api/platform/v1/generation/alibaba/wan-v2-1/video'\-H'Content-Type: application/json'\-H'x-api-key: YOUR_API_KEY'\-d'{ "input": { "prompt": "A cinematic shot of a golden retriever running through a field of sunflowers at sunset, warm rim light, shallow depth of field", "mode": "fast", "duration": 5, "aspectRatio": "16:9" }}'
Optionally include webhookUrl to receive a callback when the task succeeds or fails.
Optional request field
{"input":{"prompt":"A cinematic shot of a golden retriever running through a field of sunflowers at sunset, warm rim light, shallow depth of field","mode":"fast","duration":5,"aspectRatio":"16:9"},"webhookUrl":"https://example.com/webhooks/pollo"}
Schema
Input fields for the selected endpoint and mode.
Field
Type
Required
Description
prompt
string
Yes
Text prompt describing the content to generate.
mode
string
No
Generation mode of the model (e.g. a faster or higher-quality tier).
Wan 2.1 is an advanced AI model developed by Alibaba's Wan team for generating high-quality videos from text descriptions or images.
With Wan 2.1 on Pollo API, developers add text-to-video and image-to-video to their apps, picking a fast draft mode or a higher-fidelity pro mode for five-second clips across several aspect ratios.
Key Features of Wan 2.1 API
Bilingual Text Effects
Wan 2.1 draws Chinese and English characters directly into a shot, so titles, labels, and captions hold their shape instead of dissolving into noise as frames advance.
Realistic Physics
Motion respects gravity and momentum, so falling objects, flowing fabric, and shifting weight behave believably, keeping generated action grounded rather than floaty or rubbery under scrutiny.
Complex Motion Handling
Multi-part movement stays coherent, from a crowd crossing a street to limbs and tools moving at once, without the subject smearing or losing structure mid-clip.
Text and Image Guidance
Start from a written prompt or animate an existing still, letting Wan 2.1 serve both idea-first concepts and asset-first workflows built around a fixed opening frame.
Fast and Pro Modes
Choose a quick pass for rapid iteration or a refined pass when detail matters, so drafting and final renders can share one model without switching integrations.
Use Cases of Wan 2.1 API
Localized Video Captions: Generate clips with on-screen Chinese or English text for bilingual campaigns, subtitled explainers, and region-specific storefront content.
Signage and Storefront Scenes: Produce shots where shop signs, menus, or posters carry real readable words for retail and hospitality mockups.
Physics-Driven Action: Animate pouring liquids, tumbling props, or fabric in motion for demos where believable weight sells the moment.
Image Reanimation: Feed a still frame and set it moving while preserving the original composition, useful for photo-to-clip and reference-led tools.
Title Card Generation: Create short intro or outro sequences where kinetic text and motion appear together in a single generated shot.
Crowd and Multi-Subject Shots: Render busy scenes with several moving elements that stay distinct, suited to establishing shots and background plates.
Rapid Concept Drafting: Use fast mode to preview a described scene quickly before committing compute to a refined pro render.
Vertical Social Clips: Output portrait-format short video for feeds and stories without cropping or reframing a landscape source.
How to Use Wan 2.1 API
Get an API Key: Create a Pollo API account and generate your API key from the developer dashboard.
Choose the Model: Send requests to the Wan 2.1 endpoint at /generation/wanx/wanx-v2-1.
Add Your Inputs: Provide a prompt for text-to-video plus an aspectRatio (16:9, 9:16, 4:3, 3:4, or 1:1), or pass an image URL for image-to-video, then set mode to fast or pro. Clip length is 5 seconds.
Generate and Retrieve: Submit the task, poll the returned task status, and download the finished video once processing succeeds.
Prompting Best Practices for Wan 2.1 API
Treat the prompt as a scene brief. Name the subject, the motion, any on-screen text with its exact wording, the framing, and the setting so Wan 2.1 has clear targets to render.
A Simple Prompt Formula
Subject + specific motion + exact on-screen text (if any) + camera framing + setting and lighting
What You Should Notice
Quote Your Text Exactly: Put the words you want on screen in the prompt, since Wan 2.1 renders the characters it is given in Chinese or English.
Describe Physical Motion: State how things move, such as tips, pours, or sways, so the physics engine has a concrete goal.
Keep One Focused Action: A single clear movement holds together better across the five-second clip than several competing events.
Set The Aspect Ratio Early: Match framing to the destination, choosing vertical for social and wide for landscape scenes.
Use Mode By Intent: Draft in fast mode, then re-render the keeper in pro mode when you need cleaner detail.
Example Prompts
Bilingual Shop Sign
"A small noodle shop at dusk, steam rising from bowls, a glowing sign above the door reads 'Open' in English and Chinese. Warm lantern light, slow push-in, people passing in soft focus, realistic style."
Physics-Led Kitchen Shot
"Honey pours slowly from a jar onto a stack of pancakes, the stream twisting and pooling with visible thickness. Top-down camera, soft morning window light, clean realistic style focused on believable flow."
Animated Title Card
"Bold English text 'CHAPTER ONE' forms letter by letter over a windswept grassland, blades of grass bending in the breeze. Steady wide shot, overcast light, minimal graphic style with subtle motion."
Street Crowd Scene
"A busy market street where several vendors gesture and shoppers weave between stalls, fruit and fabric in view. Handheld medium shot, bright midday light, natural realistic style with many small moving details."
Wan 2.1 vs Kling 1.6 vs Hunyuan Video
Capability
Wan 2.1
Kling 1.6
Hunyuan Video
Bilingual on-screen text
✅ Chinese and English rendering
❌ Limited text fidelity
❌ Limited text fidelity
Physics realism
✅ Strong weight and motion
✅ Good motion control
✅ Smooth realistic motion
Complex multi-subject motion
✅
✅
✅
Input types
Text and image
Text and image
Text and image
Clip length
5 seconds
Short directed clips
Short clips
Recommended For
Text-in-video and physical motion
Directed character motion
Open-model realistic video
Why Choose Wan 2.1 API?
Wan 2.1 stands out when a shot must carry real readable words and believable motion at once, from bilingual captions and signage to physics-driven action in a single clip.
Through Pollo API, you reach Wan 2.1 with one API key alongside 300+ leading video and image models, with clean docs and task status polling.
Start with the Wan 2.1 endpoint on Pollo API to ship bilingual, physics-aware video features, then compare it against Kling and Hunyuan without touching your integration.
Wan 2.1 API FAQs
What is Wan 2.1?
Wan 2.1 is Alibaba Cloud's multimodal video foundation model, known for realistic physics, complex motion, and rendering Chinese and English text inside generated video from text or image inputs.
Does Wan 2.1 API support text and image inputs?
Yes. You can generate from a written prompt or animate an image URL, so it fits both concept-first and asset-first video workflows.
Can Wan 2.1 render text inside a video?
Yes. It draws Chinese and English characters directly into the scene, keeping signs, titles, and captions readable as the shot moves.
What clip length and formats are available?
Wan 2.1 produces five-second clips and supports several aspect ratios, including landscape, portrait, and square, selected per request.
What is the difference between fast and pro mode?
Fast mode favors quick iteration for drafting, while pro mode spends more compute on refined detail for final renders.
Why run Wan 2.1 through Pollo API?
Pollo API gives one integration for Wan 2.1 and 300+ other models, with documentation, task tracking, and lower-cost generation than comparable providers.