curl-X POST 'https://pollo.ai/api/platform/v1/generation/openai/gpt-image-2/image'\-H'Content-Type: application/json'\-H'x-api-key: YOUR_API_KEY'\-d'{ "input": { "prompt": "A cinematic shot of a golden retriever running through a field of sunflowers at sunset, warm rim light, shallow depth of field", "aspectRatio": "1:1", "resolution": "1K", "quality": "low", "background": "auto" }}'
Optionally include webhookUrl to receive a callback when the task succeeds or fails.
Optional request field
{"input":{"prompt":"A cinematic shot of a golden retriever running through a field of sunflowers at sunset, warm rim light, shallow depth of field","aspectRatio":"1:1","resolution":"1K","quality":"low","background":"auto"},"webhookUrl":"https://example.com/webhooks/pollo"}
Schema
Input fields for the selected endpoint and mode.
Field
Type
Required
Description
prompt
string
Yes
Text prompt describing the content to generate.
aspectRatio
string
No
Output aspect ratio, as width:height (e.g. 16:9).
resolution
string
No
Output resolution of the generated media (e.g. 720p, 1080p, 2K, 4K). Values are case-insensitive.
quality
string
No
Output quality tier of the generation.
background
string
No
Background of the generated image: auto lets the model decide, transparent produces an alpha background (PNG), opaque forces a solid background.
Legible text inside an image is where GPT Image 2 separates itself. Headlines, labels, and small print render cleanly, so posters, packaging, and diagrams come out ready to use.
With GPT Image 2 on Pollo API, developers add text-to-image, image editing, and multi-image composition to their apps, generating up to 4K output with aspect ratios chosen per request.
Key Features of GPT Image 2 API
Reliable In-Image Text
Typography holds together across headlines, captions, and dense labels, making the model dependable for menus, infographics, and any layout where wording has to be readable, not decorative filler.
Extreme Instruction Following
Long, detailed prompts are respected closely, so specific object counts, placements, colors, and written phrases land as described instead of drifting into a loose interpretation.
Pixel-Level Editing
Feed an existing image and change a region precisely, swapping objects, adjusting text, or refining details while the untouched parts of the frame stay intact.
Grounded World Knowledge
Scenes reflect real relationships between objects, materials, and settings, so a described tool, dish, or storefront looks plausible rather than assembled from mismatched parts.
4K Output
Generations scale to 4K, giving enough resolution for print layouts, large banners, and detailed product shots without a separate upscaling step in the workflow.
Multi-Image Composition
Combine several reference images in one request to merge subjects, place a product into a new scene, or build a composite that keeps each source recognizable.
Use Cases of GPT Image 2 API
Packaging and Label Mockups: Generate boxes, bottles, and pouches with legible brand names, ingredient lists, and small print positioned as specified in the prompt.
Infographics and Diagrams: Produce charts, labeled schematics, and step layouts where the text and structure carry the meaning, not just the visuals.
Menu and Signage Boards: Create restaurant menus, price boards, and shop signage with accurate wording and consistent typography for hospitality and retail tools.
Targeted Photo Edits: Let users select a region and change one element, such as recoloring a garment or replacing a logo, while preserving the rest.
Product Compositing: Drop a product from one image into a new setting, keeping its shape and finish for ecommerce catalog and lifestyle imagery.
Poster and Ad Layouts: Build promotional layouts where a headline, subhead, and call to action must appear exactly as written and stay readable.
App and UI Concepts: Render interface mockups and screen ideas with real button labels and menu text for design and prototyping apps.
Localized Asset Variants: Regenerate the same layout with different text to produce language or campaign variants at scale.
How to Use GPT Image 2 API
Get an API Key: Create a Pollo API account and generate your API key from the developer dashboard.
Choose the Model: Send requests to the GPT Image 2 endpoint at /generation/openai/gpt-image-2-0/image.
Add Your Inputs: Pass a prompt for text-to-image, or supply imageUrl or images for editing and composition, then set aspectRatio, resolution (1K, 2K, 4K), and quality.
Generate and Retrieve: Submit the task, poll the returned taskId and status, and download the image once it succeeds.
Prompting Best Practices for GPT Image 2 API
Write the prompt like a spec sheet. Quote the exact text you want rendered, state where elements sit, and name colors and materials so instruction following has clear targets.
A Simple Prompt Formula
Subject and setting + exact text in quotes + layout and placement + style and color + material or lighting detail
What You Should Notice
Quote The Text: Put wording you want rendered in quotation marks so the model treats it as literal copy, not a loose theme.
Name Placement: Say where each element belongs, such as top banner or lower-left corner, to control layout.
Describe The Edit Region: For image editing, specify exactly what to change so untouched areas stay as they were.
Set Aspect Ratio Early: Pick the ratio that fits the final format so composition is framed correctly from the start.
Match Resolution To Use: Choose 4K for print and large formats, lower resolutions for faster drafts and iteration.
Keep One Focus Per Edit: Change a single element per editing pass to keep results predictable and clean.
Example Prompts
Product Packaging
"A matte kraft coffee bag standing on a light wooden shelf, front label reading 'MORNING ROOT — Single Origin' in bold serif, with 'Net 340g' in small text at the base. Soft daylight, shallow depth, 4K product shot."
Labeled Infographic
"A clean flat-style infographic titled 'How Composting Works' with four numbered steps arranged left to right, each with a simple icon and a short caption. Muted green and beige palette, crisp readable sans-serif text."
Region Edit
"Using the uploaded storefront photo, replace the blank hanging sign with one reading 'AZURE BOOKS' in warm hand-painted lettering. Keep the building, awning, and street unchanged, matching the original daylight and shadows."
Multi-Image Composite
"Combine the provided sneaker and the studio backdrop image so the sneaker sits centered on a reflective pedestal, spotlight from upper left, subtle floor reflection, keeping the shoe's exact colorway and stitching detail."
GPT Image 2 vs Google Imagen 4 vs FLUX.1 Kontext
Capability
GPT Image 2
Google Imagen 4
FLUX.1 Kontext
In-image text rendering
✅ Strong on long and small text
✅ Good typography
⚠️ Editing-focused, weaker text
Instruction following
✅ Handles detailed multi-element prompts
✅ Solid adherence
✅ Strong for edit instructions
Pixel-level editing
✅
⚠️ Limited
✅ Core strength
Multi-image composition
✅
❌
⚠️ Reference-based
Max resolution
Up to 4K
Up to 2K class
Up to 4MP class
Recommended For
Text-heavy assets and precise edits
Photoreal generation
Iterative image editing
Why Choose GPT Image 2 API?
GPT Image 2 fits products that need readable text and faithful edits, from packaging and infographics to targeted region changes, all at resolutions high enough for print.
Through Pollo API, you reach GPT Image 2 with one API key alongside 300+ leading image and video models, with clean docs, task status polling, and generation that costs less than Fal AI.
Start with the GPT Image 2 endpoint on Pollo API to ship text-accurate image features, then compare it against Imagen, FLUX, and others without touching your integration.
GPT Image 2 API FAQs
What is GPT Image 2?
GPT Image 2 is an OpenAI image model known for accurate in-image text, strong instruction following, and pixel-level editing, supporting text-to-image, image editing, and multi-image composition up to 4K.
Does GPT Image 2 API render text reliably?
Yes. It is one of its defining strengths, handling headlines, labels, and small print, which makes it well suited to packaging, menus, posters, and infographics.
Can it edit an existing image?
Yes. Provide an image and describe the change, and the model edits a targeted region while keeping the rest of the frame intact.
What resolutions and aspect ratios are supported?
Output scales to 1K, 2K, or 4K, with aspect ratios from square and portrait to widescreen and ultrawide chosen per request.
Can it combine multiple images?
Yes. You can pass several reference images in one request to merge subjects or place a product into a new scene while keeping each source recognizable.
Why run GPT Image 2 through Pollo API?
Pollo API gives one integration for GPT Image 2 and 300+ other models, with clean documentation, task tracking, and lower-cost generation than comparable providers.