
GPT Image 2
OpenAI's GPT Image 2 raw image model can generate high-quality images based on natural language prompts. It provides a ready-to-use REST inference API, offering excellent performance, no cold start, and affordability.
Read Me
OpenAI GPT Image 2 API
Introduction
GPT Image 2 is OpenAI’s advanced model for high-quality image generation and editing. It is designed for production workflows that require detailed instruction following, readable text, complex compositions, photorealistic detail, and controlled changes to existing images.
On iCreat, GPT Image 2 supports text-to-image generation, reference-based creation, and natural-language image editing through one model page. Start with a prompt to create a new image, or upload up to 16 reference images to guide the subject, product, style, scene, or edit. Choose from 1K, 2K, or 4K output sizes and Low, Medium, or High quality based on your visual requirements and budget. The model is available through both the online Playground and the iCreat API.
Key Features
High-quality image generation and editing
GPT Image 2 creates new visuals from natural-language prompts and edits existing assets through the same model. It can replace backgrounds, add or remove objects, adjust visual details, and build new compositions from references. This unified workflow reduces the need to switch between separate generation and editing tools.
Complex prompt control
The model can follow layered instructions covering subject, setting, composition, camera angle, lighting, materials, text, and elements that must remain unchanged. This makes it useful for advertisements, product scenes, and structured layouts where several visual requirements need to work together clearly within one controlled, production-ready final output.
Readable text and structured visuals
GPT Image 2 is well suited to images where text and layout are part of the final composition. It can create poster headlines, product labels, signs, interface text, infographic labels, diagrams, editorial layouts, and multi-panel explanations, reducing the amount of manual design and post-production work required afterward.
Up to 16 reference images
The iCreat endpoint accepts up to 16 reference images for reference-based creation or editing. Users can supply separate images for a product, person, packaging, clothing, lighting, background, or style. Clear instructions about each image’s role help the model combine visual information more deliberately, accurately, and consistently overall.
Flexible 1K, 2K, and 4K outputs
GPT Image 2 supports 1K, 2K, and 4K output tiers with Low, Medium, and High quality options. Users can test ideas at lower cost, produce standard assets at balanced settings, or generate larger final visuals when extra detail, cropping space, print readiness, or delivery resolution matters most.
How to Use
1. Choose a workflow. Enter a prompt without an image for text-to-image generation. Upload one or more references when you want to guide a new composition or edit an existing image.
2. Describe the result. State the subject, setting, composition, lighting, style, and exact text required. For edits, clearly separate what should change from what must remain intact.
3. Add references when needed. Upload up to 16 images and explain the role of each one, such as product shape, packaging colors, clothing, background, or lighting.
4. Select the output. Choose a supported 1K, 2K, or 4K size and set Low, Medium, or High quality. Use the Playground directly or call the openai/gpt-image-2 API.
Pricing
GPT Image 2 is priced per generated image based on output resolution and quality. Text-to-image generation, reference-based creation, and image editing use the same rates.
1K — From $0.03 per image
| Quality | Unit price |
|---|---|
| Low | $0.03 |
| Medium | $0.06 |
| High | $0.22 |
Supported sizes: 1024×1024 · 1280×720 · 720×1280 · 1152×864 · 864×1152 · 1248×832 · 832×1248 · 1120×896 · 896×1120 · 1456×624
2K — From $0.06 per image
| Quality | Unit price |
|---|---|
| Low | $0.06 |
| Medium | $0.12 |
| High | $0.44 |
Supported sizes: 2048×2048 · 2560×1440 · 1440×2560 · 2304×1728 · 1728×2304 · 2496×1664 · 1664×2496 · 2240×1792 · 1792×2240 · 3024×1296
4K — From $0.09 per image
| Quality | Unit price |
|---|---|
| Low | $0.09 |
| Medium | $0.18 |
| High | $0.66 |
Supported sizes: 2880×2880 · 3840×2160 · 2160×3840 · 3264×2448 · 2448×3264 · 3504×2336 · 2336×3504 · 3200×2560 · 2560×3200 · 3696×1584
Model Comparison
GPT Image 2 vs Nano Banana Pro
Choose between them based on reference capacity, prompt length, output controls, and pricing structure, since both target professional image generation and editing.
| What matters | GPT Image 2 | Nano Banana Pro |
|---|---|---|
| Best for | Text-heavy, complex briefs | Product and design workflows |
| Text and layout | Strong | Strong |
| Image editing | Natural-language editing | Reference-based editing |
| Reference images | Up to 16 | Up to 11 |
| Prompt limit | 32,000 characters | 8,192 characters |
| Output control | Exact size + quality | Ratio + resolution |
| Lowest price | $0.03 | $0.14 |
| 2K price | $0.06–$0.44 | $0.14 |
| 4K price | $0.09–$0.66 | $0.24 |
| Budget flexibility | Higher | Lower |
| Price predictability | Varies by quality | Fixed by resolution |
Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
prompt |
string | Yes | Image instruction with a maximum length of 32,000 characters |
image |
string[] | No | One to 16 reference-image URLs for reference generation or editing |
size |
string | Yes | Output dimensions selected from supported 1K, 2K, or 4K sizes |
quality |
string | Yes | low, medium, or high; controls quality, detail, and price |
| Specification | Value |
|---|---|
| Model ID | openai/gpt-image-2 |
| Input | Text and optional images |
| Workflows | Text-to-image, reference generation, image editing |
| Reference-image limit | 16 |
| Resolution tiers | 1K, 2K, 4K |
| Quality levels | Low, Medium, High |
| Billing unit | USD per image |
| API pattern | Asynchronous task submission |
Use Cases
Posters and text-heavy advertising
Create campaign visuals that combine headlines, supporting copy, branded colors, and clear visual hierarchy, keeping written content inside the original composition instead of adding it later in a separate design tool.
Product visuals from multiple references
Combine product, packaging, lighting, background, and style references in one request to create commercial visuals that reflect several existing assets without relying only on a long written description.
Background and scene editing
Keep the main subject while replacing the setting, lighting, props, or surrounding environment, such as moving a studio product into a lifestyle scene or converting daylight into night.
Infographics and visual explanations
Generate diagrams, labeled processes, multi-panel explanations, and visual summaries by describing the information order, panel structure, labels, and relationships between elements in the final composition.
UI and editorial concepts
Create early interface ideas, landing-page visuals, magazine-style layouts, presentation graphics, and editorial concepts that teams can evaluate before rebuilding the selected direction in a production design tool.
Creative asset variations
Adapt one concept into square, vertical, horizontal, and ultrawide versions while testing different environments, lighting choices, styles, or supporting elements around the same central subject.
Prompt Tips
Structure the request clearly. Describe the subject, scene, composition, lighting, materials, exact text, style, and intended use in a consistent order.
Quote required text. Put exact wording in quotation marks and state where it should appear, especially for posters, labels, interfaces, and infographics.
Assign each reference a role. Tell the model which image defines the product, person, packaging, background, lighting, or style instead of uploading references without guidance.
Separate edits from preserved details. State what should change and what must remain unchanged, and remove conflicting style or composition instructions.
Notes
- The 16-image limit does not mean every reference will be preserved or weighted equally.
- Resolution and quality are separate controls; selecting 4K does not automatically select High quality.
- The current iCreat endpoint does not expose mask-based or multi-turn conversational editing.
- Important text, logos, packaging details, and factual content should be checked before publication.
Related Models
A professional image model for design-heavy workflows, product mockups, reference-based creation, and fixed pricing by resolution.
A more efficiency-focused option for users who prioritize faster iteration and lower-cost image production over Pro-level controls.
