Nano Banana Pro Text-to-Image

google/gemini-3-pro-image/text-to-image
OfficialText-to-Image

Google Nano Banana Pro (Gemini 3.0 Pro Image) supports text-to-image and can output 4K resolution results. It offers a ready-to-use REST inference API, excellent performance, no cold start, and affordable price.

Read Me

Google Gemini 3 Pro Image Text-to-Image

Introduction

Gemini 3 Pro Image Text-to-Image is Google’s professional endpoint for generating images directly from written prompts. It is suited to complex creative instructions, readable text, structured layouts, and high-resolution visual production. The endpoint focuses on creating new images from text rather than combining generation with reference-image editing in the same workflow.

On iCreat, users can enter prompts of up to 8,192 characters, choose from ten common aspect ratios, and generate at 1K, 2K, or 4K resolution. The endpoint uses fixed resolution-based pricing and works well for posters, advertisements, product concepts, menus, information graphics, and presentation assets through the Playground or the google/gemini-3-pro-image/text-to-image API.

Key Features

Professional text-to-image generation

Gemini 3 Pro Image Text-to-Image turns detailed written briefs into polished visuals without requiring reference images. It is suited to professional posters, product concepts, campaign graphics, and presentation assets where the model must coordinate subject, setting, composition, lighting, style, and final delivery format inside one complete image.

Readable text and structured layouts

The model is useful when text and layout are part of the generated result, including headlines, labels, menus, signs, infographics, and editorial graphics. Clear instructions about wording, placement, hierarchy, and spacing help it organize written content and visual elements inside one coherent composition for real publishing workflows.

Complex prompt understanding

Gemini 3 Pro Image is designed for complex professional instructions rather than short subject-only prompts. Users can define camera angle, materials, lighting, visual hierarchy, exact text, brand colors, negative space, and intended use, giving the model more context for producing a deliberate result that matches a detailed creative brief.

Flexible ratios and 4K output

The iCreat endpoint supports ten common aspect ratios and 1K, 2K, or 4K output. This makes it practical to prepare square posts, vertical ads, landscape banners, presentation graphics, and ultrawide visuals from the same model while selecting the resolution that fits drafting, publishing, or final high-resolution delivery.

Simple resolution-based pricing

Pricing is fixed by resolution with no separate quality tier: $0.07 for 1K, $0.105 for 2K, and $0.14 for 4K. The structure is easy to forecast for repeated production, and the moderate step between tiers lets teams choose more output space without introducing several additional quality and cost combinations.

How to Use

1. Choose the workflow

This endpoint is dedicated to text-to-image generation. Enter a written prompt to create a new visual from scratch, without uploading reference images or switching to a separate generation workflow at any stage.

2. Describe the result

Define the subject, setting, composition, lighting, style, exact text, and intended use. Clear priorities help the model organize several visual requirements inside one coherent image instead of treating them as unrelated details.

3. Select ratio and resolution

Choose the aspect ratio for the final publishing channel, then select 1K, 2K, or 4K based on whether the image is for drafting, online use, cropping, presentation, or final delivery.

4. Generate the image

Create the image directly in the Playground or submit a request through google/gemini-3-pro-image/text-to-image. The API returns a task ID that you can poll before retrieving the completed image result.

Pricing

Gemini 3 Pro Image Text-to-Image uses fixed per-image pricing based on output resolution, with no additional quality tier.

Resolution Unit Price
1K $0.07/image
2K $0.105/image
4K $0.14/image

The 4K tier costs only $0.035 more than 2K, making it practical when the workflow needs more cropping room, higher display clarity, or a larger final delivery size.

Model Comparison

Gemini 3 Pro Image Text-to-Image vs Nano Banana Pro Endpoint

Both endpoints use the same underlying model, but the current iCreat implementations separate direct text-to-image generation from the complete reference-image and editing workflow.

What matters Gemini 3 Pro Image Text-to-Image Nano Banana Pro
Best for Direct text-to-image Product and design workflows
Input method Prompt only Prompt + references
Editing workflow Not supported Prompt + references
Reference limit Not supported Up to 11
Prompt limit 8,192 characters 8,192 characters
Output control Ratio + resolution Ratio + resolution
Quality tiers None None
1K price $0.07 $0.14
2K price $0.105 $0.14
4K price $0.14 $0.24
Budget planning Lower text-to-image cost More complete workflow

Gemini 3 Pro Image Text-to-Image vs GPT Image 2

Gemini 3 Pro Image Text-to-Image uses simple resolution-based pricing, while GPT Image 2 adds reference inputs, image editing, and tiered quality controls.

What matters Gemini 3 Pro Image Text-to-Image GPT Image 2
Best for Professional text-to-image Complex generation and editing
Text and layout Strong Strong
Input method Prompt only Prompt + references
Editing workflow Not supported Prompt + references
Reference limit Not supported Up to 16
Prompt limit 8,192 characters 32,000 characters
Output control Ratio + resolution Exact size + quality
Quality tiers None Low / Medium / High
1K price $0.07 $0.03–$0.22
2K price $0.105 $0.06–$0.44
4K price $0.14 $0.09–$0.66
Budget planning Fixed by resolution Flexible by quality

Parameters

Parameter Type Required Description
prompt string Yes Image prompt with a maximum length of 8,192 characters
aspect_ratio string Yes Selects 1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, or 21:9
image_size string Yes Selects 1K, 2K, or 4K output
Specification Value
Model ID google/gemini-3-pro-image/text-to-image
Workflows Text-to-image
Prompt limit 8,192 characters
Resolution tiers 1K, 2K, 4K
Aspect ratios 1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9
Quality tiers None
Billing unit USD per image
API pattern Asynchronous task submission

Use Cases

Posters and campaign graphics

Generate posters, launch graphics, event visuals, and promotional assets that combine a clear headline, central subject, supporting elements, branded styling, and deliberate visual hierarchy in one finished composition.

Product concept visuals

Turn written product descriptions into polished concept images for presentations, landing pages, campaign exploration, packaging directions, product positioning, or early commercial mockups before photography and final production begin.

Menus, signs, and information graphics

Create menus, signs, announcements, simple infographics, and labeled visuals where readable wording, balanced spacing, icons, visual hierarchy, and supporting imagery must work together clearly for the audience.

Social and advertising formats

Produce square, vertical, landscape, and ultrawide assets for social posts, stories, display ads, banners, campaign landing pages, and channel-specific placements without changing to a different image model.

Presentation and editorial covers

Generate presentation covers, article headers, report visuals, and editorial concepts that need a strong title area, polished composition, supporting imagery, and clear space for additional information or branding.

High-resolution final assets

Use 4K output when the final image needs closer cropping, larger placement, high-resolution presentation, or additional visual detail for downstream design, campaign adaptation, printing, and final delivery.

Prompt Tips

Set the information priority first. Identify the main subject, required text, and visual focus so several instructions do not compete for attention.

Write exact image text. Put headlines, labels, or slogans in quotation marks and specify their position, size, and hierarchy.

Describe composition for the selected ratio. Vertical, square, and landscape formats need different layouts, so choose the ratio before defining subject placement and negative space.

State the final use. Tell the model whether the image is for a poster, social post, advertisement, or presentation to guide the overall composition.

Notes

  • The API documentation includes a reference-image example with an image field, but the current schema does not define that parameter; do not rely on that example until the endpoint contract is updated.
  • aspect_ratio is required in the current schema, so text-to-image requests should pass it explicitly even though the description mentions a default based on an uploaded image.
  • The schema lists 1K, 2K, and 4K, while the request example uses lowercase 1k; confirm the accepted casing before production integration.
  • The current endpoint does not expose a visible watermark setting, but images generated by Gemini image models include an invisible SynthID watermark.

Nano Banana Pro

Choose it when you need up to 11 reference images, image editing, and a more complete professional design workflow.

Nano Banana 2

Choose it for more reference capacity, everyday image production, and a balanced mix of speed, cost, and capability.

GPT Image 2

Choose it for longer prompts, up to 16 reference images, exact output dimensions, and tiered quality controls.